What safety means here

When we say a model is "safe enough to release," we mean something specific: that we have tested it against a set of published criteria, that the Safety & Evaluations Committee signed off, and that we can describe what we did not test and why. We do not mean that risk is zero. Anyone who promises that is trying to sell you something.

Safety at Curos rests on four pillars, each of which is someone's full-time job:

  • Measure. An evaluation suite that every model must pass before release — capability, safety, and honesty checks. This is our approach to model evaluations.
  • Understand. Interpretability research that tries to explain what models are doing, not just test what they output.
  • Oversee. Independent committee sign-off and published release reports. No model ships on a single person's say-so.
  • Listen. A community that can report problems, and a commitment to treat reports seriously — including when they are inconvenient.

Release gates

Before any model is deployed to Ichnus users, it has to clear the same set of gates. The gates are not optional, and the Safety & Evaluations Committee has the power to delay a release. In practice it has delayed two.

Incident response

We track incidents the way we track everything else: in the open, with a postmortem. If a model does something harmful, that becomes a published incident record with the trigger, the fix, and the change to our process. The status page covers availability; incidents about model behavior get reported to the Safety Committee and summarized in the newsroom.

Limits of this page

Curos is a concept project. There is no real evaluation suite, no real committee, and no real product being gated. This page is a design proposal for how an accountable AI organization could run itself.

Read about our approach to evaluations →