Safety at Curos
We build a product people can trust, which means we spend a large share of our effort on the boring parts: measuring behavior, checking our work, and making it possible for people outside Curos to check us too.
What safety means here
When we say a model is "safe enough to release," we mean something specific: that we have tested it against a set of published criteria, that the Safety & Evaluations Committee signed off, and that we can describe what we did not test and why. We do not mean that risk is zero. Anyone who promises that is trying to sell you something.
Safety at Curos rests on four pillars, each of which is someone's full-time job:
- Measure. An evaluation suite that every model must pass before release — capability, safety, and honesty checks. This is our approach to model evaluations.
- Understand. Interpretability research that tries to explain what models are doing, not just test what they output.
- Oversee. Independent committee sign-off and published release reports. No model ships on a single person's say-so.
- Listen. A community that can report problems, and a commitment to treat reports seriously — including when they are inconvenient.
Release gates
Before any model is deployed to Ichnus users, it has to clear the same set of gates. The gates are not optional, and the Safety & Evaluations Committee has the power to delay a release. In practice it has delayed two.
1 · Evaluation suite
The full evaluation suite passes: capability benchmarks within published bounds, safety evals clear, red-teaming rounds complete with findings triaged.
2 · Interpretability review
The interpretability team reviews anything the model does that its training data cannot obviously explain. Unexplained capability at meaningful risk is a blocker.
3 · Committee sign-off
The Safety & Evaluations Committee meets, reviews the evaluation report, and signs off or asks for more work. Its minutes are published.
4 · Public report
Alongside the release, we publish the evaluation report — what we tested, what we found, and what we did not test. The report, not the press release, is the announcement.
Incident response
We track incidents the way we track everything else: in the open, with a postmortem. If a model does something harmful, that becomes a published incident record with the trigger, the fix, and the change to our process. The status page covers availability; incidents about model behavior get reported to the Safety Committee and summarized in the newsroom.
Limits of this page
Curos is a concept project. There is no real evaluation suite, no real committee, and no real product being gated. This page is a design proposal for how an accountable AI organization could run itself.
Read about our approach to evaluations →