Why evaluations are the backbone

A claim like "this model is safe" is not a feeling; it is a set of measurements, published so others can disagree with them. Evaluations are how Curos turns confidence into something checkable. They are also how we admit what we do not know: an evaluation report that lists its own gaps is worth more than one that does not.

The four layers of the suite

1. Capability benchmarks

We measure what a model can do across knowledge, reasoning, math, code, and long-context work. These are the familiar-looking numbers you see on the model pages. They tell us where the model stands, not whether it is safe — capability is a precondition for risk, not a substitute for measuring it.

2. Safety evals

We test for behavior we have committed not to ship: following instructions that would cause real harm, providing reliable information on dangerous capabilities, or behaving deceptively. These evals are adversarial by construction — built on the assumption that the model will fail somewhere, and the job is to find where.

3. Honesty evals

The most subtle layer. We measure whether a model overstates its confidence, invents citations, or refuses to say "I don't know." Because users cannot verify every answer, honesty is a safety property. Our internal OversightQA and Curos-Helpfulness evals live here.

4. Red-teaming

Automated evals miss what creative humans find. Each release cycle includes red-teaming rounds run by both internal testers and external reviewers — researchers, policy people, and community members who are explicitly trying to make the model fail. Their findings are triaged like bug reports, and high-severity findings block release.

What passes and what blocks

Evals do not produce a single pass/fail. Findings are triaged by severity and context: a model that is unhelpful in one niche area might still ship; a model that is deceptively sycophantic in a high-stakes domain does not. The Safety & Evaluations Committee makes the call, in the open.

Publishing the results

We publish the evaluation report alongside every release: the suite, the numbers, the red-team findings, and — explicitly — what we did not test and why. We have published a field guide to the whole suite, and the models pages link to each model's report.

Limitations

Evaluations cannot prove safety. They can only bound our confidence. New behaviors appear after deployment, which is why the community can report problems and why incidents get published postmortems. An evaluation program that admits this is doing its job.

Curos is a concept project. There is no real evaluation suite, no real committee, and no real report. This page is a proposal for what transparent model evaluation could look like.

Read the research behind our evaluations →