A field guide to our evaluation suite
A practical how-to for reading our evaluation reports and benchmark tables — what the numbers mean and where they come from.
We publish an evaluation report with every model release, and the model pages carry benchmark tables. They are meant to be readable, not impressive, but “readable” still takes a little translation. This is the field guide.
Where the numbers come from
Every number in a Curos evaluation report comes from a run we can reproduce, with the suite version and the configuration recorded. If you cannot tell where a number came from, we have failed at documentation, and you should say so.
Reading the benchmark tables
The model pages show scores on public benchmarks — MMLU, GPQA, MATH, HumanEval — plus two internal evaluations, OversightQA and Curos-Helpfulness. A few rules of thumb:
- Higher is better, but not by a lot. Differences of a point or two on a public benchmark are usually within noise. We do not make release decisions on benchmark deltas that small, and you should not either.
- Internal evals are the more interesting ones. OversightQA measures whether the model can be overseen — how detectable its errors are when someone reviews its work. Curos-Helpfulness measures whether answers are actually useful, not just correct. These are closer to the experience of using the product than any public benchmark.
- The note is part of the number. Every table on this site carries a note that the figures are illustrative and part of a concept project. That is not a disclaimer at the bottom of the page; it is the metadata of the whole site.
Reading the reports
A release report has a fixed shape:
- What we tested — the suite, the version, the configuration.
- What we found — the numbers, plus the red-team findings and how we triaged them.
- What we did not test — the honest section. Every report lists gaps, because a report with no gaps has not been read carefully.
Start with the third section. It is the one that tells you how much to trust the first two.
Reproducing our work
For researchers, the evaluation harness is published with each report, and the contributing guide explains how to run it against our models. If a number does not reproduce, that is a bug in our reporting, and we want the issue filed.
The short version
Treat every number like a question, not an answer. “What did they test, how did they test it, and what did they skip?” The numbers are the beginning of the conversation — the report is the whole of it.
All news →