Best for

  • Deep research and synthesis
  • Complex, multi-step reasoning
  • Code, math, and high-stakes decisions
  • Working across very long documents

Why this tier exists

Corus exists for the questions that are worth waiting for: a research synthesis across dozens of sources, a hard proof, a codebase-wide refactor, a decision with real consequences if it goes wrong. It gets more time to think, and it uses it.

The fair-use limits are the tradeoff for making the frontier tier free rather than metered. We would rather cap how much any one person can use for free than put a price tag on the model that matters most for hard problems.

Benchmarks

Scores on public benchmarks and two internal evaluations. Figures are illustrative — see the note below.

Evaluation Corus Krus Mantus Corus
MMLU 88.2 78.9 85.6 88.2
GPQA (diamond) 76.4 58.2 67.8 76.4
MATH (competition) 79.1 62.4 71.0 79.1
HumanEval 89.3 76.5 84.1 89.3
OversightQA (internal) (ours) 84.6 71.8 79.4 84.6
Curos-Helpfulness (internal) (ours) 88.7 80.2 86.9 88.7

Benchmark scores are illustrative and invented for this concept project. Curos is not a real organization and these are not measurements of a real model.

Model card

The sections below mirror the model-card format we publish for every release. For Corus, the release report is published alongside the evaluation suite it passed.

Overview

  • Version: Corus 1, released July 2026
  • Context window: 1M tokens
  • Availability: available to all Ichnus users, free of charge
  • Safety classification: consumer general-purpose assistant, not for high-stakes uses without review

Training and data

Corus was trained on a curated corpus, with a strong emphasis on removing low-quality and duplicated data. We publish an outline of our data practices in the model documentation and the details behind each release in its evaluation report.

Known limitations

  • May confidently answer outside its knowledge — we are working on honesty, not pretending it is solved
  • Not a substitute for professional advice in medicine, law, or finance
  • Can produce plausible but wrong code; always review generated code

Safety evaluations

Before release, Corus passed our full evaluation suite — capability, safety, honesty, and red-teaming — as reviewed by the Safety & Evaluations Committee. The public report lists what we tested, what we found, and what we did not test. See our approach to evaluations for how the suite works.

Where to use it

  • Strongest reasoning and research capability
  • Best for research, code, and high-stakes decisions
  • Available free, with fair-use limits
  • 1M-token context window
All models →