Most people think of interpretability as a research topic, which it is. At Curos it is also part of the release process: before a model ships, the interpretability team has to answer a question the evaluation suite cannot — does the model do anything we cannot explain?

That framing changes what we work on. It means our roadmap is driven less by what is intellectually interesting and more by what the release gates need. Here is where we are.

What shipped this year

Two lines of work made it into production.

Refusal circuits. We identified a small, localized set of features that implement refusal in our models — a sparse circuit we can watch. That gives us a mechanistic check alongside the behavioral one: a model release can be reviewed not just for what it does, but for whether its refusal behavior has a structure we understand. The work is on the research index.

Feature trajectories. Our attention-geodesics work showed that features move along paths across layers, and that path shape carries information single-point analyses miss. We are turning this into a diagnostic: unstable trajectories flag features worth a second look before release.

What we are building next

Three things, in rough order.

A monitoring layer. The refusal circuit is a start, but we want a standing set of “vital signs” — interpretable circuits that are cheap to watch on live traffic, so a release doesn’t just pass its checks at launch but keeps passing them. This is the bridge between interpretability as research and interpretability as operations.

Interpretability as a community discipline. Our public red-teaming work taught us that outsiders find things we miss. We are opening the same door to interpretability: published tooling, reproducible analyses, and a standing invitation for researchers outside Curos to test our claims against our models.

Honesty about scale. Sparse autoencoders on frontier-scale models are approximate, and our dictionaries are a choice, not a ground truth. The roadmap includes getting better at saying what our tools do not see.

Why this matters

Interpretability is not a nice-to-have at Curos. It is one of the four pillars of safety, alongside measurement, oversight, and listening. A model whose behavior we can watch, not just test, is a model we can govern. The roadmap is how we keep widening the part we can watch.

All news →