Four questions we keep returning to
Our agenda is narrow on purpose. Each thread has a team that owns it end to end,
from the training run to the paper.
Interpretability
Reading circuits, not outputs
We reverse-engineer the internal structure of trained models — features,
circuits, and the representations that survive scale — so that claims about
model behavior can be checked against mechanism rather than inferred from
benchmarks.
Reasoning
Where chains of thought break
Long-horizon reasoning fails in structured, reproducible ways. We build
evaluations that isolate the step where a trajectory goes wrong, and study
which training interventions actually move that step.
Scaling
Discontinuities in the curve
Some capabilities arrive gradually and some arrive at once. We characterize
which is which, and look for the measurable precursors that appear in a run
before a capability does.
Alignment
Oversight that scales past us
As models take on tasks their supervisors cannot fully check, evaluation
becomes the bottleneck. We work on protocols that stay reliable when the
system under review is stronger than the reviewer.