Independent research lab · San Francisco

Capability is not
a straight line.

We study how reasoning, generalization, and failure emerge in large models — the discontinuities that show up between one checkpoint and the next. We publish what we find, including the results that complicate our own work.

Run 4417 · 12B · reasoning eval step 0 loss
Research

Four questions we keep returning to

Our agenda is narrow on purpose. Each thread has a team that owns it end to end, from the training run to the paper.

Interpretability

Reading circuits, not outputs

We reverse-engineer the internal structure of trained models — features, circuits, and the representations that survive scale — so that claims about model behavior can be checked against mechanism rather than inferred from benchmarks.

Reasoning

Where chains of thought break

Long-horizon reasoning fails in structured, reproducible ways. We build evaluations that isolate the step where a trajectory goes wrong, and study which training interventions actually move that step.

Scaling

Discontinuities in the curve

Some capabilities arrive gradually and some arrive at once. We characterize which is which, and look for the measurable precursors that appear in a run before a capability does.

Alignment

Oversight that scales past us

As models take on tasks their supervisors cannot fully check, evaluation becomes the bottleneck. We work on protocols that stay reliable when the system under review is stronger than the reviewer.

Approach

How work moves through the lab

01

Instrument first

Before a question becomes a training run, it becomes a measurement. If we cannot say what result would change our mind, the project is not ready to start.

02

Run small, then scale

Findings are reproduced across at least three model scales before we believe them. Most results that look clean at 1B do not survive the second replication.

03

Publish the negatives

Failed replications and null results go out with the same care as headline findings. The field is slower when only the wins are legible.

31
Researchers
19
Papers since 2024
6
Open-source releases
100%
Findings published
Papers

Recent publications

2026 · 07
Precursor signals precede capability jumps in reasoning models
Preprint
2026 · 05
Feature stability across checkpoints: a replication at three scales
ICML 2026
2026 · 03
Where chain-of-thought fails: step-level attribution for long trajectories
Preprint
2025 · 11
Null result: sparse probes do not recover planning circuits at 12B
Technical report
2025 · 09
Protocols for evaluating systems stronger than their evaluators
NeurIPS 2025