Research
Research for faster model improvement
Benchmarks, papers, and field notes on where models fail, how to measure the gap, and what it takes to move the frontier.
Research areasUser SimulationVerifiers & RewardPost-Training & DistillationAdversarial RobustnessLong-Horizon AgentsCybersecurity
Benchmarks
Evaluations we build to find where models fail
Publications
Papers, preprints and collaborations with frontier researchers
Field Notes
Recent technical writing from our team


Put the research to work
Turn a failure surface into
measurable model improvement.
Bring us the capability gap. We’ll shape the tasks, environments, and evaluation program needed to measure it—and improve it.