Turn a benchmark ceiling into a renewable training surface.
Each family starts from a capability contract, then expands into original tasks, executable environments, graders, trajectories, failure labels, repairs, and private holdouts. Pages distinguish supplied evaluation evidence from delivery blueprints.
GDPval-AA v2
Professional work products across 44 occupations: source packs, workbooks, reports, decision artifacts, rubrics, and reviewer evidence.
τ³-Banking
Stateful customer and tool interactions with banking policy, hidden account state, exact actions, and escalation boundaries.
Terminal-Bench v2.1
Containerized terminal worlds for debugging, recovery, data processing, systems work, and safe state-changing execution.
SciCode
Research-shaped scientific programming tasks with staged subproblems, reproducible numerical outputs, and test-backed code.
Humanity’s Last Exam
Broad expert reasoning across text and multimodal prompts, with calibrated answers, deterministic grading, and auditable derivations.
GPQA Diamond
Graduate-level biology, chemistry, and physics question families with difficult distractors and dense reasoning feedback.
CritPt
Research-level physics programming checkpoints with multi-attempt RL, public/hidden/stress tests, and repairable failure traces.
AA-Omniscience
Factual recall, confidence, abstention, source discipline, and hallucination-aware rewards across economically relevant domains.
AA-LCR
Long-document worlds that require evidence retrieval, cross-document joins, contradiction handling, and claim provenance.
Ulam Math Data
Exact-answer AIME-style RLVR, graduate and research Harbor environments, and open research problems for rollout-based post-training.