GDPval-AA v2
Professional work-product data built from source-grounded briefs that end in editable documents, spreadsheets, slides, models, and decisions.
Benchmark-shaped, training-ready data for agents and reasoning models that need to act, recover, explain, and verify—not merely imitate.
Each family converts an evaluation-shaped capability into renewable original tasks, observable trajectories, critiques, repairs, and a sealed assessment boundary.
Professional work-product data built from source-grounded briefs that end in editable documents, spreadsheets, slides, models, and decisions.
Stateful banking and customer-service simulations with policy constraints, hidden account state, multi-turn interaction, and exact action checks.
Long-horizon terminal tasks in containers with real filesystem and service state, hidden tests, recovery branches, and complete execution traces.
Research-code tasks spanning implementation, debugging, numerical experiments, and scientific libraries, with tests and result-level checks.
Broad expert-level reasoning and knowledge tasks with adversarial distractors, auditable solutions, and carefully sealed answer keys.
Graduate-level science question families with difficult distractors, explanation scoring, and expert-reviewed provenance.
Physics problem-solving and critique data focused on derivations, assumptions, units, counterexamples, and repairable reasoning traces.
Knowledge-reliability data balancing correct answers, calibrated abstention, source checks, uncertainty labels, and hallucination penalties.
Long-context reasoning corpora with distributed evidence, contradiction tracking, retrieval traps, structured citations, and answer provenance.
AIME-style exact-answer RLVR, graduate and research Harbor environments, and open research problems for multi-rollout post-training.
Every task starts with a capability contract and ends with evidence that can survive execution, review, replay, and adversarial checking.
Start from a failure surface, not a prompt genre. Specify what success changes in the world.
Create structural variants with new state, evidence, tools, constraints, and adversarial branches.
Run code, rebuild worlds, inspect artifacts, grade criteria, and retain the full trajectory.
Separate training packs, repair pairs, pilot sets, and private holdouts with traceable provenance.
The output is a renewable data product: tasks, environments, source packs, event traces, labels, verifier bundles, and a changelog.
Inspect every environmentVolume matters only after novelty, solvability, verifier coverage, and contamination boundaries are measurable.
Change mechanisms, state, evidence, tool topology, and failure opportunities—not only wording.
Prefer executable assertions, reproducible artifacts, and criterion-level review to opaque judge scores.
Preserve diagnosis, partial progress, critique, repair, and the earliest consequential mistake.
Separate prompts, source metadata, graders, and exposure logs from the training path.
Evaluate the document, codebase, environment state, proof object, decision, or tool action the model leaves behind.
A reusable editorial layer for dataset releases, methodology notes, verifier audits, and benchmark-inspired research.
How to turn terminal capability into renewable worlds, state contracts, hidden checks, failure labels, and repair trajectories.
Read field noteExact-answer AIME-style RLVR, graduate and research Harbor environments, and open problems for multi-rollout training.
Read field noteSee which programs include supplied blind runs and which describe the complete delivery structure without a model score.
Browse all pagesWe provide high-quality off-the-shelf datasets, RL tasks, and custom data.