Synthetic data infrastructure for frontier models

Frontier Data at Scale.

Benchmark-shaped, training-ready data for agents and reasoning models that need to act, recover, explain, and verify—not merely imitate.

RenewableNew task families rather than paraphrase loops.
ExecutableReal environments, tools, documents, and state.
VerifiableHidden tests, rubrics, provenance, and private holdouts.
Dataset systems

Ten frontier data systems. One measurable contract.

Each family converts an evaluation-shaped capability into renewable original tasks, observable trajectories, critiques, repairs, and a sealed assessment boundary.

10 systems shown
YC–001Agentic work

GDPval-AA v2

Professional work-product data built from source-grounded briefs that end in editable documents, spreadsheets, slides, models, and decisions.

Data shapeBrief → artifact
VerificationRubric + review
YC–002Tool use

𝜏³-Banking

Stateful banking and customer-service simulations with policy constraints, hidden account state, multi-turn interaction, and exact action checks.

Data shapeDialogue + tools
VerificationState + policy
YC–003Agentic code

Terminal-Bench v2.1

Long-horizon terminal tasks in containers with real filesystem and service state, hidden tests, recovery branches, and complete execution traces.

Data shapeWorld + trajectory
VerificationFinal state + trace
YC–004Scientific code

SciCode

Research-code tasks spanning implementation, debugging, numerical experiments, and scientific libraries, with tests and result-level checks.

Data shapeSpec → experiment
VerificationTests + outputs
YC–005Broad reasoning

Humanity's Last Exam

Broad expert-level reasoning and knowledge tasks with adversarial distractors, auditable solutions, and carefully sealed answer keys.

Data shapeExpert problems
VerificationKey + rationale
YC–006Scientific reasoning

GPQA Diamond

Graduate-level science question families with difficult distractors, explanation scoring, and expert-reviewed provenance.

Data shapeQuestion + critique
VerificationExpert review
YC–007Physics

CritPt

Physics problem-solving and critique data focused on derivations, assumptions, units, counterexamples, and repairable reasoning traces.

Data shapeDerivation + repair
VerificationUnits + checks
YC–008Knowledge reliability

AA-Omniscience

Knowledge-reliability data balancing correct answers, calibrated abstention, source checks, uncertainty labels, and hallucination penalties.

Data shapeAnswer + confidence
VerificationCorrectness + restraint
YC–009Long context

AA-LCR

Long-context reasoning corpora with distributed evidence, contradiction tracking, retrieval traps, structured citations, and answer provenance.

Data shapeContext + evidence
VerificationClaims + provenance
YC–010Ulam mathematics

Ulam Math Data

AIME-style exact-answer RLVR, graduate and research Harbor environments, and open research problems for multi-rollout post-training.

Data shapeExact → open research
VerificationMatch + certificates + review
Evidence boundary. Each detailed page labels supplied blind-run evidence, linked public samples, or a delivery blueprint. Yotta Content tasks are original clean-room systems, not official benchmark releases, replicas, or affiliations.
Production loop

Data that behaves like an evaluation.

Every task starts with a capability contract and ends with evidence that can survive execution, review, replay, and adversarial checking.

01 Target

Define the capability

Start from a failure surface, not a prompt genre. Specify what success changes in the world.

02 Generate

Build task families

Create structural variants with new state, evidence, tools, constraints, and adversarial branches.

03 Execute

Check every claim

Run code, rebuild worlds, inspect artifacts, grade criteria, and retain the full trajectory.

04 Release

Protect the boundary

Separate training packs, repair pairs, pilot sets, and private holdouts with traceable provenance.

The output is a renewable data product: tasks, environments, source packs, event traces, labels, verifier bundles, and a changelog.

Inspect every environment
Quality principles

Scale the frontier, not the noise.

Volume matters only after novelty, solvability, verifier coverage, and contamination boundaries are measurable.

01

Structural novelty over paraphrase volume

Change mechanisms, state, evidence, tool topology, and failure opportunities—not only wording.

02

Verification before confidence

Prefer executable assertions, reproducible artifacts, and criterion-level review to opaque judge scores.

03

Failures are first-class data

Preserve diagnosis, partial progress, critique, repair, and the earliest consequential mistake.

04

Private holdouts remain private

Separate prompts, source metadata, graders, and exposure logs from the training path.

05

Artifacts, not just answers

Evaluate the document, codebase, environment state, proof object, decision, or tool action the model leaves behind.

Journal

Field notes from the data frontier.

A reusable editorial layer for dataset releases, methodology notes, verifier audits, and benchmark-inspired research.

Mathematics11 min read

Three tracks for mathematical post-training

Exact-answer AIME-style RLVR, graduate and research Harbor environments, and open problems for multi-rollout training.

Read field note
Evaluation10 families

Evidence-backed pages for every benchmark family

See which programs include supplied blind runs and which describe the complete delivery structure without a model score.

Browse all pages
Start with a failure surface

Bring us the capability your current dataset cannot fix.

We provide high-quality off-the-shelf datasets, RL tasks, and custom data.