GPT-5.6 Sol Pro · judged blind run

Broad expert reasoning with a sealed answer boundary.

Original, subject-diverse tasks that require exact answers, concise derivations, calibrated confidence, and—where appropriate—multimodal inspection across the frontier of academic knowledge.

New closed-ended problems, not recycled exam text.

The system creates original expert tasks with known, auditable solutions while preserving the subject breadth, multimodal demands, and answer discipline that make frontier academic evaluations useful.

Each problem is authored or generated inside a controlled provenance workflow, independently solved, normalized into an exact or multiple-choice contract, and checked for ambiguity before entering an RL environment. Dense reward can combine correctness, rationale coverage, output format, and confidence calibration while strict accuracy remains the authoritative headline.

Evaluated model · GPT-5.6 Sol Pro

GPT-5.6 Sol Pro produced 50 frozen blind responses. Exact-answer verifiers and dense criterion graders then judged correctness, rationale coverage, formatting, and calibration. This is a synthetic-suite result, not an official Humanity's Last Exam leaderboard score.

One schema, many fields of expertise.

Mathematical sciences

Exact derivation

Algebra, probability, topology, numerical analysis, coding theory, and related mathematical domains.

Natural sciences

Cross-disciplinary precision

Physics, chemistry, biology, medicine, and scientific inference with unit and convention checks.

Technical systems

Structured reasoning

Computer science, engineering, algorithms, control, and distributed systems, including image-backed prompts.

Human domains

Knowledge plus interpretation

Humanities, social science, linguistics, formal semantics, policy, and other expert knowledge areas.

The judged GPT-5.6 Sol Pro run was broad, multimodal, and frozen before grading.

Strict accuracy48 / 5096.00%
Prompt modality7 + 43multimodal + text-only
Calibration0.0391Brier loss
SliceTasksAccuracyMean reward
Multimodal7100.00%0.947593
Text-only4395.35%0.915367
Multiple choice2095.00%0.904615
Short answer3096.67%0.930055

The candidate files were SHA-256 locked before the verifier and remained byte-identical afterward. The bundle retains per-item scores and concise derivations while excluding private hidden chain-of-thought.

Misses become targeted post-training material.

Two recorded missesFailure attribution
  1. DFA minimization: the submitted reachable-state interpretation produced 3 while the packaged verifier expected 6; the audit preserves the disagreement rather than silently changing the key.
  2. Normal-load reliability: the numerical calculation reached 6.889986…, but the selected multiple-choice option did not match the keyed rounded representation 6.8900.
  3. Training views: derive an ambiguity audit for the first and an answer-selection/format repair for the second.

This separation matters. One result is a potential task-quality or convention issue; the other is a clean agent error. They should not produce the same label or training response.

What an HLE-shaped program includes.

  • Original text and multimodal problems with internally authored solutions, field labels, difficulty metadata, and answer contracts.
  • Independent solve, ambiguity, duplication, leakage, image-rights, and formatting review before release.
  • Exact, numeric-tolerance, multiple-choice, or structured graders plus optional rationale and calibration dimensions.
  • Blind-run response files, concise derivations, confidence values, grader receipts, per-item labels, critiques, and repairs.
  • Private holdouts designed around underrepresented fields and model-specific failure clusters.

Benchmark relationship and source boundary.

This is an independent clean-room task family inspired by the broad expert and multimodal capability tested by Humanity's Last Exam. It is not official HLE data, an affiliation, or a leaderboard submission.

Primary reference: Humanity's Last Exam by the Center for AI Safety and Scale AI. Evidence summary: JSON.

Broad reasoning data

Map the fields where your model is still brittle.

Start a program