Measurable lift on APEX-Agents.
Hundreds of files, threaded emails, and a one-line, ambiguous prompt — that’s the real job. We build process-level data in the same distribution as the work, trainable and evaluable.
Agent evaluation is moving from isolated Q&A toward real, long-horizon, professional-workflow tasks — and Mercor’s APEX-Agents is the leading example: no model passes half of it. We don’t build data to game this board; we build it to train the general capability real work demands — third-party-verified to transfer to unseen tasks.
Axis 0–100 · dashed line = 50%. Every model tops out below it.
Official dataset: 33 Worlds · 480 Tasks, spanning investment banking, consulting and law.
Authored by 256 senior practitioners (avg. 12.9 yrs — BCG, McKinsey, Morgan Stanley, Citi) around real workflows, across banking / consulting / law.
Frontier labs keep investing in evals like this; Artificial Analysis independently reproduced the results.
Third-party (Applied Compute): models trained on this data gained +7.7pp on GDPVal and +8.0pp on Toolathalon — transfer to unseen tasks, not board-fitting.
A World is the simulated work computer of an investment-banking manager — a file system of 150+ attachments, an inbox, and chat threads. Each Task is a one-line prompt with the judging rubric and a gold response behind it.
The manager’s work computer.
One line — with the answer key behind it.
What the agent actually did.
Sampled statistics track the official distribution: 150–200 attachments per World, difficulty validated across 90 model runs.
Even the strongest frontier models pass only ~34% on the first try — and the failures cluster on a handful of real capability gaps, not random noise.
RL turns “right after several tries” into “right on the first” — the ~20pp Pass@1→Pass@3 gap is its room.
One task ≈ 149 files and ~88 tool calls. The slower the climb, the harder the task.
Difficulty spreads evenly; green (solved) grows from Pass@1 to Pass@3.
Models can’t reliably identify the authoritative data source, reconcile conventions across files, trace the full calculation chain, or converge on a single final answer among candidates.
Public boards forbid training on them, and few vendors can synthesize new items in the same distribution. We can — with system synthesis, a three-tier expert bench, and a loop that keeps compounding.
Real tasks, gold responses and rubrics — straight from the APEX-Agents dataset (Investment Banking).
Tell us what you’re training; we’ll match it with same-distribution, process-level data and evaluation.