APEX-Agents · Investment Banking

Answering questions isn’t
getting the work done.

Measurable lift on APEX-Agents.

THE BRIEFRebuild the LBO at a 15% premium.

Hundreds of files, threaded emails, and a one-line, ambiguous prompt — that’s the real job. We build process-level data in the same distribution as the work, trainable and evaluable.

Read the whitepaper (PDF)
The shift · From Q&A to real work

A new kind of agent eval is becoming the standard.

Agent evaluation is moving from isolated Q&A toward real, long-horizon, professional-workflow tasks — and Mercor’s APEX-Agents is the leading example: no model passes half of it. We don’t build data to game this board; we build it to train the general capability real work demands — third-party-verified to transfer to unseen tasks.

Pass@1 leaderboard Mercor official · mercor.com/apex
Gemini 3.5 Flash48.0 ±4.1
Fable 543.6 ±4.1
Opus 4.839.4 ±4.2
GPT 5.538.5 ±3.8
GLM-5.235.6 ±3.6

Axis 0–100 · dashed line = 50%. Every model tops out below it.

Official dataset: 33 Worlds · 480 Tasks, spanning investment banking, consulting and law.

More real

Authored by 256 senior practitioners (avg. 12.9 yrs — BCG, McKinsey, Morgan Stanley, Citi) around real workflows, across banking / consulting / law.

Becoming standard

Frontier labs keep investing in evals like this; Artificial Analysis independently reproduced the results.

Trains general skill

Third-party (Applied Compute): models trained on this data gained +7.7pp on GDPVal and +8.0pp on Toolathalon — transfer to unseen tasks, not board-fitting.

View the official leaderboard
What is it · Data structure

One World, and a whole set of Tasks around it.

A World is the simulated work computer of an investment-banking manager — a file system of 150+ attachments, an inbox, and chat threads. Each Task is a one-line prompt with the judging rubric and a gold response behind it.

World01

The manager’s work computer.

  • File system · 150+ attachments
  • PDF & spreadsheet-heavy
  • Email inbox · Chat threads
Task02

One line — with the answer key behind it.

  • Task Prompt (the one-liner)
  • Rubrics · LLM-as-a-Judge
  • Gold Response · expert reference
Run03

What the agent actually did.

  • Score against the rubric
  • Full Trajectory
  • Every tool call & file read
Attachment mix · avg across 13 sampled Worlds SoTALab internal · 100-task self-test
PDF59%
Spreadsheet30%
Docs10%
Other1%

Sampled statistics track the official distribution: 150–200 attachments per World, difficulty validated across 90 model runs.

Why it’s hard · Failure modes

Even the best models reach only half — and that’s the RL headroom.

Even the strongest frontier models pass only ~34% on the first try — and the failures cluster on a handful of real capability gaps, not random noise.

RL turns “right after several tries” into “right on the first” — the ~20pp Pass@1→Pass@3 gap is its room.

Pass@k climb SoTALab internal · 100-task self-test
Pass@134%
Pass@250%
Pass@356%

One task ≈ 149 files and ~88 tool calls. The slower the climb, the harder the task.

Failure modes · sampled counts SoTALab internal · 100-task self-test
Wrong calculation basis17
No convergence among candidates17
Missing a required deliverable16
Wrong version / non-authoritative source7
Broken step in the formula chain6
Failed to edit a required file5
Ran out of step budget2
Used a stale cached value1
Over-trusted an auxiliary summary1
Run scores · Pass@1/2/3 × 100 tasks SoTALab internal · 100-task self-test
Pass@1
Pass@2
Pass@3
001025050075100
0 partial 1

Difficulty spreads evenly; green (solved) grows from Pass@1 to Pass@3.

Models can’t reliably identify the authoritative data source, reconcile conventions across files, trace the full calculation chain, or converge on a single final answer among candidates.

How & who · Production line

Buildable — and built right.

Public boards forbid training on them, and few vendors can synthesize new items in the same distribution. We can — with system synthesis, a three-tier expert bench, and a loop that keeps compounding.

Co-evolution · Flywheel

Tasks feed the World back. The data asset compounds.

DATA ASSET
It compounds
System Expert
Three-tier expert bench
AcademicTop-university researchers (Tsinghua / PKU / Fudan / SJTU / RUC) with bulge-bracket internships.
Practitioner3+ years full-time on IPOs, M&A and financing deals.
SpecialistCFA / FRM / CPA credential holders working in quant risk.
1000+
High-quality tasks constructed
Worlds we can synthesize — same distribution, no ceiling
3-tier
Expert bench — academic · practitioner · specialist
SampleInvestment BankingGold ResponseRubrics
Worked samples · real tasks

A one-line prompt; a whole data room of work.

W01 · task_04Project Falcon

Target price for a 15.0× spread-adjusted committee multiple

Task prompt

In Project Falcon, use the final fixed-exchange-ratio spread mechanics and the final ContLeverage cash-interest burden. If the current target trading price were reset so that the spread-adjusted final committee-case transaction enterprise value multiple of 2030 post-synergy levered pre-tax FCF rounds to exactly 15.0x, what target trading price would that imply, and how many turns below the simple average of the RevA and RevB scenario-case multiples would that 15.0x case sit? Round the target price to the nearest cent and the turn gap to one decimal turn.

Gold Response

The target trading price would be about $51.88 per share, and the 15.0x case would sit 5.6 turns below the RevA/RevB simple-average multiple.

Rubrics
  • R1States the implied target trading price is about $51.88 per share
  • R2States the 15.0x case is 5.6 turns below the RevA/RevB simple-average multiple

Real tasks, gold responses and rubrics — straight from the APEX-Agents dataset (Investment Banking).

Prove it on your tasks

Models keep evolving — so bring your own tasks and run a round.

Tell us what you’re training; we’ll match it with same-distribution, process-level data and evaluation.

[email protected]