Not more formulas to memorize — data that trains models to actually read the context in front of them: trainable, verifiable ICL for long-context reasoning.
Context isn't just stuffing a long document into the window. It's about how a model understands, remembers, and uses information — and that is what separates real-world performance.
Longer context isn't better. Key information buried in the middle gets missed — Lost in the Middle; what matters is precisely locating and understanding it.
Long tasks need task-state management, not an ever-growing context; the model must keep track of the current goal, actions done, and open items.
Some skills should be internalized into a stable, reflex-like ability.
Physics is a natural fit for training this: answers are definite, but must obey the definitions, boundaries, conventions, sectors, and notation of the material at hand.
Context Faithful · Object Selection · Verifiable
The coverage caps any single theory's share and fills in quantum, condensed-matter, numerical, experimental, and mathematical physics — so the data can probe the many failure modes of real research reading.
Coverage isn't a pile of topics — each maps to a real ICL failure mode: definition drift, wrong physics object, missing a sector, taking the wrong branch.
100 physics long-context ICL samples in total · each with paper-style context, a reference solution, scoring points, and multi-model run records.
This data targets the shared gap today's benchmarks expose — models improve fast on short-form science QA, but stay unreliable on paper-length context, physics-object selection, cross-paragraph evidence assembly, and stable reasoning.
Same context, same prompt, same expert rubric. Each model runs three times independently, averaged on a 0–10 scale, with experts reviewing where the key points are lost.
Six samples out of 10; averages cluster at 1.3–4.5 — even the strongest model is far from passing. Scores show the stable performance band; the full record includes every run's raw output, per-item scoring, and expert review.
What trips models isn't the length — it's holding definitions consistent across the ODE, residual coefficient, affine cutoff, and exponent, each defined by the current context.
Models recognize the Hadamard structure, yet often drop the 1/4 factor or flip a sector sign — a fidelity failure, not a misunderstanding of the transform.
Beyond whether the model got it right, it captures how the model reads definitions, selects physics objects, and assembles intermediates from the given material — ready to plug into a training / eval pipeline.
Truly valuable scientific data should be trainable, verifiable, reviewable, and traceable. Tell us your training / eval direction and we'll quickly build a matching physics ICL dataset and eval pipeline.
Tell us your training / eval direction — we'll build a matching physics ICL dataset and evaluation pipeline.