OPUS selects training data in optimizer space and builds a 30M-token benchmark proxy
VoidAsuka · x · 2026-07-25
The authors describe OPUS, a data-selection method that scores examples in the geometry actually used by the optimizer instead of raw gradient space.
They also introduce BENCH-PROXY: rather than using benchmark validation data directly, they embed the benchmark, retrieve similar documents from the pretraining corpus, and build a 30M-token proxy pool to keep the target direction on the pretraining manifold and reduce noise.
More from Research
- Hamel Husain says evals beat vibes after Opus 5 costs 6× more and scores worse — HamelHusain · 2026-07-25
- ByteDance- and Monash-led paper turns task experience into weights for software agents — imjustnewatai · 2026-07-25
- Snorkel AI says agent benchmarks should be rebuilt from production traces — AI Engineer · 2026-07-25
- Statistical physics paper studies optimal MLP learning near interpolation — burny_tech · 2026-07-25
- ICML paper says regularized learning often looks Hebbian, while noise turns it anti-Hebbian — burny_tech · 2026-07-25
- AI Autonomously Generates 9,000 Lines of Math Proofs for Fluid Dynamics — burny_tech · 2026-07-25