One Layer Deeper launches a one-H100 benchmark for architecture-optimizer co-design
marksaroufim · x · 2026-08-04
One Layer Deeper is officially live as a competition about whether a model can learn to spend more compute on harder problems.
The organizers say many optimizer benchmarks assume the model is already trainable; this one instead asks participants to design the architecture, optimizer, loss, and training setup together.
Key details from the competition page:
- one file, one H100, one fixed model-state ceiling,
- you can use the training-time budget however you want,
- final ranking is based on a single score,
- the task centers on repeated modular squaring: x^(2^T) mod N,
- the leaderboard currently shows 16 ranked participants, all still at T=1.
The project is framed as a test of architecture-optimizer co-design for adaptive computation and harder test-time reasoning.
Related event: Tilde Launches 'One Layer Deeper' Compute Competition(3 posts)→
More from Research
- RL researchers debate whether Dyna-style systems should be called models or world models — tw_killian · 2026-08-04
- METR defines an “expenditure horizon” for AI optimization, using NanoGPT as a test case — gleech · 2026-08-04
- Lightbot 0 learns parkour-style whole-body contact skills for rescue scenarios — CyberRobooo · 2026-08-04
- Lean formalization of Mochizuki’s abc proof fails at the same step again — MarioKrenn6240 · 2026-08-04
- Oxford study finds constitutional midtraining cuts blackmail behavior without hurting benchmark scores — Oxford · 2026-08-04
- GoodfireAI pitches interoperability research as a way to understand how LLMs think — Scobleizer · 2026-08-04