Elo-per-token Analysis Explains Why LLM Agents Scale Fast Then Slow Down
Kaiyuan Liu · hf · 2026-09-15
This paper introduces Elo-per-token analysis to study LLM agents' test-time strategies. Agents initially scale faster than independent sampling but eventually slow down; parallel short sessions outperform a single long run.
More from Research
- BVB benchmark tests agent video understanding by rebuilding 288 real videos in Blender — SonglinYang4 · 2026-09-15
- Stanford Bioengineering opens tenure-track faculty search, applications due Sept 30 — anshulkundaje · 2026-09-15
- Terence Tao's inverse Galois AI challenge completes its first stage — tak3sh8 · 2026-09-15
- Training from scratch on a single H100 hits 76% on ARC-AGI-1 in ~4 hours — GregKamradt · 2026-09-15
- First cryptanalytic extraction of neural networks without knowing their architecture — chaumian · 2026-09-15
- Poison set choice swings LLM backdoor attack success from 3% to 80%, SAILS paper shows — chaumian · 2026-09-15