NVIDIA paper: model accuracy drops 62.8% as context grows from 4K to 128K on long agent tasks
dair_ai · x · 2026-10-02
A new NVIDIA paper on long-running agents introduces Long-Transduction, a benchmark setup where models must keep reading, updating, and emitting state-dependent outputs over thousands of tokens, varying three factors independently to isolate what breaks long tasks.
Across seven open-weight models, accuracy drops 62.8% when context grows from 4K to 128K, 36.5% from input format changes alone, and 39.9% when per-step operations get harder. The takeaway: a model accepting 128K tokens still makes more mistakes the longer it works—explaining why agents lose their place mid-way through long tables or ledgers.
Related event: NVIDIA Paper: Model Accuracy Drops 62.8% on Long Tasks as Context Grows(2 posts)→
More from coding & agent
- Anthropic engineer: future models will get much better at code deletion and simplification — simpsoka · 2026-10-03
- The Best AI Workflows Keep Friction Exactly Where Mistakes Matter — alifcoder · 2026-10-03
- Dev builds browser 3D game from scratch with Opus 5.5, Blender and Three.js — jason_mayes · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- Dev tired of nbviewer crashing builds serverless browser Jupyter notebook renderer, MIT-licensed — cneuralnetwork · 2026-10-03
- Monitoring Can't Keep Up: What to Auto-Detect When Your AI System Scales — goyalshaliniuk · 2026-10-03