NVIDIA paper: model accuracy drops 62.8% on 128K-token tasks vs 4K
rohanpaul_ai · x · 2026-10-02
A new NVIDIA paper tested 7 open models on simple repetitive tasks (adding numbers, sorting lists) and found reliability degrades sharply with task length even within the context window.
Key findings:
- Average accuracy was 62.8% lower on 128K-token jobs than 4K-token ones
- Even the best model got every item right in only 17.1% of the longest jobs
- Models seem to understand the task but lose their place, especially when items lack IDs
Practical advice for agent builders: number every item, split big jobs into small batches, and verify every line of output.
More from coding & agent
- Arena.ai Launches HarnessTax: Quantifying How Much the Harness Matters for Coding Agents — solyarisoftware · 2026-10-02
- Berkeley paper: LLMs know your preference changed but still use the old one — rohanpaul_ai · 2026-10-02
- 1,565-email benchmark: cheapest Perplexity/Cloudflare decision model is also most accurate — michellechen · 2026-10-02
- Codex + FreeCAD MCP recreates JWST sunshield deployment step by step — burhop · 2026-10-02
- Frozen model, evolving harness: ModularRSI lifts Terminal-Bench 2.0 from 47.57 to 52.43 — jiqizhixin · 2026-10-02
- Hands-on with OpenAI Dots: proactive AI butler shows promise, but latency and UX fall short — cedric_chee · 2026-10-02