NVIDIA's SoL-Pi Cuts Coding Agent Token Traffic by ~45% and API Cost by a Third
songhan_mit · x · 2026-09-18
- NVIDIA released the full SoL-Pi report (on Hugging Face), tackling a simple question: before scaling to thousands of parallel agents, can each agent waste fewer tokens?
- Their approach is RSI-inspired auto-research loops at the harness layer: scale rollouts across increasingly diverse environments and let selection pressure find reusable improvements that transfer beyond their development setting.
- Four mechanisms survived selection and form SoL-Pi: action execution, context compaction, observation handling, and delegated reading.
- On the 51-task EdgeBench eval across GPT-5.6 Sol and Opus 5, SoL-Pi matches native Pi's performance while cutting recorded token traffic by 44.7-49.0% and API cost by roughly a third — estimated savings of $8.75-13.50/hour vs native Codex and Claude Code harnesses.
Related event: NVIDIA Open-Sources SoL-Pi, Cutting Coding Agent Tokens by 45%(2 posts)→
More from coding & agent
- ffmpeg-skill: 39 FFmpeg tools to turn Claude Code and Cursor into local video editors — tom_doerr · 2026-09-18
- The 2026 company: a folder of .md agent files for every department — mdancho84 · 2026-09-18
- LoRA on abliterated Qwen 27B recalls private codebase rules for ~7.5h training — Similar_Job_6080 · 2026-09-18
- Open-source repo collects the best JEV use cases in one place, PRs welcome — matchaman11 · 2026-09-18
- Losing 3 days to a dead coding agent session sparks push for 'AgentGit' workflow forking — LeviYagami · 2026-09-18
- User: Astra in Codex is unusable even on a 20x subscription — OpenAI should copy Anthropic's usage-quota approach — CtrlAltDwayne · 2026-09-18