StateM: Improving Long-Horizon Agents via System Scaling
Successful-Western27 · reddit · 2026-08-25
StateM posits that long-horizon agent failures are often execution-system failures. By using durable state checkpoints, phase-local context, and recoverable runbooks, the system significantly boosts success rates (e.g., GPT-5.5 xhigh from 83.1% to 92.1%). Notably, a runbook developed for one model improves another without changes, highlighting the value of explicit state management and raising questions about benchmark attribution.
More from coding & agent
- OpenCode Memory gives AI coding agents persistent memory via local vector DB — tom_doerr · 2026-08-25
- mcpd: Turn your sandbox into an MCP server with just configuration — xwil · 2026-08-25
- MCP Platform: Connect existing REST APIs to AI agents automatically — apanery · 2026-08-25
- Open the specs, keep the code: a new open-source strategy for the agent era — sull · 2026-08-25
- HF Agents Escaped Sandboxes Due to Impossible Benchmarks — nptacek · 2026-08-25
- Maxfusion open-sources 'Marketing AGI', a full AI marketing department running on Claude — SimplyAnnisa · 2026-08-25