StateM Framework Reaches 95.3% Accuracy on Terminal-Bench via Harness Scaling Without Model Weight Changes

liuziwei7 · x · 2026-08-18

The paper introduces StateM, an agent-native runtime that improves agent execution accuracy via Harness Scaling (scaling the control flow around the model) without modifying model weights. It organizes execution around durable states, phase-local context, checked transitions, and recoverable runbooks.

Original post →

More from coding & agent

coding & agent channel →