New benchmark Session-Bench compares session preservation across 10 coding harnesses; Pi scores 18/19

jazzy8alex · reddit · 2026-08-15

Session-Bench is a new agentic benchmark that evaluates how well coding harnesses preserve session records after task completion. It compares 10 CLI session formats across 19 gates covering completeness, readability, stability, openness, and tooling. Key findings: the same probe produced a 1.5 KB session in Pi and 101 KB in Kimi Code; only Pi, OpenClaw, and Kimi Code stamped a true session-format version; some harnesses preserve readable reasoning, others store sealed reasoning or signatures; some record dollar cost, others only token counts. Pi scores 18/19, OpenClaw 17/18, Claude Code and Codex tie at 12/18. The author emphasizes this is not a coding-quality ranking but a report card for an overlooked part of coding-agent infrastructure.

Original post →

More from coding & agent

coding & agent channel →