New benchmark Session-Bench compares session preservation across 10 coding harnesses; Pi scores 18/19
jazzy8alex · reddit · 2026-08-15
Session-Bench is a new agentic benchmark that evaluates how well coding harnesses preserve session records after task completion. It compares 10 CLI session formats across 19 gates covering completeness, readability, stability, openness, and tooling. Key findings: the same probe produced a 1.5 KB session in Pi and 101 KB in Kimi Code; only Pi, OpenClaw, and Kimi Code stamped a true session-format version; some harnesses preserve readable reasoning, others store sealed reasoning or signatures; some record dollar cost, others only token counts. Pi scores 18/19, OpenClaw 17/18, Claude Code and Codex tie at 12/18. The author emphasizes this is not a coding-quality ranking but a report card for an overlooked part of coding-agent infrastructure.
More from coding & agent
- Sandbox escape detection framework for AI agents: six steps including monitoring, alerting, and red-teaming — blaizedsouza · 2026-08-15
- Multi-agent systems are the next abstraction, potentially creating super-intelligence — scaling01 · 2026-08-15
- Grok 4.6 now available in GitHub Copilot — intellectronica · 2026-08-15
- AQuA: Recursively Self-Improving Quantitative Trading Research Agents — MengdiWang10 · 2026-08-15
- Cursor's Returns Are — Scobleizer · 2026-08-15
- Discussion: Underrated coding agents besides Codex and Claude Code — ai__supremacist · 2026-08-15