Benzi: compiler-backed coding agent claims 78.2% SWE-bench Verified at under 10¢ per fix

DonkeyTheKing · reddit · 2026-10-02

A developer released Benzi, an AI coding agent built around a compiler and static analysis instead of reading source code.

The idea: Current agents (Claude Code, OpenCode) fetch code snippets or build embeddings to approximate program structure, inflating tokens, causing context drift, forgetting after compaction, and breaking on multi-file refactors when line numbers shift. Benzi instead gives the model deterministic intelligence via tool calls: before any edit, the compiler reports the blast radius and runs full static checks; the model rarely needs to ask 'what feeds this function?'

Numbers: Benzi Sonnet reads only 9,125 lines to complete the same tasks vs Claude Code Sonnet at 20,704, DeepSeek's harness at 43,598, and OpenCode at 65K+ (disqualified for repeated failures).

Features: 'truth tiers' separating statically analyzable from unanalyzable facts, bridged by a runtime tracer; syntax- and semantically-verified writes; context-aware model-written repros; mid-task model upgrades. Supports Python, JS/TS, Java, C#, C++, C, Go, Rust, Ruby.

Benchmarks: 78.2% SWE-bench Verified at under 10¢ per fix using V4flash. The bet: indexing a repo beats reading raw code, and that's more economical than the industry's context-window arms race.

Original post →

More from coding & agent

coding & agent channel →