Cline benchmarks Ox Alpha: same bug fix with ~3x fewer output tokens

burny_tech · x · 2026-08-25

Cline compared Ox Alpha vs Fable on a real bug from its repo. Both fixed it correctly, but Ox Alpha used far fewer thinking tokens — Fable re-derived its conclusion repeatedly, saying "I found the root cause" 7 times before editing, while Ox stated it once and wrote the fix, with roughly 3x lower output tokens.

Cline notes most reasoning models rely on re-verification for reliability, whereas Ox seems to trust its first conclusion — a fundamentally different post-training philosophy. Ox Alpha had already been made free in Cline, with early benchmarks showing only marginal gains over Fable and GPT.

Original post →

More from coding & agent

coding & agent channel →