Coding Evaluation Differences: Sol vs. Fable

ZainHasan6 · x · 2026-07-15

This post compares the capabilities of two coding/agent configurations, **Sol** and **Fable**. The core conclusion is that they "solve different problems," with a correlation of only **0.48**. ### Key Results - Both solved **88** tasks under optimal configurations - Solved only by Sol: **9** tasks - Solved only by Fable: **12** tasks - Solved by neither: **4** tasks ### Different Failure Modes - When Sol fails, there is a **20%** chance it breaks existing tests in the repo - When Fable fails, it's more commonly a "near miss," with about **65%** of cases missing the final new test ### Stability and Limits - Across 26 previous configurations, no model exceeded **80% reliability** (the ratio of passing the same task 4/4 times) - Sol's peak reliability is **84.3%**, with **61** tasks achieving a 4/4 pass rate - Fable has a higher peak at **88%**, beating Sol's **85.8%** - However, Sol is described as a "more reliable tool," while Fable is more likely to produce occasional but spectacular results

Related event: Sol vs Fable: Coding Agent Benchmarks Show Specialization, Not Dominance(9 posts)→

Original post →

More from coding & agent

coding & agent channel →