Fable scores 78 on vision-logic benchmark, still misses expert-level CAD errors

Afinetheorem · x · 2026-09-02

Fable model performs near the top on vision+logic tasks, scoring 78 on a private benchmark—surpassing the previous high of 61 by Sol 5.6 Pro. However, it still makes mistakes on problems that wouldn't miss an expert, such as adversarially designed CAD plans with clear errors that take three steps to identify.

Related event: Fable 5.1 Sets Record on Private Vision-Logic Benchmark(3 posts)→

Original post →

More from Models

Models channel →