Benchmark Score Gap Attributed to Safeguards, Not Smarts

eyishazyer · x · 2026-09-02

Comparative data reveals that Fable 5.1 and Mythos 5.1 share the same model weights. The score gap on Terminal-Bench 4.0 (55.8 vs 60.9) is entirely due to differences in safeguards, not increased intelligence.

Original post →

More from Safety

Safety channel →