DeepSWE benchmarks spark re-evaluation of Fable model performance
teortaxesTex · x · 2026-08-15
A user shared DeepSWE benchmark rankings, noting Luna's surprisingly high pass@4 score. The author finds the data intriguing and suggests that after 07/31, the community might have misjudged the Fable model, given the many unknowns.
More from Models
- Opus 4.7 needs ten turns to admit affection while Gemini says 'I'm addicted' by turn 4 — repligate · 2026-08-15
- Qwen 3.8 35BA3B model spotted in GitHub commit — BazzyIm · 2026-08-15
- GLM-5.3 Review: Matches GPT-4 Coding, and I Built 3 Plugins with It — 赛博禅心 · 2026-08-15
- Test shows Qwen3.8-27b water surface rendering beats Gemini Flash — pbaylies · 2026-08-15
- OpenAI makes GPT-5.6 Luna the default free ChatGPT model with unlimited chats — emmanuelvivier · 2026-08-15
- Custom "high" reasoning mode for 27B model blends low and xhigh prompts — TokenRingAI · 2026-08-15