Mystery Model Scores Over 80% on DeepSWE, Beating Fable and GPT-5.6-sol
kimmonismus · x · 2026-08-21
A mystery model achieved over 80% on 10 DeepSWE tasks, significantly outperforming Fable (65%) and GPT-5.6-sol (52%). The author speculates it might be from a Chinese company, possibly a new GLM or Kimi model.
More from Models
- DeepSeek V4-vision-exp launches API with ultra-fast speed and low cost — teortaxesTex · 2026-08-21
- Fastest NVFP4 quant of Qwen3.8 27B released, 50% faster than Q4 on compatible hardware — ionsago · 2026-08-21
- DeepSeek launches V4-Flash-Vision-Exp, multimodal agent performance nears Opus 4.8 — deepseek_ai · 2026-08-21
- Jie Tang on scaling history: FLOPs were intelligence, parameters were knowledge — cedric_chee · 2026-08-21
- Leaked Benchmark Suggests MIMO V3 Pro Performance Rivals Fable — teortaxesTex · 2026-08-21
- Leaked GLM-5.4 Reportedly Surpasses GPT Thanks to Rapid RL Iteration — kimmonismus · 2026-08-21