Open-source Baba Is You benchmark reruns Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash
pmigdal · reddit · 2026-07-29
The author reran an open-source Baba Is You benchmark against several newly released models: Claude Opus 5, Kimi K3, Grok 4.5, and Gemini 3.6 Flash.
They say July was unusually productive for model launches and are using the benchmark to compare the new releases. The thread specifically asks whether Claude Opus 5 is cheaper than Fable 5 and which model ends up being the most expensive in practice.
More from Models
- User reverses course and says GPT 5.6 Sol is actually a really good model — TheZachMueller · 2026-07-29
- Sam Altman teases GPT-5.6 Sol on Cerebras at 750 tokens/sec in July — daniel_mac8 · 2026-07-29
- Grok 4.5 Medium tops LaurenBench with 56.9%, ahead of Claude Sonnet 5 and GLM 5.2 — elonmusk · 2026-07-29
- A viral chart compares 12 paid AI tools with free replacements — nikola_mr64990 · 2026-07-29
- Verdent Partners with Moonshot to Optimize Agentic Coding for 2.8T-param Kimi K3 — PrajwalTomar_ · 2026-07-29
- Best Local Models Under 120B: Are Qwen Series the Only Answer? — Possible_Grocery8079 · 2026-07-29