Three New Models Added and Benchmarked
testingcatalog · x · 2026-07-11
The AI/ML platform has added Grok 4.5 and Muse Spark 1.1 for testing, placing three recently launched models side-by-side for comparison in the Playground and API.
They were tested using three playable prompts:
- Fruit Ninja-style slicing game
- Angry Birds-style tower toppling
- Crossy Road clone
In the results, Grok 4.5 performed the best across all three tasks, delivering cleaner outputs with more stable physics and visuals. GPT Sol handled the first two tasks well but struggled with the Crossy Road clone. Meta Muse Spark 1.1 was the cheapest option, but it was noticeably slower and less stable on the Crossy Road task. The author concludes that relying solely on single-run costs is misleading; the expenses from failures, rework, and undeliverable results can make a seemingly "cheap" model quite costly.
Related event: Aggregator Platform Adds Grok 4.5 and Muse Spark 1.1 for Testing(2 posts)→
More from Models
- Google publishes a migration guide for Gemini 3.6 Flash and 3.5 Flash-lite — _philschmid · 2026-07-21
- Google is said to have started training Gemini 4 — thesaraharminta · 2026-07-21
- Hands-on: Gemini 3.6 Flash is Slightly More Token-Efficient Than 3.5 — scaling01 · 2026-07-21
- Google Launches Gemini 3.6 Flash and Others, Targeting Agents and Security — koraykv · 2026-07-21
- Artificial Analysis chart compares Gemini 3.5 Flash-Lite with 3.6 Flash — Expensive_Syrup_6529 · 2026-07-21
- Compute Allocation Limits: The Root Cause of Missing Architecture Innovation in European LLMs — IgorCarron · 2026-07-21