Three New Models Added and Benchmarked

testingcatalog · x · 2026-07-11

The AI/ML platform has added Grok 4.5 and Muse Spark 1.1 for testing, placing three recently launched models side-by-side for comparison in the Playground and API.

They were tested using three playable prompts:

In the results, Grok 4.5 performed the best across all three tasks, delivering cleaner outputs with more stable physics and visuals. GPT Sol handled the first two tasks well but struggled with the Crossy Road clone. Meta Muse Spark 1.1 was the cheapest option, but it was noticeably slower and less stable on the Crossy Road task. The author concludes that relying solely on single-run costs is misleading; the expenses from failures, rework, and undeliverable results can make a seemingly "cheap" model quite costly.

Related event: Aggregator Platform Adds Grok 4.5 and Muse Spark 1.1 for Testing(2 posts)→

Original post →

More from Models

Models channel →