Gemini and Inkling Underperform on WeirdML, Suggesting Overfitting to Agentic Settings

xeophon · x · 2026-07-30

In the newly released WeirdML v2 benchmark, Gemini 3.5 Flash Lite and Inkling scored only 39.0% and 32.3% respectively, falling far below expectations.

Analysis suggests that while these models possess strong underlying capabilities, they struggle with slightly different text-only settings. This indicates they may be overtrained on agentic workflows, hindering their ability to generalize to other reasonable scenarios—a potential red flag for their overall adaptability.

Furthermore, the WeirdML v2 update introduces API cost tracking, revealing a clear scaling relationship between performance and cost, alongside a diverse Pareto frontier across various price points.

Original post →

More from Models

Models channel →