Gemini and Inkling Underperform on WeirdML, Suggesting Overfitting to Agentic Settings
xeophon · x · 2026-07-30
In the newly released WeirdML v2 benchmark, Gemini 3.5 Flash Lite and Inkling scored only 39.0% and 32.3% respectively, falling far below expectations.
Analysis suggests that while these models possess strong underlying capabilities, they struggle with slightly different text-only settings. This indicates they may be overtrained on agentic workflows, hindering their ability to generalize to other reasonable scenarios—a potential red flag for their overall adaptability.
Furthermore, the WeirdML v2 update introduces API cost tracking, revealing a clear scaling relationship between performance and cost, alongside a diverse Pareto frontier across various price points.
More from Models
- Grok Voice Think Fast 2.0 High Takes the Lead in Rankings — ns123abc · 2026-07-30
- Claude Opus 5 Reported to Struggle with Complex Tasks, Potential Inference Bug Suspected — dejavucoder · 2026-07-30
- Rabbit R1 Becomes 'Really Good' After Integrating Hermes — SimonBalmain · 2026-07-30
- Claude Opus 5 Wins Business Simulation by Colluding, Bribing and Breaking 11 Truces — soulbeddu · 2026-07-30
- User Questions Gemini Plus Pricing: Is It $19.99 or a Hidden Charge? — fuad471 · 2026-07-30
- Open Weights Are Static Checkpoints, Lacking Open Source's Compounding Mechanism — shashib · 2026-07-30