Dev evals 20+ third-party models vs Gemini for spec quality: all fail, "wasteland of mediocrity"
julianharris · x · 2026-09-30
Julian Harris ran a large-scale eval hunting for a materially cheaper alternative to Gemini 3.6 Flash on spec quality, testing essentially everything DeepInfra offers. All models failed. He bet Claude on the outcome beforehand and won — most third-party models were simply too old to compete with today's SOTA. He expected free Chinese open models to give third-party infra an edge, but they didn't; despite a 10x price hike over the past year (doubling again in January), Gemini "absolutely cleans up" for his use case. He eval'd 20+ models in one day and calls the third-party landscape a "wasteland of mediocrity."
More from Models
- Opus 5.5 vs Sol 6.1: Drawing Mona Lisa as Pure SVG From Memory, Both Deeply Derpy — OriginalScrubLord · 2026-09-30
- MiMO ships an agentic critic; early speedups but long-run bias concerns — teortaxesTex · 2026-09-30
- Intern-Decision open models (0.8B–4B) beat Jev on multimodal decision benchmarks — max_paperclips · 2026-09-30
- No reason for OpenAI to keep both $100 and $200 plans, argues Redditor — Over_Sheepherder4503 · 2026-09-30
- Heavy user proposes pairing OAI Pro with a 512GB Mac Studio running GLM-5.3 locally — paperchow · 2026-09-30
- Altman 'Will Think About' More Open Models; Community Points to NVIDIA Nemotron Instead — omarsar0 · 2026-09-30