Massive Divergence in Recommendations Across Nine Model Tiers
anon1959 · reddit · 2026-07-13
The author pitted multiple tiers of Claude, GPT, Gemini, and Grok against each other on 10 "best dev tool" questions to verify: will different tiers from the same model family give different recommendations?
Test scope includes:
- Claude: Haiku / Sonnet / Opus
- OpenAI: GPT 5.5 / 5.6
- Gemini: Flash / Pro
- Grok: Fast / Expert
Questions cover vector databases, coding assistants, LLM observability, RAG frameworks, GPU clouds, TTS APIs, gateways, agent frameworks, evaluations, and API providers. The results are quite interesting:
- Out of 10 questions, not a single #1 recommendation was unanimously agreed upon by all 9 tiers
- Splits exist even within the same family: Claude 4/10 consistent, Gemini 5/10, GPT 6/10, Grok 7/10
- Clear "self-preference" or anti-preference phenomena emerged: some tiers recommend their own products, while others recommend competitors
- Flagship tiers generally lean towards CoreWeave; lower-priced tiers prefer cheaper options like Lambda and pgvector
The author acknowledges this is just a single sample and could be affected by randomness, but it highlights an often-overlooked point: when people test "what ChatGPT thinks of my product," they might only be interacting with the flagship model, whereas free/lower-tier users face an entirely different recommendation logic.
More from Venture
- Dimension launches an $800M third fund and says it now manages $1.65B — chaitjo · 2026-07-21
- AI is cutting costs faster than it is creating new revenue — kevinkern · 2026-07-21
- Déjà View looks up earlier startups for any idea and how they ended — Sea-Assignment6371 · 2026-07-21
- Open-source AI could capture enterprise spending as closed-model pricing keeps eroding — bigdata · 2026-07-21
- TSMC’s 3nm utilization reportedly tops 120% as AI demand drives a $190B capex cycle — tengyanAI · 2026-07-21
- Chinese AI startups rush to raise capital as U.S. rivals still pull in more cash — KateClarkTweets · 2026-07-21