Tired of Cheesed Benchmarks: Redditors Ask Where to Find Trustworthy LLM Comparisons
Sarlo10 · reddit · 2026-09-15
A Reddit user vents that content creators hype every new model as AGI while benchmarks keep getting cheesed, and asking AI itself just surfaces unreliable sources. They ask where to find trustworthy info on how LLMs actually compare in real usage — including whether Chinese models are genuinely that good or just prompt-dependent, unlike Claude which excels even with bad prompts. The thread probes the reliability of benchmark-driven coverage versus real user experience.
More from Models
- Twin prime bound pushed to 186 as GPT-6 Astra launch fuels lab math race — RexDouglass · 2026-09-15
- OpenAI cuts desktop voice pricing ~60%, 2.4x more ChatGPT Voice in Codex — athyuttamre · 2026-09-15
- Lobehub bench shows high reasoning beats max, blogger slams vendor marketing spin — karminski3 · 2026-09-15
- Qwen 27B Passes the Pelican Test for $0.21 vs $3 on Paid Models — DivideHorror3217 · 2026-09-15
- LlamaIndex hits back: Cohere Parse scores only ~50% on ParseBench despite $1.50/1k pages pricing — ravithejads · 2026-09-15
- DeepSeek engineer who wrote v4.1's core Attention operator reflects on building his own replacement — WebAssemblyMan · 2026-09-15