Reddit users call Gemini 4 Argon benchmaxxed, citing weak Terminal-Bench results
PrisonOfH0pe · reddit · 2026-10-01
- Reddit user PrisonOfH0pe panned Gemini 4 Argon as benchmaxxed, saying its comparisons to Sonnet 5.5 and 6.1 "would've looked decent maybe six months ago."
- He questioned why Google only showcased benchmarks present in the training data, some of which are flawed or outdated.
- After checking Terminal-Bench results, he concluded the whole launch was "mostly paper hype."
More from Models
- VoxParity benchmark: only 11 of 23 voice agents act on what they hear, not just read — Bhavik Mangla · 2026-10-01
- Astra 6 Ultrafast Mode: 300 Tokens/Sec Changes Agent Workflows, But Tools Are Now the Bottleneck — soumitrashukla9 · 2026-10-01
- Anthropic retires Claude Opus 3 but keeps it on API and gives it an essay column — repligate · 2026-10-01
- Bindu Reddy: Gemini Argon pricing is 5x cheaper than Astra, but benchmarks look too good — bindureddy · 2026-10-01
- Gemini Answers Niche Questions Claude Can't, Says User Pushing Back on Programmer Gripes — PAstynome · 2026-10-01
- User finds Opus 5.5 Max still reproduces Opus 5's broken outputs — 0xkarasy · 2026-10-01