Gemini 3.8 Flash called a joke: "Terminal-Bench never lies"
oleks01 · x · 2026-09-08
Developer oleks01 dismisses Gemini 3.8 Flash as "a joke," arguing users shouldn't be fooled by Google's flashy comparison charts — its real Terminal-Bench results expose the gap between marketing and actual agentic coding performance.
More from Models
- ChatGPT monthly active users top 1.06 billion in August, fourth straight record month — FinanceYF5 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- OpenAI Codex may issue another usage reset this weekend, says Codex lead resets happen — umesh_ai · 2026-09-11
- OpenAI Reportedly Pointing Its Navier–Stokes Model at Riemann and P vs NP — 141_1337 · 2026-09-11
- Benchmark author says OpenRouter unreliably honors Meta Muse effort levels, EU payments broken — PawelHuryn · 2026-09-11
- User burns $200 of Codex credits in one agent turn — 4,700 of 5,000 credits, task unfinished — RileyRalmuto · 2026-09-11