Is LMArena quietly testing a Gemini 4 stealth model? SVG quality sparks suspicion
Background-Youth7289 · reddit · 2026-09-20
A Reddit user compared anonymous LMArena models on an SVG 'horse riding a bike' prompt. The model labeled Gemini 3.8 Flash High took 30 minutes to render vs 5 minutes for Claude Fable 5.1 High — opposite of earlier reports — and produced dramatically better output than expected from a Flash-tier model. The author suspects LMArena was quietly testing an experimental, possibly Gemini 4-generation model under that label, though it's unverified speculation. Worth watching for corroborating reports.
More from Models
- Fable 5.1 laps Astra in hands-on test: it architects and fixes, Astra stalls and burns tokens — altryne · 2026-09-20
- Jev matches Claude Haiku 4.5 at podcast ad detection — cheaper and much faster — ttlequals0 · 2026-09-20
- DiffusionGemma denoises in a single parallel pass, ~0.2s structured decisions on DGX Spark — inductionheads · 2026-09-20
- Redditor pleads with FP4 inference engine builders: small dense models at FP4 are cooked — buttplugs4life4me · 2026-09-20
- Ex-OpenAI dev tries GPT-5.6 Luna on Extra High: can't even find a link in the conversation — yuntiandeng · 2026-09-20
- Early user: gpt-6 astra follows instructions far better than 5.6-sol, no more tangents — haider1 · 2026-09-20