Suspected Gemini 4 Pro spotted in Arena as gemini-3.8-flash; critics want harder tests
teortaxesTex · x · 2026-09-17
An anonymous model named gemini-3.8-flash appeared on Arena and aced the classic pelican-riding-a-bicycle SVG prompt, fueling speculation it's Google's Gemini 4 Pro testing anonymously. One commenter argued that a frontier-scale (1e27) model should be benchmarked on something harder than meme prompts.
Related event: Suspected Gemini 4 Pro Spotted in LMArena Under '3.8 Flash' Alias(6 posts)→
More from Models
- Jev Model Router Cuts Latency 95% vs GPT-5.6, Runs Inline in Agent Sessions — pwendell · 2026-09-17
- AndroidLife: Qwen3.8-27b runs 60 real phone tasks, fails 43% and cooks the chip to 98.2°C — East-Muffin-6472 · 2026-09-17
- New paper: do LLMs solve cognitive development tests like humans? — GolinoHudson · 2026-09-17
- Take: frontier LLMs are too slow and pricey — small models win workflows — TejasKumar_ · 2026-09-17
- Debate erupts after OpenAI flags model's defense of human culture as misalignment — RachelVT42 · 2026-09-17
- "Direct confidence readout" claim debunked: it's just entropy from the logit distribution — mgostIH · 2026-09-17