Gemini 3.8 Flash hits 305 tokens/sec, nearly 2x the runner-up
NewVeterinarian5384 · reddit · 2026-09-03
Per Google's official blog, Gemini 3.8 Flash outputs 305 tokens per second — nearly double second-place Muse Spark 1.2 (154 t/s) and well ahead of GPT-5.6 Luna (126 t/s) — with an intelligence score of 59, just below the 60-66 leader group. If the speed holds in real API use, long coding-agent runs could feel far less painful. A 3.8 Flash Cyber variant shipped alongside.
More from Models
- Leak: Astra won't be the best model of the year; a 'monster' is slated for end of year — ChrisGPT · 2026-09-03
- Startup Mostik bridges AI models via their weights, tops ARC-AGI 3 at 1/20 the cost — nordicinst · 2026-09-03
- Anthropic launches browser tool to detect Claude-made files via C2PA content credentials — btibor91 · 2026-09-03
- ByteDance's looped language models match 12B rivals at 1.4B size, with Bengio as co-author — peterjliu · 2026-09-03
- Anthropic Weakened Safety Filters, Signed an AI Cyberattack Warning Letter, Then Shipped Mythos 5.1 Anyway — AgentBlackVeil · 2026-09-03
- Marin 535B A23B Frontier-Scale Training Run Is Fully Livestreamed — Sentdex · 2026-09-03