Gemini 3.8 Flash benchmarked so fast the engineer thought his Slurm job crashed
_lewtun · x · 2026-09-25
Hugging Face engineer Lewis Tunstall praised Google's serving infrastructure: while benchmarking Gemini 3.8 Flash on a new eval, it ran so fast he assumed his Slurm job had crashed — an anecdotal signal of the model's notably low inference latency.
More from Infra
- Jensen Huang Pushes Back on AI Doom: "AI Is Software, It's Math" — DavidLinthicum · 2026-09-25
- Elon Musk: space compute will obviously round up to 100% of all compute — CurieuxExplorer · 2026-09-25
- Tessara's falsification test for its Micron call: watch DRAM price change on Sept 30 — tengyanAI · 2026-09-25
- Tessara predicts Micron Q3 revenue of $56.2B, beating the highest of 22 analyst estimates — tengyanAI · 2026-09-25
- Your data stack is about to get less forgiving: agents turn stale data into wrong actions — bigdata · 2026-09-25
- Bending Spoons runs 99% of AI traffic on self-hosted open models, thanks to evals — alex_verem · 2026-09-25