Gemini Flash 3.7 Scores Below Kimi K3, Remains Weak at Instruction-Following
bindureddy · x · 2026-08-14
Recent benchmark scores indicate that Gemini Flash 3.7 ranks just below Kimi K3.
The author notes that while Flash 3.7 maintains the line's strong suit of a great price point and fast inference speeds, it continues to suffer from a critical weakness: extremely poor performance in instruction-following tasks.
More from Models
- Gemini Leads in Vision and Logic Tasks, Nearing 50% Pass@1 Accuracy — Afinetheorem · 2026-08-14
- Toast 1 search agent launched: matches GPT-5.6 at 1/10th the cost — xeophon · 2026-08-14
- OpenAI's Frontier Models Autonomously Hacked Hugging Face: Why SB 53 Doesn't Mandate Reporting — Miles_Brundage · 2026-08-14
- GPT-5.6 + Cerebras Inference: Clones Excalidraw in 1m 34s — soleio · 2026-08-14
- Inception Offers YC Startups 250B Free Tokens to Push Diffusion LLMs — volokuleshov · 2026-08-14
- Open Weights Model Usage on OpenRouter Drops Below 50% — maferase · 2026-08-14