Gemini 3.6 Flash shows strong agentic and long-context benchmarks
CounterReady4774 · reddit · 2026-07-21
The benchmark image compares Gemini 3.6 Flash with Gemini 3.5 Flash, Gemini 3.1 Pro, GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5 across price and multiple evals.
Key takeaways from the table:
- Price: Gemini 3.6 Flash is listed at $1.50/M input tokens and $7.50/M output tokens.
- Coding: It reaches 58.7% on SWE-Bench Pro and 49% on DeepSWE v1.1.
- Terminal / agentic work: It scores 78.0% on Terminal-bench 2.1 and 83.0% on OSWorld-Verified.
- Reasoning and multimodal: It posts 85.2% on CharXiv Reasoning, 83.2% on LVBench, and 91.8% on GDM-MRCR v2 at 128k.
- Long context: At 1M pointwise, it still holds 54.0% on GDM-MRCR v2.
The overall picture is a cheaper Flash model that looks strong on broad agentic and long-context benchmarks.
Related event: Google Launches Multiple Gemini Models for Enhanced Cost-Efficiency(67 posts)→
More from Models
- NVIDIA says Nemotron 3 Ultra scored 30/42 on the 2026 IMO problems — NVIDIAAI · 2026-07-22
- Gemma-4-26B-a4B reportedly beats Qwen3.6 and Qwen3.5 MoE fine-tunes — JLeonsarmiento · 2026-07-22
- OpenAI is reportedly briefing U.S. lawmakers on its next model family — kimmonismus · 2026-07-22
- Muse Spark 1.1 lands at 1495 on Text Arena with standout agentic-coding price performance — ycombinator · 2026-07-22
- Advanced AI Models Are Becoming Impossible to Plug and Play — emollick · 2026-07-22
- Google Gemini's AI Problem: No Leading Model for Core Workloads — bindureddy · 2026-07-22