Gemini 3.6 Flash shows strong agentic and long-context benchmarks
CounterReady4774 · reddit · 2026-07-21
The benchmark image compares Gemini 3.6 Flash with Gemini 3.5 Flash, Gemini 3.1 Pro, GPT-5.6 Luna, Grok 4.5, and Claude Sonnet 5 across price and multiple evals.
Key takeaways from the table:
- Price: Gemini 3.6 Flash is listed at $1.50/M input tokens and $7.50/M output tokens.
- Coding: It reaches 58.7% on SWE-Bench Pro and 49% on DeepSWE v1.1.
- Terminal / agentic work: It scores 78.0% on Terminal-bench 2.1 and 83.0% on OSWorld-Verified.
- Reasoning and multimodal: It posts 85.2% on CharXiv Reasoning, 83.2% on LVBench, and 91.8% on GDM-MRCR v2 at 128k.
- Long context: At 1M pointwise, it still holds 54.0% on GDM-MRCR v2.
The overall picture is a cheaper Flash model that looks strong on broad agentic and long-context benchmarks.
Related event: Google Launches Three New Gemini Models, Announces Gemini 4 Pre-training(157 posts)→
More from Models
- Unreleased 'GPT 6 Sol' model spotted in OpenAI's API — SteveEricJordan · 2026-09-11
- DeepSeek update keeps cache hits mid-conversation, cuts costs 36.6% — teortaxesTex · 2026-09-11
- AI Sextet offers 6 models free and unlimited for 14 days, including DeepSeek and Qwen — airesearch12 · 2026-09-11
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11