Gemini 3.6 Flash holds steady overall but drops 14% on chart understanding
llama_index · x · 2026-07-22
A ParseBench benchmark of Gemini 3.6 Flash and Gemini 3.5 Flash Lite suggests Google’s newer Flash variants have not clearly improved document understanding.
Key findings
- Gemini 3.6 Flash is roughly on par overall, but drops 14% on chart understanding versus Gemini 3.5 Flash.
- Gemini 3.5 Flash Lite improves layout detection by 11%, but regresses on tables by about 12%.
- The author argues the Flash line appears to have been post-trained more for coding and reasoning, creating a plateau in visual/document recognition.
The benchmark image shows the same pattern: stronger reasoning-oriented models can outperform on agentic or layout tasks, while visual understanding lags in some subsets.
Related event: Gemini 3.6 Flash Benchmarks and Tests: Faster and Cheaper, But Not Smarter(9 posts)→
More from Models
- Cisco says two small open security models beat GPT-5.5 on vulnerability detection cost — The Decoder · 2026-07-23
- Laguna-S-2.1 stumbles on a 100-meter walk-or-drive sanity check — logic_prevails · 2026-07-23
- BTL-3 launches as a 27B open-weight agent model in an 8.39GB file — QuixiAI · 2026-07-23
- Claude Opus 4.8 now drives 40% of Anthropic usage on OpenRouter — maferase · 2026-07-23
- GLM misses only “Computer locked” in nearly perfect ProgramBench cmatrix run — jyangballin · 2026-07-23
- GLM-5.2 hits 8.5% on ProgramBench, ranks 3rd overall — parth007_96 · 2026-07-23