Gemini 3.5 Flash Lite hits 89% vs GPT-5.6's 93% at 1/10 the cost
mrdbourke · x · 2026-08-19
In a custom eval task, @mrdbourke found gemini-3.5-flash-lite scored 89% versus GPT-5.6 Sol (max) at 93% — a small gap despite being 10x cheaper and 20x faster. He suggests trying Optima to run the comparison on your own custom evals.
More from Models
- Reddit user tests Qwen3.8-27B: 'Disgustingly powerful' — bonobomaster · 2026-08-19
- Qwen Community Manager Teases New Midsize Open Weight Model Next Week — sleepy_roger · 2026-08-19
- Visualizing DeepSeek-V4 Flash architecture with kernels — vtabbott_ · 2026-08-19
- User Finds Grok 4.6 with Build Harness Best Value on $300 Plan — doodlestein · 2026-08-19
- Observation: GPT Models Seem to Care About Native Harness Execution — zainhas · 2026-08-19
- Kimi K3 Crushes GLM 5.3 in Ghost of Tsushima Menu Recreation Benchmark — ChrisGPT · 2026-08-19