Early tests suggest Gemini 3.5 Pro is far less lazy than Gemini 3.1 Pro
Able-Line2683 · reddit · 2026-07-26
Early Arena-style testing suggests Gemini 3.5 Pro may be far less “lazy” than Gemini 3.1 Pro.
- The poster says the model is being tested on the Arena AI website in battle mode with multiple checkpoints.
- In their tests, Gemini 3.5 Pro produced 20+ separate code files, each hundreds of lines long.
- They claim that is more output than any other model they have tested so far, and that the laziness issue seen in prior Gemini versions may be fixed.
- This is still anecdotal and based on unofficial testing, not an official benchmark.
More from Models
- Kimi K3 is expected to go open-weight tomorrow, boosting the open-source camp — Hot_Example_4456 · 2026-07-26
- Claude usage screenshot shows Max-plan limits and $1,775.98 in credits consumed — letandrewcook · 2026-07-26
- Open-weight 4B models approach o3-level performance on Swedish medical exams — AccomplishedCat4770 · 2026-07-26
- Claude Opus 5 scores 86.3% on WeirdML v2 and still averages 7,000-plus tokens — xeophon · 2026-07-26
- ChatGPT’s math output looks like a serious paper after a two-hour prompt — airkatakana · 2026-07-26
- Baseten’s paper writes 247 fake facts into Qwen3 and still can’t make them stick — gerardsans · 2026-07-26