Gemini 3.5 Flash Lite reportedly graded classwork worse than Gemini 2.5 Flash Lite
mazdarx2001 · reddit · 2026-07-24
A user reports that Google’s newer Gemini models did worse than their predecessors on grading classwork.
The post says the test compared Gemini-3.5-flash-lite against Gemini-2.5-flash-lite, and the newer model performed worse on the day of testing despite claims of newer data and stronger reasoning. The linked writeup is framed as a practical evaluation, not a benchmark paper, but the main point is clear: a newer LLM does not automatically mean a better one for real-world grading.
More from Models
- Unlimited-OCR returns to No. 1 on Hugging Face trending models — _akhaliq · 2026-07-24
- V4 update slips, with K3 weights possibly landing in the same release window — teortaxesTex · 2026-07-24
- MacaronV1 review says UI, agents, and 2.1M-token context are the real story — 卡尔的AI沃茨 · 2026-07-24
- Kimi K3 Max matches GPT 5.6 Sol Max on software tasks at 55% of the price — togethercompute · 2026-07-24
- Open-weight models are winning by running on hardware people can actually buy — max_paperclips · 2026-07-24
- Fable 5 appears to be tying an API into Slack for workflow automation — beechinour · 2026-07-24