Gemini 3.8 Flash eval: strong at numerical methods, mogged by Qwen 3.8 27B on checkpoints

AnuranBuilds · x · 2026-09-03

PhyseraAI benchmarked Gemini-3.8-Flash on its TB bench across 5 random tasks: it excels at deriving and implementing coherent numerical methods, but is inconsistent on interacting edge-case semantics, overbuilds static-analysis solutions while missing the hardest coverage cases, and long trajectories often mean speculative scope expansion rather than better outcomes. Another developer confirms Gemini's random-quality checkpoints make it lose badly to smaller models like Qwen 3.8 27B on checkpoint/sequence-based work.

Original post →

More from Models

Models channel →