Gemini 3.8 Flash flops on warden bench: worst cost-per-finding of tested models
zeeg · x · 2026-09-04
Contrary to the hype, developer grichadev found Gemini 3.8 Flash performed poorly on the warden bench, ranking worst among tested models on a cost-per-finding basis. He calls it a good litmus test for real applied use, and notes he couldn't test the cyber-focused Flash variant because it's gated.
More from Models
- Leak claims GPT-6 Astra scores 98.6% on ARC-AGI-3 and tops most benchmarks — yuwen_lu_ · 2026-09-04
- 'We're living in the singularity': researcher stunned by ARC AGI 3 score — rand_longevity · 2026-09-04
- Meta's long-context MRCR scores flagged as overfit: 1k samples can lift 60% to 90%+ — eliebakouch · 2026-09-04
- Unverified Leak: OpenAI Reportedly Rolling Out GPT-6 'Astra', Brockman Says 'Welcome to the AGI Era' — rohanpaul_ai · 2026-09-04
- OpenAI rolls out GPT-6 Astra to vetted cyber customers at 2.5x GPT-5.6 pricing — rohanpaul_ai · 2026-09-04
- GPT-6 Astra pricing revealed: $10/M input, $50/M output, to drop compaction — koltregaskes · 2026-09-04