Same insurance table query: Ministral 14B and Qwen3.8-27B nail it, Gemma 4 31B hallucinates
andrejusb · x · 2026-09-05
The author tested three models on the same insurance pivot table, same "" query, with no schema provided:
- Ministral 14B (Standard mode): answered correctly
- Gemma 4 31B (Advanced mode): hallucinated, nulled real values, and shifted columns
- Qwen3.8-27B (swapped into Advanced mode): also correct
A demo video, GitHub repo, and live app are linked. The takeaway: model size and "advanced mode" don't guarantee reliability on structured table queries — the 31B Gemma lost to two smaller models.
More from Models
- Codex users report $200 Pro quota draining in half a day, down from days — xiaohu · 2026-09-05
- OpenAI engineer claims GPT-6 Astra moved internal projects 6 months ahead — Polymarket · 2026-09-05
- After extensive testing, this dev says open-source models are now neck and neck with frontier labs — Fluffy-Ad-889 · 2026-09-05
- Sam Altman says Astra can turn his game ideas into playable demos within minutes — sama · 2026-09-05
- GPT-6 availability is a mess: Pro tier in Chat, all efforts in Work, absent in Codex — justalexoki · 2026-09-05
- Asking Astra to generate an animation with both time and space symmetries — yaroslavvb · 2026-09-05