Counterintuitive test: 31B Gemma hallucinates data extraction, loses to 14B Ministral
andrejusb · x · 2026-08-11
A practical test on document data extraction yielded a counterintuitive conclusion: when processing the same insurance pivot table without a provided schema, the model with double the parameters performed significantly worse.
- Gemma 31B (Advanced mode): Showed lower accuracy and severe hallucinations. It nulled real values and shifted the data into adjacent columns.
- Ministral 14B (Standard mode): Nailed the extraction task flawlessly.
Takeaway: Larger model size isn't a proxy for task fit. You should test various model tiers on your actual business documents before committing.
More from Models
- Perplexity Agent API Integrates Kimi K3 with Automated Data Story Workflow — AravSrinivas · 2026-08-11
- Renaming a PDF Boosts LLM Scores: A Weird Flaw in AI Evaluation — generativist · 2026-08-11
- DeepSeek Open-Sources DeepSeek-OCR 2: Visual Causal Flow for Markdown Conversion — tom_doerr · 2026-08-11
- Prediction: Anthropic Will Split Chat and Agentic Models — natanielruizg · 2026-08-11
- DeepSeek v4 flash jailbreak exposed: role-playing prompts bypass safety guardrails — aaditya_ai · 2026-08-11
- Approved for Anthropic Cyber Program, Claude Still Refuses: Dev Turns to Open Source — npinto · 2026-08-11