6 LLMs tested on 97 hand-labeled cells: the misses were ASR errors, not the models
Sufficient_Flower860 · reddit · 2026-09-30
The author benchmarked extraction LLMs on a real production pipeline (video ASR transcription followed by LLM-based structured extraction), using a gold set of 6 videos with 97 hand-labeled cells at temp=0 with real validation logic.
Key results
- Gemma 4 31B: 23/30 on the largest video at 14.6 min per video — too slow.
- DeepSeek V4 Flash (weekend discount): 71/97 = 73.2%, still slow at 690s for one video.
- Gemini Flash free tier (5 rpm / 20 rpd): constant 503s across 3.x versions; a polling script proved all versions were intermittently available — the bottleneck was availability, not capability.
- Local CLI via Google One Pro: gemini-3.6-flash hit 84/97 = 86.6% (25s/video); 3.7 and 3.8 both scored 91/97 = 93.8% with identical per-video scores and identical errors, but 3.8 was 5-6x slower. High reasoning on 3.7 scored slightly worse (90/97) and slower.
Key insight: 3.6, 3.7-medium and 3.7-high all missed the exact same 4 cells on one video — when misses stay fixed across models and reasoning levels, the information likely isn't in the transcript at all. The residual errors traced back to ASR mangling model names, not LLM extraction quality.
More from coding & agent
- EpiCon: shared multimodal memory lifts agent scores 1.7-4.9 points across 11 benchmarks — Ziyun Zeng · 2026-09-30
- EngiWorld: top model scores just 44.3 on professional engineering agent benchmark, 3.6% multi-software success — zhiman-ai · 2026-09-30
- xLLM training infra open-sourced with xattn attention backend and xBridges toolkit — HongyiWang10 · 2026-09-30
- Auto-research loop on 120 B300s finds 40% Kimi K3 inference gain for $9,176 — bookwormengr · 2026-09-30
- 8 research agents, one 30B model, 144 hours: RSIArena tests AI-driven post-training research — my_cat_can_code · 2026-09-30
- Free Open-Source AI Engineering Course: 500+ Lessons, 340 Hours, Math-First Curriculum — ghumare64 · 2026-09-30