6 LLMs tested on 97 hand-labeled cells: the misses were ASR errors, not the models

Sufficient_Flower860 · reddit · 2026-09-30

The author benchmarked extraction LLMs on a real production pipeline (video ASR transcription followed by LLM-based structured extraction), using a gold set of 6 videos with 97 hand-labeled cells at temp=0 with real validation logic.

Key results

Key insight: 3.6, 3.7-medium and 3.7-high all missed the exact same 4 cells on one video — when misses stay fixed across models and reasoning levels, the information likely isn't in the transcript at all. The residual errors traced back to ASR mangling model names, not LLM extraction quality.

Original post →

More from coding & agent

coding & agent channel →