What Didn't Work: LLM Won Nuanced Emails 2:1, Three Tasks Where Small Model Failed

mikegiannulis · x · 2026-09-18

The part that won't make the hype reel: for nuanced author emails, their LLM won about 2:1; for subjective "too salesy" judgments, Jev wasn't useful; prospect scoring with thin data showed no improvement.

The takeaway: they found jobs for the small model — and jobs to keep it away from.

Related event: Small model beats LLM on reranking and routing; real win is latency, not the 20-cent bill(8 posts)→

Original post →

More from coding & agent

coding & agent channel →