What Didn't Work: LLM Won Nuanced Emails 2:1, Three Tasks Where Small Model Failed
mikegiannulis · x · 2026-09-18
The part that won't make the hype reel: for nuanced author emails, their LLM won about 2:1; for subjective "too salesy" judgments, Jev wasn't useful; prospect scoring with thin data showed no improvement.
The takeaway: they found jobs for the small model — and jobs to keep it away from.
More from coding & agent
- 4 deployment strategies explained via 4 visuals: feature toggle, blue-green, canary — _jaydeepkarale · 2026-09-18
- ServerKit: open-source server panel for apps, DBs and Docker hits 1.3k stars — tom_doerr · 2026-09-18
- Gary Marcus: Agents hold production credentials in a security gap nobody owns — GaryMarcus · 2026-09-18
- Amp's design philosophy: simple primitives, no tricks, let models improve it — HankYeomans · 2026-09-18
- You.com wraps AI Agentic Hackathon with self-improving agents challenge — PolarBearby · 2026-09-18
- Muse for Mac Launches With Cross-App Access; User Wires Up Perplexity Deep Research API — ChrisUniverse · 2026-09-18