Sentdex: 'instant' LLM hype rests on pedigree — 150ms is trivial with pure function calling
Sentdex · x · 2026-09-17
Responding to a discussion about low-latency "instant" models, Sentdex argues that much of the trust in them comes purely from the team's pedigree rather than demonstrated value: he hasn't seen a use case where they clearly make sense.
He also points out that "instant" is really 150ms latency, which he says a pure function-calling LLM could match easily — implying the claimed technical advantage is overstated.
More from Models
- YC built AI versions of its partners on GLM-5.2, cutting latency 31% vs OpenAI — ycombinator · 2026-09-17
- Astra isn't GPT-6 itself — it's just one tier alongside sol and luna — flowersslop · 2026-09-17
- Open weights are not open source: why AI's favorite label is under dispute — StanfordHAI · 2026-09-17
- New Class of AI 'Judgment Models' Like Jev Could Reshape Business Automation — The AI Daily Brief · 2026-09-17
- Parallel decoding vs. structured outputs: devs speculate on a closed-source release — ricklamers · 2026-09-17
- fx's safety classifier benchmarked: ~5-18x faster and more accurate than GPT-5.6-Luna — cramforce · 2026-09-17