AI price-performance up 300x-1,000x in 4 years, podcast summarization now 1% of cost
CedricMakes · x · 2026-09-09
Four years after his viral project summarizing every Lex Fridman podcast with AI, developer CedricMakes reports the entire stack has gotten dramatically cheaper:
- Transcription improved 300x+; LLM price-performance, especially for large context windows and agents, improved 1,000x+
- This year he scaled the original demo hundreds of times larger at under 1% of the per-episode cost from 4 years ago
- Stack includes Grok Build (code + orchestration), NVIDIA Parkeet V2 (300x real-time transcription), GPT 5.6-Luna (chapter timestamps), GPT 6-Astra (security & scaling)
His takeaway: the cost of summarizing all the world's valuable information is trending to $0.
Related event: Dev builds automated AI podcast summary site as costs plummet(3 posts)→
More from Infra
- vLLM x AgentX: Full-Stack Optimizations for Real-World Agentic Serving — jfiance · 2026-09-09
- Desert Ant Labs introduces on-device intelligence for every product — Arcuru · 2026-09-09
- Speculative Decoding With Qwen3-30B-A3B Yields 1.5x Local Speedup, Up to 5x — Arindam_1729 · 2026-09-09
- Explainer: Speculative Decoding Speeds Up LLM Inference by ~100% — blaizedsouza · 2026-09-09
- Cerebras CTO's chip architecture deep dives—WSE-3, Hot Chips 34, Cornell lectures—barely get any views — blaizedsouza · 2026-09-09
- Cosmos3 (64B) INT4 Quants Bring Local Image and Video Gen to Mac and CUDA — Formal-Swordfish-228 · 2026-09-09