YC hosts inference-focused Paper Club: naive vs tuned inference can differ 100x in cost
ycombinator · x · 2026-10-09
Y Combinator's next Paper Club (Oct 21, 5-9pm, Mountain View) focuses on inference, arguing it has become as complex and important as training — naive vs well-tuned inference can differ by up to 100x in latency and cost. Speakers include Dan Fu (Together AI), Jeffrey Morgan (ollama), Steven Arellano (Wafer), and a TBD speaker from SGLang/RadixArk.
More from Infra
- Nebius launches spot auctions with H100s at $0.87 per GPU-hour — kevinsxu · 2026-10-09
- Oracle's $18B Project Jupiter debt trades below 90 cents as AI data center risk seeps into retail products — MatthewChang · 2026-10-09
- Air Street's Benaich: AI labs are shifting compute from pretraining to RL — nathanbenaich · 2026-10-09
- Musk: internet access is the top poverty fix; SpaceX to launch ~90% of orbital mass this year — elonmusk · 2026-10-09
- AI inference to hit ~$350B by 2027, surpassing databases as software's biggest market — demian_ai · 2026-10-09
- Dev joins NVIDIA Inception for GPU discounts to fund a 4x RTX 6000 TP8 local build — net_termina · 2026-10-09