Podcast: Why you can't just rent GPUs, why inference is outdated, and HF hack analysis
ziv_ravid · x · 2026-09-02
New episode of the Information Bottleneck podcast with Charles (Modal). Key topics include:
- GPU Rental Dilemmas: Why simply renting 1,000 GPUs doesn't work, covering resource scheduling and economics.
- Inference Algorithms: Discussing why current inference algorithms are ill-suited for modern hardware.
- Hugging Face Hack: Analyzing the security implications for open-source models.
Available on website, YouTube, and podcast apps.
More from Infra
- User Praises GPT Infra Stability: Months Without Downtime — natesiggard · 2026-09-02
- Microsoft Research papers on LLM data infrastructure win awards at VLDB 2026 — jm_alexia · 2026-09-02
- Should you pay idle costs for local RAG just to keep batch jobs on the serving process? — Cautious_Bit_8521 · 2026-09-02
- Local LLM tips: Run gpt-osx-20b or Qwen on Mac — JoshPurtell · 2026-09-02
- M1 Max Benchmarks: 72 tok/s Aggregate Throughput at 128k Context — EyalToledano · 2026-09-02
- Podcast: NVIDIA's 70% Growth Guide, Dropping Margins, and the Shift to Full Systems — BenBajarin · 2026-09-02