Turbopuffer halves queue time after fixing deceptively hard autoscaling
DanielLockyer · x · 2026-07-23
Turbopuffer says autoscaling is deceptively hard: its indexer fleet scaled nodes based on queue wait time, and when a queue went quiet it could scale down too aggressively.
That caused new jobs to wait for nodes to come back. The team adjusted the HPA signal and cut queue time by about 2×.
More from Infra
- AI video dubbing costs about $5–7 per finished minute once lip sync is included — Madmahi25 · 2026-07-23
- Nebius shows its first NVIDIA Vera Rubin NVL72 rack in Finland — demian_ai · 2026-07-23
- A Bittensor subnet launches inference at roughly half the usual price — markjeffrey · 2026-07-23
- Ascend SuperPOD optimization lifts DeepSeek-V4 post-training MFU to 34.22% — pmttyji · 2026-07-23
- For a 1 GW data center, build 2 GW into the grid — anderssandberg · 2026-07-23
- Google is spending $200B+ on cloud and compute, Beff Jezos says — beffjezos · 2026-07-23