Dev questions whether OpenAI's prompt_cache_key design wastes massive compute
YouJiacheng · x · 2026-09-11
Developer YouJiacheng argues that OpenAI's promptcachekey design may be wasting significant compute, noting that the /fork and /side endpoints appear widely used and likely suffer poor prompt cache hit rates.
More from Infra
- China's CXMT plans four new DRAM fabs by H2 2028, closing in on Samsung and SK hynix — zephyr_z9 · 2026-09-11
- OpenAI CFO Sarah Friar: compute bought a year ago could sell for 3-5x today — rohanpaul_ai · 2026-09-11
- NVIDIA and Red Hat team up with vLLM on a bounty for real builds with small open models — NVIDIAAI · 2026-09-11
- SF Compute signs $245M in take-or-pay contracts for NVIDIA Blackwell B300 capacity — mattshumer_ · 2026-09-11
- SpaceX CFO: vertical integration is core, Starship paves way for orbital compute — elonmusk · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11