cache-proxy Ships Exact and Semantic LLM Caching with x402 USDC Payments on Base
modelcontextprotocol · reddit · 2026-09-29
cache-proxy is a new LLM caching proxy offering both exact and semantic caching to cut inference costs on repeated or similar requests. Billing uses the x402 protocol with USDC on Base, pay-per-use, and a free health endpoint is available.
More from Infra
- Celesto: open-source persistent microVM computers for AI agents, boots in 500ms — aniketmaurya · 2026-09-29
- Redditor runs gpt-oss-120b across a phone, three Macs and two Windows PCs — ANR2ME · 2026-09-29
- BioNeMo team boosts Mixtral-8x7B training throughput 2.21x vs HF BF16 baseline — AllThingsApx · 2026-09-29
- DeepSeek's elastic compute team is hiring heavily, shares sandbox infra for large-scale agent training — teortaxesTex · 2026-09-29
- Developer slams third-party inference providers: Gemini up 10x, Luna 15s latency — julianharris · 2026-09-29
- Nereus: adaptive parallelism boosts 8B PPO throughput up to 7.27x over OpenRLHF — Songlin Jiang · 2026-09-29