LLM Output Cost Control Proxy

Atriou2 · reddit · 2026-07-10

After seeing API bills inflated by a few extremely long LLM outputs, the author built a proxy layer to sit in front of model providers and automatically limit abnormally long responses.

Key approaches include:

This tool uses a BYO key model, stores no prompts or responses, and only logs token counts. The author has opened a free beta, hoping users will route real traffic to test its stability.

Original post →

More from Infra

Infra channel →