OneTriangle Launches KV Cache Transfer Engine for Faster, Cheaper Inference
brianryhuang · x · 2026-08-25
OneTriangle (YC S26) launched an inference engine optimized via Cross-Model KV Cache Transfer, aiming to reduce costs and latency.
Technical Mechanism:
- Small Model Prefill, Large Model Decode: Input is processed on a small model; its KV tensors are projected into the large model's attention space. The large model then decodes directly from this populated cache.
- Cost Reduction: 61% cheaper prefill costs.
- Speed Boost: 2.5× faster Time to First Token (4.3s → 1.7s in Qwen3 4B→14B tests).
Availability:
- Currently hosts DeepSeek V4 Flash at $0.15/M (Input) and $0.35/M (Output).
- Supports open-weight models including Llama, Qwen, Mistral, and Gemma.
- Users only need to change one line of config to integrate.
More from Venture
- IT giants quietly landed $7.5B in AI transformation contracts in a year — dhruv2038 · 2026-08-25
- Can frontier model giants survive the commoditization by DeepSeek and open source? — jfiance · 2026-08-25
- Pika launches API Club: aggregating gen media models at up to 88% cost savings — gizakdag · 2026-08-25
- Joseph Jacks Urges Hugging Face Not to Sell for $13B — JosephJacks_ · 2026-08-25
- NVIDIA's strategy: investing in open source to counter OpenAI and Anthropic — markjeffrey · 2026-08-25
- How to get your first 100 paid users: A free 5-day playbook — kylegawley · 2026-08-25