Relace cuts Deepseek v3 inference costs by nearly 50% on OpenRouter
stuffyokodraws · x · 2026-08-25
Relace AI claims to have reduced serving costs for Deepseek v3 Flash on OpenRouter by nearly 2x compared to competitors. The optimizations are built on high-speed inference capabilities, reaching 10k tok/s Fast Apply and 50k tok/s Compact models. After two weeks of effort, Relace now tops the OpenRouter charts for Deepseek and is soliciting partnerships with coding agent developers offering free tiers.
Related event: Relace Cuts DeepSeek v4 Flash Pricing to a Fraction of Official Rates(2 posts)→
More from Venture
- 10 revenue-driving Grok Bot workflows: outreach, ads, clipping, running 24/7 — EXM7777 · 2026-08-25
- WonderTx targets 'left behind' chemistry molecules with AI drug discovery — ycombinator · 2026-08-25
- PromptQL founder shares 6 lessons from closing a $10M+ AI deal at a Fortune 500 — prasanna_says · 2026-08-25
- Hugging Face Annualized Revenue Jumps 50% to Over $150M in Two Months — rohanpaul_ai · 2026-08-25
- Nvidia-Backed AI Cloud Lambda in Talks to Raise $3B — SumitGup · 2026-08-25
- Dev case: Building a Grok bot with Cursor earned $999 profit in 24 hours — jeff_weinstein · 2026-08-25