Relace Is Now the Cheapest DeepSeek v4.1 Flash Provider on OpenRouter
ilyasu · x · 2026-09-25
Relace is now the cheapest DeepSeek v4.1 Flash provider on OpenRouter. EBorgnia's breakdown: agent traces are 96%+ input tokens so GPU utilization is prefill-bound; DeepSeek's encoder/decoder split lets a lossier cheap submodel handle prefill without hurting decode quality, with a redesigned attention cutting KV cache 4x. Fully exploiting the gains remains a deep engineering problem DeepSeek "left as an exercise." Ilya Sutskever reshared it.
More from Infra
- Perplexity Launches Fast Search API: 95% of Results in Under 230ms on Rust-Based Photon — perplexity_ai · 2026-09-25
- 100B Model Trained Across 5 Data Centers on Plain Internet Links at 30.8% MFU — markjeffrey · 2026-09-25
- AMD gaining 10 points of GPU share would be 'transformational', analyst argues — Beth_Kindig · 2026-09-25
- Google sees orbital AI data centers reaching cost parity with terrestrial ones by mid-2030s — McDonaghMatthew · 2026-09-25
- Goldman Sachs hikes AI power forecasts: 2030 data center capacity raised to 217GW — McDonaghMatthew · 2026-09-25
- Puro-2B: an open recipe trains a Qwen2-1.5B-beating LLM on RTX 5090s for just $4.4K — IgorCarron · 2026-09-25