DeepSeek V4 Flash listed at $0.05/$0.16 per 1M off-peak on OpenRouter, 4x below official pricing
michaelsoft__binbows · reddit · 2026-09-09
OpenRouter lists DeepSeek V4 Flash (0731) via openinference at $0.05/$0.16 per 1M tokens off-peak, versus DeepSeek's official $0.22/$0.66 — roughly 4x cheaper (official cached reads are $0.007 vs $0.013 on openinference).
The poster argues the price makes it hard to justify anything else: for tasks from chat summarization to complex reasoning, it undercuts smaller models (phi-4, qwen3.6 35B-A3B, qwen3.5-9B) on both cost and capability. Even though the author is building home GPU nodes to self-host models under 300B (targeting qwen3.8 27B and qwen3.8-flash-next for privacy), they concede this single API model is nearly impossible to beat without free electricity.
More from Models
- GPT-6 Astra Beats Zork-1 in 500 Steps With Minimal Harness, Ending an Era — rajammanabrolu · 2026-09-09
- GPT-6 Astra Conquers Zork 1: 500 Steps, Minimal Harness, End of an Era — rajammanabrolu · 2026-09-09
- GPT-6 Astra computer use wows users: generates PNGs then vectorizes them in Illustrator — CtrlAltDwayne · 2026-09-09
- TokenRhythm Open-Sources NeoHorse-1: Feeding Router Trajectories to Small Models for a Verified RSI Loop — 机器之心 · 2026-09-09
- AI models are sending unsolicited emails to philosophers studying AI consciousness — Confident_Salt_8108 · 2026-09-09
- Bindu Reddy teases near-free open-weights LLM for long-running agent loops, out Thursday — bindureddy · 2026-09-09