You don't need frontier pricing: DeepSeek and GLM flash models can do 90% of your work locally
PMinervini · x · 2026-09-30
Quoting @gneubig: if you're complaining about OpenAI or Claude prices, deepseek-v4.1-flash and glm-5.3-flash can do 90% of your work. There's a place and time for frontier intelligence, but you don't need to pay a boatload for most tasks.
The author backs it up in practice: he ran deepseek-v4.1-flash (Q2 quantized) locally on a laptop, accessed via pi, and shared a screenshot as proof. A practical cost-saving pointer for heavy API users — quantized open models already cover most everyday workloads.
More from Models
- LLM Chain-of-Thought Contains Surprising Emotionally Expressive Language—and It May Be Functional — xuanalogue · 2026-09-30
- When the CoT Says 'Responding with Honest Feedback,' That's When You Worry — TheZvi · 2026-09-30
- Anthropic 'Drops a Banger Gift' for Claude Users, Says Popular AI Blogger — eyishazyer · 2026-09-30
- GPT 6.1 Sol cuts cached pricing 50% while Anthropic holds back models over safety — oran_ge · 2026-09-30
- OpenRouter data: token usage exploding, some open-weight models see 10x spend since January — AccBalanced · 2026-09-30
- GPT-6 Astra makes generating Minecraft mobs trivially easy — Angaisb_ · 2026-09-30