31 local AI model releases you missed in August 2026
vramkickedin · reddit · 2026-09-01
A Reddit roundup of August 2026's local-model releases, covering 31 projects:
- Big open models: Motif-3 (open 314B for long agentic tasks), Qwen3.8-2.4T-A95B (Qwen Max-class locally), LG's K-EXAONE-2.0-750B-A37B (262K context, 10 languages), Solar-Open2-250B shrunk to 153GB via 4-bit MoE
- Sparse/efficient MoE: Ornith-1.5-35B-A3B, NVIDIA Nemotron-3.5-Lightning-30B-A3B-NVFP4, AMD Instella-MoE-16B-A3B, Meituan LongCat-Flash-Lite-Sparse (million-token context)
- On-device/tiny: LiquidAI LFM2.5-2.6B for phones, plus 100M–150M-class micro models and TinyTitle, a 1.98MB chat-title model
- Notable directions: DeepSeek-V4-Pro-0813 (faster tool actions), Poolside's private coding model Laguna-S-2.1, Danish ethical model DFM-Mimir, Maple-Preview solving Olympiad problems at 200 tok/s, multiple Ling-3.0 quants, and a Gemma-4 finetune cutting repetitive AI phrasing
More from Infra
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- 2 engineers + AI designed a working LLM chip in 2 weeks, no human in the loop — 新智元 · 2026-09-01
- Why did increasing context size increase speed in Llama.cpp? — satnl · 2026-09-01
- AI inference demand surges again, supply brutally outpaced by token growth — Baconbrix · 2026-09-01
- Warp founder predicts cloud-based collaborative factories for all companies within a year — charlieholtz · 2026-09-01
- JPM: 1GW of AI Infrastructure Costs $40-45B, Frontier Labs Make ~$30B per GW — zephyr_z9 · 2026-09-01