Local AI roundup: 27B reasoning in 5.9GB, phone-class 35B, and dozens more
vramkickedin · reddit · 2026-10-02
A Reddit roundup of September 2026 local/open AI releases highlights: ternary quantization squeezing a 27B reasoning model into 5.9GB (Ternary-Bonsai-2), OrcaSAQ2 27B fitting high-fidelity reasoning in 12.3GB (plus an uncensored 15.7GB variant), and a phone-class Edge0-35B-A3B running under 3GiB. Agent-focused drops include Xiaomi's self-improving MiMo-V2.6-Pro-RL and Nex-N2.5-Pro/mini for coding and computer use. Other notables: Jev-Omni for joint text/image/audio/video decisions, diffusion-based UI generation in 1s (OUI-1), Google's benchmark-topping TimesFM-3.0, and a Qwen-2.5-1B-RLCD variant 7x faster at structured extraction on Apple Silicon.
More from Infra
- 180B Qwen model runs on one DGX Spark: 2.39-bit quant keeps 95.5% of BF16 scores — TheZachMueller · 2026-10-02
- Google: Starship must launch 1,600 times before space data centers work — TechCrunch AI · 2026-10-02
- Open-Source Local AI Avatar App Uses LM Studio and Chatterbox Voice Cloning — TheRedHairedHero · 2026-10-02
- How to cost AI-powered filters: roofline model puts 5k-review LLM filter floor at 6.6s on H100 — sh_reya · 2026-10-02
- Hugging Face cofounder launches a million sandboxes live on stage at Modal Runtime — graceisford · 2026-10-02
- Uno Speculative Decoding Hits 1.30x on 28k-Token Prompts, 2x DFlash, Now in vLLM — yuntiandeng · 2026-10-02