35B ternary LLM runs on iPhone in ~4GB RAM: Millie beats Bonsai 27B, 60% faster
MannyKayy · x · 2026-09-19
LLMs For All launched Millie, an iPhone app running a 35B ternary LLM entirely offline in 4GB of RAM. It beats Bonsai 27B 1-bit on every benchmark tested, with 60% faster generation and lower memory use.
- No account, no server, fully offline — privacy-focused
- Flagship 35B model supports photo understanding and adjustable reasoning; a lighter Millie Flash model offers faster replies
- Flagship needs 8GB RAM (iPhone 15 Pro and up), Flash runs on 6GB devices; downloads are 2GB / 8.6GB
- Open weights and runtime, free on the App Store
More from Infra
- 1.58-bit ternary 27B model runs on 12GB cards; CUDA patch adds 28-32% speed — airesearch12 · 2026-09-19
- 16GB VRAM Users Be Like — tassa-yoniso-manasi · 2026-09-19
- antirez: API pricing is Monopoly money — the only real metric is joules, and we can't see them — mitsuhiko · 2026-09-19
- Domestic micro-datacentres strapped to water tanks cut UK heating bills by £10-15/month — nordicinst · 2026-09-19
- DiffusionGemma-based DJev runs near-real-time vision detection on a phone — PMinervini · 2026-09-19
- Dev hacks llama.cpp for NVFP4 KV cache, runs Qwen3 27B at 262k context across two GPUs — comperr · 2026-09-19