llama.cpp team launches Llama: a 1MB app to run open models locally on Mac and Windows
mervenoyann · x · 2026-10-05
Llama (llama.app), a free, open-source app from the llama.cpp team and Hugging Face, positions itself as "OpenAI on your own computer": a menu-bar app plus a local OpenAI-compatible API.
- What it does: one-click download and local execution of the latest open models (Qwen, gpt-oss, gemma families), with the app picking models and settings that fit your machine
- Two modes: browser-based chat similar to ChatGPT, and a local API endpoint so coding agents, editors, and scripts just work
- Lightweight: 1MB download, fully offline, installable via brew/winget or from source; already at 1.5K GitHub stars
More from Infra
- Schmidhuber: compute gets 10x cheaper every 5 years, 100,000x in 25 years — SchmidhuberAI · 2026-10-06
- Musk confirms TSMC in talks to build chips at his planned Texas Terafab — Polymarket · 2026-10-06
- Running MiniMax H3 locally on a 12GB GPU: 0.6MP at 4:3 is the sweet spot — vortis23 · 2026-10-06
- AI Data Center Boom Hits Grid Limits as Planned Projects Get Pulled Back — DavidLinthicum · 2026-10-05
- Kirin 9050 reportedly uses 1.5-micron HBI die-to-wafer packaging; analyst says scaling it is the hard part — teortaxesTex · 2026-10-05
- PGlite hits 20 million downloads per week as in-browser Postgres explodes — matei_zaharia · 2026-10-05