llama.cpp v0.6.0 ships with Metal speedups, Clef support and a new desktop app
ggerganov · x · 2026-10-06
Georgi Gerganov announced llama.cpp v0.6.0 with:
- Clef support (text + vision)
- High-quality support for Qwen3.8-Flash-Next
- Major Metal performance improvements
- A new llamabatchext API
Alongside the release comes llama.app, a free open-source 1MB menu bar app for Mac (system tray on Windows) that runs the latest open models locally. It bundles a ChatGPT-like chat UI plus a local OpenAI-compatible API server on localhost, so coding agents and editors plug in directly. One-click model installs (e.g. Qwen3.8 27B, gpt-oss 20B, gemma-4), works offline, and Linux gets a one-line install script.
Related event: llama.cpp v0.6.0 Released with MTP Speculative Decoding and Metal Gains(2 posts)→
More from Infra
- SpaceX files for 32.4-mile Florida gas pipeline to make Starship methane on-site — DimaZeniuk · 2026-10-06
- Inference Companies Ship Gateways, Gateway Companies Ship Inference — the Lines Are Blurring — michellechen · 2026-10-06
- Can an 8GB Radeon 5700 Run Image-to-Video? ComfyUI Low-VRAM Question — Fluffy-Composer9675 · 2026-10-06
- Pairing a Blackwell RTX 4000 24GB with an Old RTX 3060 12GB for Local LLMs — Otherwise-Tangelo-52 · 2026-10-06
- Local LLM for VFX: 2x DGX Spark vs M5 Ultra 256GB for Blender/Houdini via MCP — abrasmel · 2026-10-06
- Modal Runtime keynote ships Endpoint Candidates, Clusters and VM Sandboxes — graceisford · 2026-10-06