Testing a 27B Model Locally on Mac Studio
MaziyarPanahi · x · 2026-07-15
- This post demonstrates running **prism-ml/Ternary-Bonsai-27B-gguf** on a **Mac Studio**. It utilizes **Hugging Face Inference Providers**, hosted by **Together AI**. The author emphasizes that "you don't need a Mac Studio to use the same model," pointing to inference-as-a-service rather than hardware dependency. - The quoted content provides more context: the model runs locally using **llama.cpp + Metal**, is about **7.2GB** in size under the **Apache-2.0** license, and processes a patient medical record spanning roughly **3 years** and comprising **292 visit notes**. - During this test, **GLM-5.2** acted as the questioner while Bonsai provided the answers. The model answered 3 questions in about **2 seconds** while maintaining a cache of **19,398 tokens**. The author notes it ultimately uncovered an issue that had been buried for **17 months**, highlighting the practical utility of long-context and local privacy scenarios.
Related event: Mac Studio Runs 27B Model Locally to Process 3-Year Medical Records(2 posts)→
More from Infra
- UK AI datacentres face backlash over heat, noise and land use — nordicinst · 2026-07-21
- Fluidstack raises $830M at $7.5B valuation as Anthropic backs a $50B compute buildout — rohanpaul_ai · 2026-07-21
- Early Krea2 Gradio WebUI targets 6GB low-VRAM local runs — Fluid_Kaleidoscope17 · 2026-07-21
- Z.AI starts running a 1GW AI data center built entirely on domestic chips — Polymarket · 2026-07-21
- Local models feel far more capable once paired with the right harness — Soft-Barracuda8655 · 2026-07-21
- Voice-agent teams should use platforms first, then own STT events when failures get weird — FollowingSuitable941 · 2026-07-21