Zeno Runs Qwen3.5-35B Locally on 16GB Mac via Offloading
close_Meal6005 · reddit · 2026-08-19
Startup Icosa released Zeno, a local AI tool running the 4-bit Qwen3.5-35B-A3B model entirely on Mac. To handle the limitation of 16GB unified memory, they built an offloading system instead of shrinking or pruning the model. The tool is free, watermark-free, and keeps data private.
More from Infra
- Distributed Locking & TLA+ Verification at Modal — tokenbender · 2026-08-19
- Alchemy Adds Fly.io Machines API Provider with Resource Bindings — samgoodwin89 · 2026-08-19
- New tool cargo-bsize helps analyze and reduce Rust binary size — charliermarsh · 2026-08-19
- Open MAX and Form Open Alliance to Unify NVIDIA, AMD, Trainium, Google TPU, and More — clattner_llvm · 2026-08-19
- Higgsfield Chooses Together AI for Inference on Dedicated Containers — togethercompute · 2026-08-19
- ModCon '26 wraps: Mojo is now open source and heterogeneous compute has a real software platform — clattner_llvm · 2026-08-19