Run 35B Models on 16GB Machines! QuarkStar Engine Hits Over 80 tok/s
Nicolodeva · reddit · 2026-08-04
A developer has released QuarkStar, a lightweight native inference engine inspired by DwarfStar, focusing on running LLMs on consumer-grade hardware.
- Core Features: Supports fully resident inference for models like Qwen3.6-35B-A3B on machines with only 16GB of RAM; includes bounded SSD expert streaming for memory-constrained setups.
- Cross-Platform: Implements native Vulkan backend on Linux and native Metal backend on Apple Silicon.
- Performance: On an AMD BC-250 (16GB GDDR6) worth about $150, it achieves 639 tok/s prefill and 81 tok/s generation at 2K context. On an M2 Pro (16GB), it hits 37 tok/s generation at 2K context.
The project aims to push the limits of low-end hardware, allowing users who can't afford $3,000+ AI rigs to run large models locally.
More from Infra
- MiniMax H3 Acceleration Benchmark: TE-Speed Delivers up to 1.785x Speedup — Commercial_Board9219 · 2026-08-04
- Cloudflare Launches CI SDK with AI Self-Healing Code Fixes — dinasaur_404 · 2026-08-04
- Gavin Baker Reveals SSI to Launch Model in August, Discusses AI Infra & GPU Prices — zephyr_z9 · 2026-08-04
- AI Trade Enters Stock-Picker Phase as Compute Supply Defies Narrative — tengyanAI · 2026-08-04
- ClickHouse Cloud Rebuilds Autoscaling Orchestration for Near Real-Time Reactivity — mgill25 · 2026-08-04
- 65-byte Malicious File Crashes llama.cpp: Open-source Library 'modelvet' Hardens Model Parsing — tetsuoai · 2026-08-04