Qwen3.8-27B Benchmarks on M2 Ultra 192GB
planetearth80 · reddit · 2026-08-18
A user shared detailed benchmarks for running Qwen3.8-27B (Q6KXL) quantized model via llama.cpp on a Mac Studio M2 Ultra with 192GB RAM. The post includes specific build details, model parameters, llama-bench flags, and throughput results. It also details the actual serving configuration with llama-server, including context size, KV cache settings, and speculative decoding parameters. The author seeks comparisons with other Mac Ultra users to optimize performance.
More from Infra
- Qwen3.8-27B uncensored quants released, FastMTP boosts inference up to 3.02x — hauhau901 · 2026-08-18
- AI Accelerator Shipments Forecast to Reach 16.3M in 2026, Up 62% — Beth_Kindig · 2026-08-18
- Ollama benchmarks: DeepSeek V3 Flash leads, Qwen wins quality but 30x slower — ollama · 2026-08-18
- Agent Boom Pushes Frontier Model Gross Margins to Over 85% — ben_j_todd · 2026-08-18
- Bittensor co-founder: building open, permissionless AI you can mine like Bitcoin — markjeffrey · 2026-08-18
- SGLang Reserves 18.5GB for GDN State, vLLM Doesn't: 5x KV Cache Gap on Qwen3.8-27B — SomeRandomGuuuuuuy · 2026-08-18