Qwen3.8-27B Benchmarks on M2 Ultra 192GB

planetearth80 · reddit · 2026-08-18

A user shared detailed benchmarks for running Qwen3.8-27B (Q6KXL) quantized model via llama.cpp on a Mac Studio M2 Ultra with 192GB RAM. The post includes specific build details, model parameters, llama-bench flags, and throughput results. It also details the actual serving configuration with llama-server, including context size, KV cache settings, and speculative decoding parameters. The author seeks comparisons with other Mac Ultra users to optimize performance.

Original post →

More from Infra

Infra channel →