Qwen 27B q8 vs bf16 on a DGX Spark: is the 1% token difference worth the memory?

superSmitty9999 · reddit · 2026-09-08

A hobbyist running Qwen 3.8 27B at q8 locally on a DGX Spark reports being happy with results, but wrestles with FOMO: measurable differences between 8-bit and bf16 are small, yet he worries the differing 1% of tokens might be the hardest, most important ones. He asks whether anyone has actually noticed a difference.

The thread captures a classic local-deployment dilemma: memory and speed vs marginal quality, with perceived differences as the open question.

Original post →

More from Infra

Infra channel →