RTX 3090 Is the #2 LLM GPU, So Why Does Nobody Use Native INT8 W8A8?

TheOnlyBen2 · reddit · 2026-09-02

Per Hugging Face hardware stats, the RTX 3090 is the second most used GPU among LLM enthusiasts. Since it has native INT8 Tensor Cores, INT8 W8A8 should offer better performance—yet the community defaults to FP8 or smaller quants. The poster asks what they're missing.

Original post →

More from Infra

Infra channel →