RTX 3090 Is the #2 LLM GPU, So Why Does Nobody Use Native INT8 W8A8?
TheOnlyBen2 · reddit · 2026-09-02
Per Hugging Face hardware stats, the RTX 3090 is the second most used GPU among LLM enthusiasts. Since it has native INT8 Tensor Cores, INT8 W8A8 should offer better performance—yet the community defaults to FP8 or smaller quants. The poster asks what they're missing.
More from Infra
- PyTorch 2.14 released with 2,995 commits from 487 contributors — PyTorch · 2026-09-02
- LeCun urges stronger cybersecurity for Neoclouds to prevent rogue AI takeovers — ylecun · 2026-09-02
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02
- Asus AI PC Price Jumps 50%, Speculating on Upcoming DGX Spark Hike — mountainyoo · 2026-09-02
- User Praises GPT Infra Stability: Months Without Downtime — natesiggard · 2026-09-02
- Microsoft Research papers on LLM data infrastructure win awards at VLDB 2026 — jm_alexia · 2026-09-02