With an RTX 3090 Ti, should you run models in fp16, fp8, or int8?
throwaway0204055 · reddit · 2026-07-27
A Reddit user asks which precision to use on an RTX 3090 Ti with 24GB of VRAM and 64GB of system RAM: fp16, fp8, or int8.
The post is essentially a local inference / quantization setup question for running models on consumer hardware.
More from Infra
- OpenAI may be hitting compute limits as Codex and ChatGPT Work jump from 2M to 10M users — JoshuaJBouw · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- NVIDIA says Vera CPU lifted selected EDA workloads by up to 1.5x — NVIDIA Blog · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- MiniBot 2.40 adds xAI, HF Studio and vLLM support with inline media tools — Creative-Type9411 · 2026-07-27
- Apple smart glasses, Nvidia-SK AI data center deal, and Ctrip’s RMB 5.179 billion fine headline a tech roundup — APPSO · 2026-07-27