With an RTX 3090 Ti, should you run models in fp16, fp8, or int8?

throwaway0204055 · reddit · 2026-07-27

A Reddit user asks which precision to use on an RTX 3090 Ti with 24GB of VRAM and 64GB of system RAM: fp16, fp8, or int8.

The post is essentially a local inference / quantization setup question for running models on consumer hardware.

Original post →

More from Infra

Infra channel →