Running Cosmos3-Nano on RTX 5090: FP8/NVFP4 Quantization Fits in 32GB VRAM
fengwang_2_718281828 · reddit · 2026-08-08
A developer successfully quantized the Cosmos3-Nano model to fp8 and nvfp4 formats, enabling it to run within a 32GB VRAM budget. They also packaged and shared a Docker image with a WebUI, making it easy for users with RTX 5090 GPUs to deploy and test the model locally.
More from Infra
- Tesla's Magnet Expertise Could Help Musk Tackle Chip EUV Lithography — beffjezos · 2026-08-08
- Ex-OpenAI Co-founder Brockman Rumored to Tackle Silicon Supply Chain — beffjezos · 2026-08-08
- TensorLens: Inspect HF Model Quantization Layouts Directly in Your Browser — Brilliant-Hall1387 · 2026-08-08
- RAMageddon: 2027 Memory Capacity is Reportedly Sold Out — johnnyApplePRNG · 2026-08-08
- Fixing MiniMax H3 Black Frames on Legacy GPUs: FP16 Mix Cuts Inference 11x — Bubbly_Lawfulness_43 · 2026-08-08
- Microsoft Open-Sources BitNet: Running 100B LLMs on a Single CPU at 1.58 Bits — JafarNajafov · 2026-08-08