Dev wrote Marlin-style FP4/FP8 kernels for RTX 3090 — NVIDIA declined to upstream them

QuixiAI · x · 2026-09-07

Developer buythedip revealed they wrote Ampere Marlin-style FP4/FP8-to-FP16 dequant/GEMM kernels for RTX 3090s, based on QuixiAI's work — and NVIDIA wouldn't upstream the contribution.

Notable for anyone running low-precision quantized inference on older consumer GPUs: community-written kernels fill the gap on Ampere, but the lack of official maintenance means users must maintain their own forks.

Original post →

More from Infra

Infra channel →