Dev wrote Marlin-style FP4/FP8 kernels for RTX 3090 — NVIDIA declined to upstream them
QuixiAI · x · 2026-09-07
Developer buythedip revealed they wrote Ampere Marlin-style FP4/FP8-to-FP16 dequant/GEMM kernels for RTX 3090s, based on QuixiAI's work — and NVIDIA wouldn't upstream the contribution.
Notable for anyone running low-precision quantized inference on older consumer GPUs: community-written kernels fill the gap on Ampere, but the lack of official maintenance means users must maintain their own forks.
More from Infra
- FP8/FP4 Quantization Delivers Only ~1.5x and 2x Real Speedups, Far Below Theoretical Gains — scaling01 · 2026-09-07
- Naura demos key etch process for 64-layer 3D DRAM without EUV, selectivity above 500:1 — pstAsiatech · 2026-09-07
- Netherlands builds 'Dutch AI' by finetuning Qwen 3.5 27B in subsidized datacenter — teortaxesTex · 2026-09-07
- This Week's AI Must-Reads: OpenAI's Research Acceleration Report and Broadcom's $16.7B AI Chip Quarter — VibeMarketer_ · 2026-09-07
- On-device Android agent with Gemma 4 E2B hits 2.6 tok/s live vs 11 tok/s on replay — HowDevelop · 2026-09-07
- Prediction: OpenAI's agent compute spend may exceed total employee wages by early next year — haider1 · 2026-09-07