One-line PyTorch tweak: set_float32_matmul_precision('high') actually pays off
code_star · x · 2026-09-05
A developer reports that PyTorch's often-ignored warning to set torch.setfloat32matmulprecision('high') is genuinely worth following.
The setting lets float32 matmuls use faster lower-precision paths (like TF32 on Ampere+ NVIDIA GPUs), delivering noticeable speedups for training and inference with negligible accuracy loss in most cases. Many users dismiss the message; the author found it does improve real-world performance.
More from Infra
- Jensen Huang: 1GW of AI data center costs $50-60B, and 'we're building 100 gigawatts' by decade's end — victor_explore · 2026-09-05
- Full recipe: running Qwen3.8 27B on AMD Strix Halo with patched ROCm llama.cpp — ilintar · 2026-09-05
- Open-sourced Lightpanda: headless browser 11x faster than Chrome with 9x less RAM — JafarNajafov · 2026-09-05
- Tencent Hunyuan preview: 770B params, 1M context, Apache 2.0, 214GiB quantized — Aiden_Tech_Ai · 2026-09-05
- NVIDIA's $1B Single-Model Training Forecast Landed Years Ahead of Schedule — IgorCarron · 2026-09-05
- Starting local AI on attic hardware: 3x NUC11 plus Ryzen 3700X rig — -markusb- · 2026-09-05