Open-source sfFFT delivers 3.8-6.1x speedup for FFT convolutions on DGX Spark GB10
IgorCarron · x · 2026-10-09
solidSF released sfFFT, an MIT-licensed fused tensor-core FFT convolution kernel tuned for NVIDIA GB10 (DGX Spark), fully open source on GitHub.
- Performance: 3.8x-6.1x end-to-end speedup over cuFFT for sequence lengths 128-8192 (memory-max mode)
- Accuracy: 4e-4 error with fp32 I/O, 2.4e-3 with bf16 (vs cuFFT's 3e-7)
- How it works: targets the core operator of long-convolution sequence models (Hyena, H3, convolution-mode S4/SSM layers, long audio filters); fuses the forward FFT, pointwise filter-spectrum multiply, and inverse FFT into one kernel launch per batch, running FFT stages as small dense DFT matrix multiplies on tensor cores within shared memory
More from Infra
- New EXLR8 quant claims to beat EXL3 on every metric, GLM 5.3 cold-start TTFT at 30s — EAccelerate_42 · 2026-10-09
- Pat Gelsinger: Memory, packaging and power—not chip speed—are AI's real bottlenecks — a16z Podcast · 2026-10-09
- Founder boosts CPU utilization with adaptive capacity planning instead of chasing users — DanielLockyer · 2026-10-09
- GPUs already within 2x of brain efficiency, and still beat human workers on energy — MikePFrank · 2026-10-09
- Do AI agents still need Kubernetes? Berlin event says yes, with agent-on-K8s cases — Al_Grigor · 2026-10-09
- Cloud Backlogs Hit $1.69T, CoreWeave Posts $2.58B Quarter as Inference Becomes the Battleground — FinanceYF5 · 2026-10-09