Swordfish Inference Kernel Released
AlpinDale · x · 2026-07-13
Introduces the Swordfish Kernel: a weight-only inference kernel designed for the Blackwell SM100 architecture. The author positions it as the successor to Marlin and Machete, aiming to maximize throughput on data center-grade Blackwell GPUs.
Two core results are highlighted:
- On the B200, performing INT4 inline dequantization achieves roughly 90% of cuBLAS BF16 performance.
- On Jetson Thor, it reaches approximately 95% of memory bandwidth utilization.
The author notes this is the result of their personal research into the Blackwell architecture, and both a related blog post and the source code have been made public.
Related event: Swordfish Inference Kernel for Blackwell Open Sourced(2 posts)→
More from Infra
- Tesla’s FSD v14 Lite is reportedly headed to 4 million older HW3 cars — MatthewBerman · 2026-07-21
- TSMC’s 3nm utilization reportedly tops 120% as AI demand drives a $190B capex cycle — tengyanAI · 2026-07-21
- Nativ brings local AI model running to Mac with a desktop app and localhost API — Simon Willison · 2026-07-21
- Octen says agent search now runs at 62ms P50 with only a 6ms P90 gap — aakashgupta · 2026-07-21
- Zhipu acquires a compiler-team spinout to optimize AI inference on domestic chips — zephyr_z9 · 2026-07-21
- Open reproduction of Meta’s REWIRE data pipeline cuts the cost to about $11 — vanstriendaniel · 2026-07-21