Triton backend pushes Falcon3-10B to 97.5 tok/s on an RTX 5070
OCV_Researcher · reddit · 2026-07-27
Triton backend for Falcon3-10B hits 97.5 tok/s on an RTX 5070
A Reddit user built an experimental GPU-only inference backend for tiiuae/Falcon3-10B-Instruct-1.58bit and reported substantial decode gains on an NVIDIA RTX 5070.
Reported results
- Hybrid packed decode: 97.51 tok/s
- Stock Transformers BitLinear decode: 9.89 tok/s
- Speedup: 9.86×
- Fully packed prefill: 426.63 tok/s vs 298.72 tok/s stock
What it uses
- K-contiguous packed ternary weights
- Packed-word DP4A decode path
- Triton kernels
- StaticCache
- CUDA Graph replay
Validation and caveats
The author says the implementation passed bit-exact logit checks and matched generated-token outputs against the baseline, but emphasizes that the benchmark is limited to one GPU and one Windows/PyTorch/Triton stack. Timings exclude loading, tokenization, repacking, JIT compilation, graph capture, and streaming.
The repo and release are public, and the author is specifically asking for independent reproductions on Ampere, Hopper, Ada, and Blackwell GPUs.
More from Infra
- Windows 11 ComfyUI user gets Sage Attention and FlashAttention 2 running on RTX 5090 FE — Left_of_Laniakea · 2026-07-27
- Baseten says its GLM-5.2 API hits 280 tokens/s peak and adds vision support — iamrobotbear · 2026-07-27
- Tech Giants Back Open Source, But Is AI Infrastructure the Real Bottleneck? — myllmnews · 2026-07-27
- Open-source GGUF VRAM calculator estimates context memory before you download — Dry_Wing_ · 2026-07-27
- Sparrow switches its Standard mode to Ministral 3 14B for local document extraction — andrejusb · 2026-07-27
- Nvidia supplier Wistron opens $700 million Texas plant for GB300 and Vera Rubin systems — Beth_Kindig · 2026-07-27