Triton 后端让 Falcon3-10B 在 RTX 5070 跑到 97.5 tok/s
OCV_Researcher · reddit · 2026-07-27
Triton 后端让 Falcon3-10B 在 RTX 5070 上跑到 97.5 tok/s
一位 Reddit 用户做了一个面向 tiiuae/Falcon3-10B-Instruct-1.58bit 的 GPU-only 推理后端,并在 NVIDIA RTX 5070 上给出了明显的解码提速数据。
结果
- 混合 packed decode:97.51 tok/s
- 标准 Transformers BitLinear decode:9.89 tok/s
- 提升:9.86 倍
- fully packed prefill:426.63 tok/s,对比标准的 298.72 tok/s
技术点
- K-contiguous packed ternary weights
- packed-word DP4A decode 路径
- Triton kernels
- StaticCache
- CUDA Graph replay
校验与限制
作者强调做了 logit bit-exact 校验,也和基线生成结果做了对比,但当前结论只来自单卡、单平台测试;而且计时不包括加载、分词、重打包、JIT 编译、graph capture 和流式输出等开销。
代码和 release 都已公开,作者也在征集 Ampere / Hopper / Ada / Blackwell 上的独立复现。
「Infra」频道最新
- Windows 11 上 ComfyUI 跑通 Sage Attention 和 FlashAttention 2 — Left_of_Laniakea · 2026-07-27
- Baseten 为 GLM-5.2 做出峰值 280 tok/s API — iamrobotbear · 2026-07-27
- 马斯克等巨头力挺开源,AI 竞争的终局在基础设施? — myllmnews · 2026-07-27
- 开源 GGUF 显存计算器可先算上下文内存占用 — Dry_Wing_ · 2026-07-27
- Sparrow 标准模式改用 Ministral 3 14B 做本地抽取 — andrejusb · 2026-07-27
- Wistron 在德州投产 7 亿美元工厂,量产 Nvidia GB300 与 Vera Rubin — Beth_Kindig · 2026-07-27