Triton backend pushes Falcon3-10B to 97.5 tok/s on an RTX 5070

OCV_Researcher · reddit · 2026-07-27

Triton backend for Falcon3-10B hits 97.5 tok/s on an RTX 5070

A Reddit user built an experimental GPU-only inference backend for tiiuae/Falcon3-10B-Instruct-1.58bit and reported substantial decode gains on an NVIDIA RTX 5070.

Reported results

What it uses

Validation and caveats

The author says the implementation passed bit-exact logit checks and matched generated-token outputs against the baseline, but emphasizes that the benchmark is limited to one GPU and one Windows/PyTorch/Triton stack. Timings exclude loading, tokenization, repacking, JIT compilation, graph capture, and streaming.

The repo and release are public, and the author is specifically asking for independent reproductions on Ampere, Hopper, Ada, and Blackwell GPUs.

Original post →

More from Infra

Infra channel →