XTC Sampling Is Already Built Into llama.cpp and Other Open-Source Inference Engines
ziv_ravid · x · 2026-09-26
The author notes that XTC sampling is already built into llama.cpp, ExLlamaV2, text-generation-webui and other open-source inference engines, so anyone running models locally can likely use it today.
The accompanying paper finally provides a proper formalization and evaluation of the technique that the open-source community has been using in practice.
Related event: XTC sampling lands in llama.cpp and other open-source engines(2 posts)→
More from Infra
- AMD publishes educational GEMM optimization ladder for Helios MI455X GPUs with HipKittens — salykova_ · 2026-09-26
- Terafab starts hiring: 1 TW/year chip output and orbital AI compute in its sights — seanmcdonaldxyz · 2026-09-26
- Link: Scaling LLM Inference from a Single Node to Millions — abhijithneil · 2026-09-26
- Blog: Scaling LLM Inference from a Single Node to Millions — abhijithneil · 2026-09-26
- Samsung, Oxford and PKU propose TrOPD to distill frontier-model reasoning into on-device small models — jiqizhixin · 2026-09-26
- Pay-as-you-go vs committed LLM API volume: real procurement questions from a scaling team — LeviYagami · 2026-09-26