Llama.cpp DSpark PC Tree fork boosts inference by up to 29.5%
RapidRaid · reddit · 2026-08-22
A developer implemented the DSpark PC Tree algorithm in llama.cpp, providing benchmark results on RTX 5090 and CPU.
Performance (Qwen 3.0, RTX 5090):
- Plain: 94.27 tok/s
- DSpark n3: 155.64 tok/s (+65%)
- PCTree k3/n16: 159.00 tok/s (+69% vs plain)
The PCTree k3/n16 configuration beat linear DSpark in 9 of 11 categories, with a peak gain of +6.56% in summarization. CPU tests showed consistent stability.
Launch Args:
--spec-type draft-dspark --spec-draft-n-max 3 --spec-dspark-pctree --spec-dspark-pctree-k 3 --spec-dspark-pctree-n 16
More from Infra
- Green Compute Launches Biogas-Powered GPU Cluster on Bittensor with 4090/5090 Rentals — markjeffrey · 2026-08-22
- Nvidia's New Vera CPU Tested: Big Bandwidth & FP8 Boost Agent Execution — pzakin · 2026-08-22
- RAG and MCP Bridge the Gap Between Foundation Models and Enterprise Utility — ingliguori · 2026-08-22
- AI data center company Nscale seeks up to $3B in US IPO — nmasc_ · 2026-08-22
- Theta Launches AI Characters 2.0 with Offline Stateless Sessions on EdgeCloud — Scobleizer · 2026-08-22
- Converting vacant US office space into data centers: A feasibility analysis — edgarpavlovsky · 2026-08-22