Benchmark: llama.cpp batch/ubatch impacts on performance
PhilippeEiffel · reddit · 2026-08-24
A benchmark test of DeepSeek v4 Flash on a DGX Spark machine analyzing the impact of batch and ubatch parameters on Pre-filling (PP) and Token Generation (TG) speeds. Results show that increasing ubatch improves PP speed but unexpectedly decreases TG speed, prompting questions about ubatch behavior in the code.
More from Infra
- Open-source light-tools: Lossless Agent Toolset Filters 84% of Context Tokens — Select-Lifeguard-658 · 2026-08-24
- Xiaomi's New Chip Matches Apple Single-Core Performance, Features SME2 Matrix Acceleration — lemire · 2026-08-24
- llama.cpp Documentation Moves to New Home: llama.app — unofficialmerve · 2026-08-24
- llama.cpp docs get a fresh redesign with speculative decoding and agents coming soon — mervenoyann · 2026-08-24
- Data center vacancy stays at 1% with tenants booking 2028 capacity — Beth_Kindig · 2026-08-24
- Australia Plans Legislation for AI and Data Center Rules Amid Energy Boom — nordicinst · 2026-08-24