Truespar Launches Paddock: High-Performance Local LLM Engine for NVIDIA GPUs
wandedob · x · 2026-08-06
Truespar has launched Paddock, a local LLM inference engine (research preview) optimized specifically for NVIDIA GPUs, supporting Windows and Linux.
Key Features & Performance:
- High Concurrency: Features continuous batching, chunked prefill, and paged KV cache, ideal for handling parallel requests from AI agents.
- Single Binary: A standalone executable with no Python/PyTorch or CUDA version sync required. Fully private by default.
- Benchmarks: On identical hardware and weights, Paddock is faster than vLLM in 11/13 scenarios (up to +76%), faster than SGLang in 13/13 (1.66x to 2.91x), and faster than llama.cpp in 13/13 (1.12x to 7.2x).
More from Infra
- Musk: Chip Capacity is the AI Bottleneck; TerraFab to the Rescue — XFreeze · 2026-08-07
- YC S26 Startup Understudy Cuts Anthropic Bills by 80% via Model Distillation — ycombinator · 2026-08-07
- AI Agent Energy Use is 600x Higher Than Standard Prompt Estimates — tobyordoxford · 2026-08-06
- Sam Altman-Backed Oklo Achieves Reactor Criticality in Under a Year — TinfoilTricorn · 2026-08-06
- Help Wanted: Porting NVIDIA's MiniMax H3 Sol-Engine Optimizations to a Single RTX 5090 — cat_trick · 2026-08-06
- Developer Builds <500 KB Physics-Informed Neural Operator for Orbital Conjunction Screening — AssociatePatient2860 · 2026-08-06