Benchmarks: Running DeepSeek Locally on 4x 5060 Ti with 128k Context
Ambitious_Fold_2874 · reddit · 2026-08-01
A developer on Reddit shared their local inference speeds running DeepSeek models using llama.cpp, powered by 4x 5060 Ti (16GB) GPUs and 4-channel DDR4 3200 RAM.
With a context window of 128,000, the setup achieved approximately 200 tps for prompt processing and 11 tps for token generation (with -ub/-b set to 4096).
Related event: Developer Tests Local DeepSeek Deployment on Four 5060 Ti GPUs(2 posts)→
More from Infra
- NXP Semiconductors in Talks to Acquire AI Chip Designer Ambarella — pstAsiatech · 2026-08-01
- Full 2.78T-parameter Kimi K3 Runs on Consumer Laptop via NVMe Streaming — rickasaurus · 2026-08-01
- CXMT's LPDDR6 Memory Nearing Mass Production with 12,800Mbps Speed — bookwormengr · 2026-08-01
- OpenAI Hits Git Perf Limits in Giant Monorepo, Upstreams Fixes — charliermarsh · 2026-08-01
- Why Chinese LLMs Struggle in AI Coding: The Hidden Costs of Compute and Quotas — 创业邦 · 2026-08-01
- Taalas Bakes Llama 3.1 into Custom Silicon, Hitting 15,000 Tokens/sec — generativist · 2026-08-01