Benchmarks: Running DeepSeek Locally on 4x 5060 Ti with 128k Context

Ambitious_Fold_2874 · reddit · 2026-08-01

A developer on Reddit shared their local inference speeds running DeepSeek models using llama.cpp, powered by 4x 5060 Ti (16GB) GPUs and 4-channel DDR4 3200 RAM.

With a context window of 128,000, the setup achieved approximately 200 tps for prompt processing and 11 tps for token generation (with -ub/-b set to 4096).

Related event: Developer Tests Local DeepSeek Deployment on Four 5060 Ti GPUs(2 posts)→

Original post →

More from Infra

Infra channel →