Berkeley/MIT Open Source Inference Engine: RTX 5090 Runs 284B Models at 25 tok/s

gnukeith · x · 2026-08-22

Researchers from UC Berkeley and MIT open-sourced a new inference engine with crazy performance. It runs the 284B DeepSeek-V4-Flash model at 25 tok/s on an RTX 5090 system, which is 1.46x faster than llama.cpp on the same test, while Ollama fails to serve it. On the Qwen3.6-35B benchmark, it is up to 3x faster than Ollama.

Related event: Berkeley and MIT Open-Source FreeToken for Running Giant Models on Consumer GPUs(3 posts)→

Original post →

More from Infra

Infra channel →