Berkeley/MIT Open Source Inference Engine: RTX 5090 Runs 284B Models at 25 tok/s
gnukeith · x · 2026-08-22
Researchers from UC Berkeley and MIT open-sourced a new inference engine with crazy performance. It runs the 284B DeepSeek-V4-Flash model at 25 tok/s on an RTX 5090 system, which is 1.46x faster than llama.cpp on the same test, while Ollama fails to serve it. On the Qwen3.6-35B benchmark, it is up to 3x faster than Ollama.
More from Infra
- Data center water use: 10GW hyperscaler consumes 120B gallons annually — SumitGup · 2026-08-22
- OpenBot: Open-source AI Coworkers with Isolated Containers and Policy Gateways — aigclink · 2026-08-22
- MiniMax-H3 INT8 Release: Why Keep FC2 in BF16 — marres · 2026-08-22
- Nuclear 'hot rock' generates immense energy vs weak passive solar needing maintenance — tawnniee · 2026-08-22
- 2-4K GPUs can serve 100T tokens daily, sparking efficiency debate — teortaxesTex · 2026-08-22
- Prediction market opens on whether Apple will announce 1TB+ unified memory chip — benfielding · 2026-08-22