RunNburn: Open-Source Engine Runs 295B MoE on 64GB RAM Desktop, Beats llama.cpp

coderredlab · hn · 2026-07-30

RunNburn is an open-source Rust-based GGUF inference engine designed to run oversized models on consumer hardware that exceed VRAM/RAM capacities.

Core Mechanics:

Benchmark Results:

Scope:: For models fitting comfortably in VRAM, llama.cpp remains faster due to years of kernel tuning. RunNburn targets models that "don't fit" and currently does not support continuous batching or multi-tenant throughput.

Original post →

More from Infra

Infra channel →