Scratch-built engine beats vendor runtime for ternary 8B models on free ARM cores

Annual_Manner_5901 · reddit · 2026-08-22

The author presents nucleo, a custom inference engine running Bonsai-8B ternary models on Oracle's free 2-core ARM instance. It outperforms the official llama.cpp fork by 14% in decode and 55% in prefill speeds.

Performance Benchmarks:

Technical Details:

Original post →

More from Infra

Infra channel →