Local Qwen3.8-27B Runs Typed Decisions in <10GB VRAM at 170ms

kyr0x0 · reddit · 2026-09-23

The author released Bonsai-Llama-Jev, a local typed-decision inference system built on Qwen3.8-27B Q264 and llama.cpp, keeping OpenAI API compatibility.

Key numbers

Technical points

Original post →

More from Infra

Infra channel →