Dev ships C-based local inference engine targeting tool calling on low-VRAM hardware
ZenZombie117 · reddit · 2026-10-04
A developer released Xyntetik-Runner 1.0.0, a self-built local inference engine written in C, aimed at agentic workflows on constrained hardware rather than raw speed versus llama.cpp:
- Fixed truncated tool calling: responses return valid JSON directly, no parse-retry loops
- Schema enforcement to eliminate invalid tokens — the author calls this a killer pain point for agent workflows
- No idle memory usage: RAM is only consumed when the model is actually called
- LoRA training directly on 4-bit files, no FP16 copy needed
- Token-identical to llama.cpp, verified token-for-token on Gemma4-Moe
- Supports Metal, CUDA, and Intel/AMD CPUs, plus an OpenAI-compatible --serve endpoint
The project began on an 8GB M1 Mac and a 3070 with 8GB VRAM; access to a 24GB MIG slice of a friend's Blackwell card expanded the scope. Runner will stay free forever, with future tracks targeting enterprise-grade verified local inference and research (the earlier Xyntetik-Kvist-14B model grew out of this work).
More from coding & agent
- Replit CEO Amjad Masad: general models should JIT-train their own smaller replacements — amasad · 2026-10-04
- Startup claims it will 'kill a $1T industry'; agent dev Jason Kneen mocks the approach — jasonkneen · 2026-10-04
- Running Claude-dev + Codex-review loop, but can't wire Gemini Pro in as reviewer — bocondo · 2026-10-04
- New edition of 'An Introduction to Programming Languages' adds browser runtime and AI chapters — alfcnz · 2026-10-04
- pg-jev: an open-source Postgres extension for querying rows in plain language — Born_Excuse_5610 · 2026-10-04
- New research prototype tests whether defensive mechanisms can stop autonomous web agents — Admin-ABC-XYZ · 2026-10-04