OntoPrune cuts 83% of context tokens, speeding local CPU TTFT by 6.7x

vigmarcarlo · reddit · 2026-10-05

OntoPrune: context pruning that makes small coding models viable on pure CPU

Targeting the two pain points of running small local coding models (Qwen 2.5 Coder 1.5B/3B, Gemma 2B, DeepSeek Coder) — slow prompt evaluation and noise-induced hallucinations — the author open-sourced OntoPrune (MIT, fully offline Python middleware):

Benchmarks (12-core CPU + Ollama, qwen2.5-coder:3b)

Engineering notes

Install with pip install ontoprune; supports CLI piping straight into Ollama and ships a reproducible benchmark script.

Original post →

More from coding & agent

coding & agent channel →