Visual quality control in 1 second of training, 65ms per frame, on laptop CPU only

Embarrassed_Bison527 · reddit · 2026-10-08

A developer demoed a fully local visual quality-control tool: point a webcam at parts, show 30 good and 30 bad examples, hit Learn, and it sorts new parts live in 65ms per frame on a laptop CPU — no GPU, no cloud, nothing leaves the machine.

The trick is that no big model is trained: a frozen SigLIP2 image encoder turns frames into descriptions, and only a thin classifier is trained on top. Retraining takes a second whenever the part changes, and text queries like "what colour is this?" work zero-shot. Compared to calling a vision API, cost, latency and privacy all favor this setup. The trained classifier clearly beats plain similarity matching, though it only hits 81% on a deliberately faint defect. The author is candid that accuracy numbers come from a single photo session, so they measure consistency, not robustness to different lighting or cameras. Built on the open-source indecis library.

Original post →

More from coding & agent

coding & agent channel →