clef-flash distilled into 0.6B model on Apple Neural Engine triages 1,000 tickets in 57s

art_zucker · x · 2026-10-11

FluidInference distilled clef-flash (9B) into a 0.6B text decision model built on Qwen3-0.6B, released as a Core ML package that runs entirely on Apple's Neural Engine.

Original post →

More from Infra

Infra channel →