clef-flash distilled into 0.6B model on Apple Neural Engine triages 1,000 tickets in 57s
art_zucker · x · 2026-10-11
FluidInference distilled clef-flash (9B) into a 0.6B text decision model built on Qwen3-0.6B, released as a Core ML package that runs entirely on Apple's Neural Engine.
- Triages 1,000 support tickets in 57 seconds on an M5 Pro (48ms each), outputting team, urgency and refund decisions within a few points of the 9B teacher on classification
- Same API contract as Clef/Jev/SystemOne: one forward pass yields probabilities for every option across typed questions
- Apache-2.0, 1.3GB of fp16 weights, ships with PyTorch weights and a Swift runtime (ClefTextManager)
- Strong on routing/classification, but lags badly on knowledge tasks like ARC-Challenge and CommonsenseQA
More from Infra
- Florida's 3 biggest utilities form alliance to pave way for data centers — 4KTV · 2026-10-11
- TensorFold 1.0.7 writes each learned fact into ~10 new neurons, 4x faster with 3D view — HankYeomans · 2026-10-11
- China building 38 nuclear reactors (~39.8GW) as energy emerges as AI's biggest bottleneck — kimmonismus · 2026-10-11
- Token prices keep falling, yet devs burn more: Jevons paradox hits AI coding agents — daniel_mac8 · 2026-10-11
- Tenstorrent Blackhole folds 4x more per dollar than H200; antibody docking hits 84% top-1 — DavidBennett__ · 2026-10-11
- Firmus Grid pulls ASX listing at $30B valuation as Altman says now ill-advised to IPO — shashib · 2026-10-11