600M decision model hits 224 moderation ops/sec on Apple Neural Engine, 95.4% human agreement
huggingface · x · 2026-10-08
FluidInference ported Liquid AI's d1-omni-600M decision model to Core ML for local Apple Silicon inference:
- Performance: moderates 224 comments/sec on an M5 Pro using GPU + ANE across 5,000 real Civil Comments
- Accuracy: 95.4% agreement with human raters, matching the PyTorch original
- Architecture: LFM2.5 encoder trunk plus a decision head answering yes/no, choice, and score questions in one forward pass with zero generated tokens; 99.1% of ops run on the Neural Engine
- Pitfall: the ANE's fused SiLU is 1.4% off, which compounded over 16 SwiGLU MLPs and flipped 41/492 decisions; rewriting SiLU in a tanh-equivalent form reduced flips to 2/492
- Model and Swift SDK are available on Hugging Face
More from Infra
- Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory — _akhaliq · 2026-10-09
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09