GLiNER2.5-Decide ported to CoreML: 4x faster, 5x less peak RAM, half the size
BLUECOW009 · x · 2026-09-26
A developer converted GLiNER2.5-Decide to CoreML, achieving roughly 4x faster inference, 5x lower peak RAM, and half the model size. Both the converted model and the conversion code are open-sourced, useful for on-device deployment on Apple hardware.
More from Infra
- Samsung, Oxford and PKU propose TrOPD to distill frontier-model reasoning into on-device small models — jiqizhixin · 2026-09-26
- Germany and the Netherlands put €40m into a challenge to design AI chips with AI — VraserX · 2026-09-26
- AMD Claims One Factory Box Can Run 2.3x More Software Workers — shashib · 2026-09-26
- Token Prices Fall via Subsidies, Moore's Law via Manufacturing — Different Drivers — HanchungLee · 2026-09-26
- KV cache transplants on Qwen3.8-27B: start at Q6, hand off to Q3, beat static quant — wadeAlexC · 2026-09-26
- Google open-sources GKE agentic migration for AI-assisted EKS-to-GKE moves — rseroter · 2026-09-26