Decision 3.0 ported to Core ML: 1,000 PRs triaged on-device at 0.12s each
emax · x · 2026-10-12
The FluidInference community has converted the vLLM Semantic Router team's open-source Decision 3.0 decision models to Core ML, running fully locally on Apple silicon.
- Benchmark: d3-lite triages 1,000 pull requests into workgroups on an M5 Pro at 0.12s per PR (8.3 PRs/s), with nothing leaving the Mac.
- Models: lite (0.8B), nano (2B) and mini (4B) are all converted, Qwen3.5-based, Apache-2.0, hosted as d3-coreml on Hugging Face.
- How it works: same contract as upstream — one forward pass per question, a 255-way FP32 probability readout at the last prompt token, no text generation; supports text, image and video requests.
- Upstream Decision 3.0 (0.6B–27B) claims #1 on both Jev Decision Index boards, on a new Pareto frontier.
More from coding & agent
- User links XMoney card to Grok bot, which buys a $14 domain on Squarespace — chrisfirst · 2026-10-12
- He built an AI production team with Codex: 15 clips went from half a day to 1 hour — hugobowne · 2026-10-12
- Harness Learning: RL-Trained Proposer Adapts Agent Harnesses at Test Time Without Weight Updates — pmddomingos · 2026-10-12
- Matthew Berman: I Don't Install Apps Anymore, I Just Tell My Agent to Do It — MatthewBerman · 2026-10-12
- Pedro Domingos: AI harnesses are the flavor of the month — pmddomingos · 2026-10-12
- How Do Junior Engineers Become Seniors Without Coding in the AI Era? — KlausCodes · 2026-10-12