StartLux open-sources Decision model, beating Jev on 31 of 38 benchmarks

机器之心 · wechat · 2026-10-03

Shanghai startup StartLux (founded less than 5 months ago) released StartLux-Decision, a fully open-source decision model in 0.8B/2B/4B/9B/27B sizes plus quantized builds. On DecisionIndex 0.2.1 (38 benchmarks), the 27B version beats Jev 1.13 on 31 items, 63.88 vs 57.91 composite (team self-tested); it also won 35 of 36 chess games against Jev at comparable latency.

What it does: instead of open-ended generation, Decision models directly pick options, answer yes/no, or score candidates — team routing, tool selection, next browser action. These high-frequency small judgments become the most-called layer in long-running agents.

Speed & engineering: on one H200 (BF16), the 4B answers three questions in 26ms average per forward pass (0.8B: 12.2ms), using CUDA Graph, fast linear-attention kernels, and merged multi-question forwards. It outputs probability distributions for confidence-based routing and human-in-the-loop fallbacks. Q4KM quantization (2.71GB for 4B) retains 98.3% decision agreement with original weights.

Layered local intelligence: StartLux stacks a 27B general model, Decision models for high-frequency calls, Agent + Memory on top, and quantization/inference systems below — all optimized for constrained local hardware.

AutoResearch & RSI: the model went from direction-setting to first validated version in 3 days via an automated research pipeline (95%+ engineering automation, 70% AI participation in core research decisions). CTO Guo Quanwei outlined a 5-level RSI framework, placing the company at Level 3. The piece notes industry consensus forming around dedicated decision models, citing OpenAI's Decisions API and Cloudflare's open-source Clef.

Original post →

More from coding & agent

coding & agent channel →