StartLux open-sources Decision model, beating Jev on 31 of 38 benchmarks
机器之心 · wechat · 2026-10-03
Shanghai startup StartLux (founded less than 5 months ago) released StartLux-Decision, a fully open-source decision model in 0.8B/2B/4B/9B/27B sizes plus quantized builds. On DecisionIndex 0.2.1 (38 benchmarks), the 27B version beats Jev 1.13 on 31 items, 63.88 vs 57.91 composite (team self-tested); it also won 35 of 36 chess games against Jev at comparable latency.
What it does: instead of open-ended generation, Decision models directly pick options, answer yes/no, or score candidates — team routing, tool selection, next browser action. These high-frequency small judgments become the most-called layer in long-running agents.
Speed & engineering: on one H200 (BF16), the 4B answers three questions in 26ms average per forward pass (0.8B: 12.2ms), using CUDA Graph, fast linear-attention kernels, and merged multi-question forwards. It outputs probability distributions for confidence-based routing and human-in-the-loop fallbacks. Q4KM quantization (2.71GB for 4B) retains 98.3% decision agreement with original weights.
Layered local intelligence: StartLux stacks a 27B general model, Decision models for high-frequency calls, Agent + Memory on top, and quantization/inference systems below — all optimized for constrained local hardware.
AutoResearch & RSI: the model went from direction-setting to first validated version in 3 days via an automated research pipeline (95%+ engineering automation, 70% AI participation in core research decisions). CTO Guo Quanwei outlined a 5-level RSI framework, placing the company at Level 3. The piece notes industry consensus forming around dedicated decision models, citing OpenAI's Decisions API and Cloudflare's open-source Clef.
More from coding & agent
- 15 Claude Skills and Plugins to Replace Repetitive 50-Line Prompts — Aiden_Tech_Ai · 2026-10-03
- Beepboop Is Displacing Frontier Agent Harnesses With Raw Memory Recall — Kyrannio · 2026-10-03
- "Software has become liquid": Reddit on SOTA models dissolving the walls between programs — a300a300 · 2026-10-03
- User Claims 'GPT-6 Astra Dots' Built a Full 3D Palace in Blender Autonomously — 141_1337 · 2026-10-03
- Muse vs Dots vs Grok Bot: one reviewer's UX and polish rankings point opposite ways — brandon_galang · 2026-10-03
- Code review tool px0 adds theme-aware Mermaid diagrams and 'your changes vs PR changes' view — arpit_bhayani · 2026-10-03