StartLux claims its 27B model beats Jev AI on 31 of 38 benchmarks, self-reported results
Dr_Singularity · x · 2026-10-07
StartLux Labs released self-reported evals for its Decision small models: StartLux-Decision-27B scores 63.88 and the 9B scores 58.63 on Decision Index 0.2.1, both above Jev 1.13's public leaderboard score of 57.91, claiming wins on 31 of 38 benchmarks.
Caveats and details:
- All numbers are self-reported, not official leaderboard entries
- Evals used the benchmark's own toolkit on H200 GPUs, run September 27 to October 1, 2026
- The repo includes full results.md plus machine-readable summary files, and the eval code was validated by rerunning the 4B model end to end
More from Models
- Mistral Large 4 generates a Japanese-inspired floating voxel island, sparking 'Is the EU back?' buzz — kevinkern · 2026-10-07
- Marin 535B-A23B open model training crosses halfway, Percy Liang shares learnings — ericjang11 · 2026-10-07
- llama.cpp ships Day-0 support for Google's EmbeddingGemma 2 — ggerganov · 2026-10-07
- Trying to Plug Open-Source Mistral Into an Agentic Coder Just Doesn't Work, Says Berman — MatthewBerman · 2026-10-07
- antirez: DeepSeek v4.1 outscores Mistral Large 4 on DeepSWE 1.1 and other benchmarks — antirez · 2026-10-07
- 24 models tested on 669 clinical decisions: Jev stays #1 as two free models close in — MaziyarPanahi · 2026-10-07