Solar Decide scores 57.8% vs Jev's 83.3% on a custom benchmark of real moderation reports
FlexDruk · reddit · 2026-09-29
Korean company Upstage released Solar Decide (based on Solar Mini 4), but the author tested it in their Roblox game's AI chat-report pipeline (mostly Korean players) and rolled it back:
- On generic JevBench: Solar Decide 88.9% vs Jev 98.6%;
- Using 90 human-labeled cases out of 352 real reports, an agent built a custom benchmark (ReportJevBench): Solar Decide scored only 57.8% vs Jev's 83.3%;
- The gap is false positives: 11/90 for Jev vs 37/90 for Solar Decide — wasted tokens and a defeated purpose; on day one Solar Decide validated a false report.
Dataset and results are open-sourced on GitHub (Korean, needs translation).
More from Models
- Leaked OpenAI 'dot' details show raising phone to ear triggers ChatGPT Voice — koltregaskes · 2026-09-29
- Google to replace Gemini Gems with Skills starting November 17 — mark_k · 2026-09-29
- Carla v0.1.0: a local llama.cpp loom TUI for growing AI characters — max_paperclips · 2026-09-29
- Leaked OpenAI DevDay reveal called 'just a Grok bot / Meta Muse rip-off' — gaganghotra_ · 2026-09-29
- Running Jev at high frame rate with full-state snap inferences makes it a true System 1 — mathemagic1an · 2026-09-29
- OpenAI model naming rumor: Dots, Orbit and 'o' said to be in the mix — mark_k · 2026-09-29