Community benchmark: Mercury Decide hits 66.7% accuracy, trailing Jev's 83.3%
FlexDruk · reddit · 2026-10-01
A developer benchmarked the newly released Mercury Decide using an uncontaminated, Korean-focused dataset (ReportJevBench) built around deciding whether Roblox chat reports violate ToS. Accuracy: Jev 83.3%, Kev 72.2%, Mercury Decide 66.7%, Solar Decide 57.8%. At a 0.5 threshold Mercury Decide produced 28 false negatives out of 90 versus Jev's 3; even after threshold tuning it leaned toward rejecting nearly every report, performing worse than Kev. The author concludes no API decision model is currently more stable than Jev, while noting the benchmark's narrow scope (Korean understanding + Roblox ToS).
More from Models
- Gemini Argon undercuts Astra 5x on price, but coders say Opus 5.5 still wins — eyishazyer · 2026-10-01
- RLHF co-author Diogo Almeida discusses the philosophy behind TypeSafe's Jev model — AxSaucedo · 2026-10-01
- Early user: GPT-6 Sol feels slower and dumber than 5.6 in Codex — burkov · 2026-10-01
- GPT-6.1 Sol plots a 3-day Earth-Mars transfer orbit, but a Codex app update locked the session — MikePFrank · 2026-10-01
- Codex Pro Users Report Credits Vanishing Before Included Usage Is Exhausted — Constant-Potential-9 · 2026-10-01
- Qwen Flash Next MTP work resumes with official GGUF quants and llama.cpp PR — jacek2023 · 2026-10-01