Community benchmark: Mercury Decide hits 66.7% accuracy, trailing Jev's 83.3%

FlexDruk · reddit · 2026-10-01

A developer benchmarked the newly released Mercury Decide using an uncontaminated, Korean-focused dataset (ReportJevBench) built around deciding whether Roblox chat reports violate ToS. Accuracy: Jev 83.3%, Kev 72.2%, Mercury Decide 66.7%, Solar Decide 57.8%. At a 0.5 threshold Mercury Decide produced 28 false negatives out of 90 versus Jev's 3; even after threshold tuning it leaned toward rejecting nearly every report, performing worse than Kev. The author concludes no API decision model is currently more stable than Jev, while noting the benchmark's narrow scope (Korean understanding + Roblox ToS).

Original post →

More from Models

Models channel →