Open Replications Miss the Secret Sauce: HF Collection Curates Datasets for Jev-Style Models

vanstriendaniel · x · 2026-09-24

The author argues open Jev replications underweight the "not-so-secret sauce"—the data—and has curated a Hugging Face collection of open datasets for training and evaluating Jev-style decision models, with notes on label provenance (humans, exact rules, or an LLM teacher). Highlights: jev-bench (166k samples, human labels, 22 public classification/NLI/rating sets recast as choice/score tasks, some with annotator distributions for calibration testing), typed-decisions (the most-used shared benchmark, soft probabilistic labels, 1.2k train/400 test, PT/ES translations), and procedural-typed-decisions (220k procedurally generated states with noise-free rule-computed labels, ideal for reasoning over state). A ready-to-use index for anyone reproducing or benchmarking Jev-style models.

Original post →

More from Models

Models channel →