Open Replications Miss the Secret Sauce: HF Collection Curates Datasets for Jev-Style Models
vanstriendaniel · x · 2026-09-24
The author argues open Jev replications underweight the "not-so-secret sauce"—the data—and has curated a Hugging Face collection of open datasets for training and evaluating Jev-style decision models, with notes on label provenance (humans, exact rules, or an LLM teacher). Highlights: jev-bench (166k samples, human labels, 22 public classification/NLI/rating sets recast as choice/score tasks, some with annotator distributions for calibration testing), typed-decisions (the most-used shared benchmark, soft probabilistic labels, 1.2k train/400 test, PT/ES translations), and procedural-typed-decisions (220k procedurally generated states with noise-free rule-computed labels, ideal for reasoning over state). A ready-to-use index for anyone reproducing or benchmarking Jev-style models.
More from Models
- Fable 5.1 and Astra both post perfect scores on Mensa Norway IQ test — Chemical-Agency-3997 · 2026-09-24
- User claims MiniMax H3 was trained on The Will Stancil Show after model adds unprompted whistle — Kyrannio · 2026-09-24
- Xiaomi's MiMo V2.6 Pro builds a habit tracker in 62 seconds for under 5 cents — socialwithaayan · 2026-09-24
- Gemini 4 Is Almost Ready, Says New Google DeepMind Chief Koray Kavukcuoglu — The Verge AI · 2026-09-24
- New chat, no refusal: user shows ChatGPT's self-image prompt only blocked in original thread — Todesluke · 2026-09-24
- GPT-6 Astra Fixes a UI That Even Opus 5.5 Couldn't Crack — haltakov · 2026-09-24