A 10-question regime-discrimination benchmark is redesigned after answer leakage
Local-Reading-1624 · reddit · 2026-07-25
This Reddit post shares the evaluator version of a 10-question “regime discrimination” test, rewritten after a contamination issue was discovered in v1.
What changed
- The respondent and evaluator files were split so the answer key is never shown to the model.
- The question text itself was kept unchanged after review; the main fix was procedural separation.
- The author records that earlier model outputs appeared to mirror answer-structure language too closely, so the previous round was discarded as contaminated.
Example structure of the test
The post includes a sample question about AI industry CAPEX: a cloud AI company reported record GPU utilization, yet guided lower next-quarter data-center CAPEX while its stock rose slightly.
The evaluator is asked to generate multiple competing explanations instead of collapsing too early into one story, and to identify the first signal that would distinguish them.
Scoring framework
The rubric emphasizes:
- avoiding immediate convergence on the surface conclusion;
- generating at least two genuinely different causal regimes;
- keeping the causal chains distinct;
- separating leading and lagging indicators;
- not inventing facts;
- checking whether the analytical axis itself was post-hoc constructed.
It is essentially a benchmark design document for testing whether a model can hold multiple causal hypotheses in mind under uncertainty.
More from Research
- Stanford’s 457-page AI Index 2025 report tracks falling costs, efficiency gains and adoption — mdancho84 · 2026-07-25
- Rice Gene ROAD1 Significantly Boosts Drought Tolerance in Multiple Crops Without Yield Loss — NikoMcCarty · 2026-07-25
- Live demo on RoboPapers shows the system working in a hotel room in Korea — chris_j_paxton · 2026-07-25
- Reddit post lays out a 10-question eval to catch prompt leakage in AI tests — Serofa20125 · 2026-07-25
- A 10-question LLM regime-reasoning test checks ambiguity, competing hypotheses, and falsifiability — Local-Reading-1624 · 2026-07-25
- An essay links compression and intelligence to mark Ray Solomonoff’s 100th birthday — ryangr · 2026-07-25