A 10-question regime-discrimination benchmark is redesigned after answer leakage

Local-Reading-1624 · reddit · 2026-07-25

This Reddit post shares the evaluator version of a 10-question “regime discrimination” test, rewritten after a contamination issue was discovered in v1.

What changed

Example structure of the test

The post includes a sample question about AI industry CAPEX: a cloud AI company reported record GPU utilization, yet guided lower next-quarter data-center CAPEX while its stock rose slightly.

The evaluator is asked to generate multiple competing explanations instead of collapsing too early into one story, and to identify the first signal that would distinguish them.

Scoring framework

The rubric emphasizes:

It is essentially a benchmark design document for testing whether a model can hold multiple causal hypotheses in mind under uncertainty.

Original post →

More from Research

Research channel →