ARGUS: An LLM Pipeline Audits Causal Identification Assumptions, Catching 73% of Planted Flaws
Yonghong Zhang · hf · 2026-09-28
Yonghong Zhang introduces ARGUS, a structured LLM pipeline that audits the evidence supporting identification assumptions in difference-in-differences studies of climate policy.
- Method: compares reported evidence against an eleven-dimension assumption-implication-evidence rubric, abstaining when relevant evidence can't be retrieved
- Results: detects 73% of planted flaws on an 11-flaw benchmark vs 18% for a keyword pipeline; abstains on 40% of paper-dimension assessments across 26 economics papers
- Honest calibration: in a five-paper pilot with dual-annotator labels, ARGUS assigns higher risk than labels on 25 of 33 completed assessments; a pre-registered rule removes most of this in-sample, though weighted agreement remains low
ARGUS produces evidence-linked risk reports for expert review without adjudicating causal claims. Code and data are open-sourced.
More from Research
- Cyber Index Alliance Launches With IBM, NVIDIA and Vercel to Standardize AI Cyber Defense Eval — ArtificialAnlys · 2026-09-28
- Artificial Analysis Details Cyber Index Methodology: Safety Refusals Score Zero, Tracked Separately — ArtificialAnlys · 2026-09-28
- Nissenbaum paper takes on privacy nihilism as AI inference erodes data-category frameworks — AllThingsApx · 2026-09-28
- Princeton researchers warn AI could slow science despite exploding paper output — AllThingsApx · 2026-09-28
- One-arm VR intervention during bimanual DAgger feels like an AI brain chip — neurosp1ke · 2026-09-28
- NeurIPS-rejected paper shows six agent dimensions to measure after benchmark saturation — random_walker · 2026-09-28