MIT economist challenges Aaru's human-simulation benchmarks: opaque method, no baseline, data leakage risks
soumitrashukla9 · x · 2026-09-27
LLM human-simulation startup Aaru claimed industry-leading results across 2,993 questions (mean TVD 7.62%, MAE 3.53%). MIT economist Andrey Fradkin ran the analysis past ChatGPT and Refine as AI referees and flagged opaque methodology, no baseline-model hurdle, potential data leakage, and unjustified statistical claims. Until peer review forces the authors to address these, skepticism is warranted.
More from Research
- Xiaomi publishes MiMo-V2.6 paper on scaling reinforcement learning toward LLM-Core — KyeGomezB · 2026-09-27
- The brain is a predictive machine: remove reality's correction signal and it hallucinates — alfcnz · 2026-09-27
- Colosseum paper on auditing collusion in multi-agent systems accepted at NeurIPS 2026 — niloofar_mire · 2026-09-27
- MIT talk explains LLMs from first principles, no transformers needed — vishalmisra · 2026-09-27
- Microsoft's ProgramDistill turns interactive web apps into verifiable SWE training tasks — _akhaliq · 2026-09-27
- MIT's pseudorandom codes survey maps the crypto primitive powering AI content watermarks — matthew_d_green · 2026-09-27