5 LLMs across 10 auctions: models preserve human bidders' mechanism-level orderings
soumitrashukla9 · x · 2026-10-03
- Researchers from MIT, Harvard, Google Research and OpenAI test 5 LLMs out of the box across 10 auction settings against human experimental benchmarks (paper: arXiv 2507.09083), focusing on GPT-4o, Claude 3.5 Haiku and Gemini 2.0 Flash.
- LLM and human deviations from theory differ in magnitude and direction: humans overbid in second-price auctions, while most models underbid.
- Surprisingly, without fine-tuning, large models robustly preserve key orderings of auction formats: first-price auctions are harder than second-price, and ascending clocks reduce deviations vs sealed bids. Kendall's τb between human and GPT-4o difficulty rankings is 0.60.
- The reasoning model bids nearly at equilibrium in observed private-value settings.
More from Research
- DeepMind's Lampinen: LMs may make 'intentional' discoveries of new math concepts — AndrewLampinen · 2026-10-03
- UT Austin professor rewrites Feature Ranking chapter in free ML e-book — GeostatsGuy · 2026-10-03
- New survey maps 237 methods for embedding physics priors in robot learning — Jan_R_Peters · 2026-10-03
- Kevin Buzzard: AI is solving math humans can't, and mathematicians are grieving — stevenstrogatz · 2026-10-03
- SkillRefiner: Offline skill refinement from your agents' existing execution traces — gregd_nlp · 2026-10-03
- Open pretraining run matches Llama 3.2 1B on ARC-C at ~10% of the cost, author details 4 pitfalls — jon_durbin · 2026-10-03