HANDRAISER's Per-Task Numbers: 48.9% Cost Cut on Debate, 32.2% on Average
lileics · x · 2026-09-23
From the HANDRAISER thread, per-task results with Llama-3.1-8B listeners: communication cost down 24.3% on Text Pictionary, 23.4% on Meeting Scheduling, and 48.9% on MMLU-Pro Debate — averaging 32.2% lower cost with comparable or better task performance. Full details in the companion post.
More from Research
- Yarin Gal: LLM-written experiments that fail to replicate are no different from any others — yaringal · 2026-09-23
- Yarin Gal: LLM experiments that don't replicate are just failures, and an AI arXiv could help — yaringal · 2026-09-23
- Oxford's Yarin Gal Proposes arXiv Ban LLM-Written Papers to Curb AI Slop — yaringal · 2026-09-23
- New paper proves fundamental confidence-efficiency bounds for transductive conformal prediction — _onionesque · 2026-09-23
- $1B and unlimited frontier tokens: where would you spend them to fix cybersecurity? — chrisrohlf · 2026-09-23
- Grok explains why DeepSeek picked DualPipe + ZeRO-1 over ZeRO-3 on 2048 H800s — TheZachMueller · 2026-09-23