MENTAT boosts reasoning-intensive regression by up to 65%, Khattab team paper shows
lateinteraction · x · 2026-09-20
- lateinteraction shares the paper Reasoning-Intensive Regression by Diane Tchuindjo and Omar Khattab, and floats two extensions: letting the model think longer before harder decisions, and supporting composite types like dicts and lists instead of only booleans.
- The paper defines RiR: deducing subtle numerical scores from text (rubric scoring, dense rewards, domain-specific retrieval) with limited training data.
- Four realistic RiR benchmarks show that prompting frozen LLMs and fine-tuned encoders both struggle. MENTAT, combining batch-reflective prompt optimization with neural ensemble learning, improves up to 65% over both baselines, with substantial headroom.
- Latimer also speculates RL-for-reasoning followed by a regression head could beat verbalized regression.
Related event: Smarter Thinking Could Reverse AI's Jevons Paradox, Meta Researcher Argues(3 posts)→
More from Research
- New research traces distillation length inflation to student-teacher EOS token mismatch — tw_killian · 2026-09-20
- XGEN Labs unveils generative world simulation JING+DAO, tops WBench leaderboard — hey_abusiddik · 2026-09-20
- Schmidhuber: LLMs aren't truly creative because they lack compression progress — SchmidhuberAI · 2026-09-20
- Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline — This_Cell_1829 · 2026-09-20
- François Fleuret nicknames his training curves; researchers admit they curse baselines too — giffmana · 2026-09-20
- HN: How OpenAI Used Its Own LLMs to Design Its Jalapeño Chip — petrusenko_max · 2026-09-20