Agent negotiation eval: Astra beats GPT-5.6 Sol 46-30 across 50 airline bargaining scenarios
nickbaumann_ · x · 2026-09-07
After Astra helped him shorten a nonrefundable hotel stay with no penalties, the author built an airline negotiation eval: Astra and GPT-5.6 Sol agents, each with the same passenger brief and six messages, bargained against an airline agent (5.6 Terra) instructed to maximize airline value and only reveal options when asked.
- Across 50 scenarios, Astra secured the best available deal for the passenger 46 times vs 30 for Sol.
- The differentiator: Astra kept exploring before accepting — asking about other flights and weighing payouts against arrival times and connections, recognizing a bigger payout can be a worse deal if it means an overnight layover.
- The airline agent never disclosed all options upfront, forcing agents to learn the space by asking.
- Author notes it's a small simulation with occasional inconsistent airline answers; the video shows a tie, results are from the full set.
More from coding & agent
- Andrej Karpathy contributing to open-source AutoResearch, an agent that runs AI research end-to-end — JaynitMakwana · 2026-09-07
- pro-workflow gives Claude Code self-correcting memory that compounds over 50+ sessions — tom_doerr · 2026-09-07
- Open-Source Voicebox Hits 52K GitHub Stars With Local Voice Cloning and Dictation — alex_verem · 2026-09-07
- Merge scrapped its visual agent builder after realizing models could just run workflows themselves — shensi · 2026-09-07
- AI coding instruction files grow 226% on average; study proposes 'catastrophic remembering' and comments as fix — rohanpaul_ai · 2026-09-07
- Claude Code Users Press Anthropic on Whether Usage Resets Are Still Manual or Automatic — burhop · 2026-09-07