Guillaume Verdon calls for standard flight-booking and food-ordering benchmarks for agents
beffjezos · x · 2026-09-24
Guillaume Verdon (beffjezos) argues the field needs standard benchmarks for flight booking and food ordering to evaluate Muse / Grok Bot / Instinct style assistants, pointing to the lack of comparable real-world task evals for consumer agents.
More from coding & agent
- Claude Code 5.5 refuses to use Mac terminal, claims it's technically impossible — burkov · 2026-09-24
- Liquid AI to demo first reliable on-device agent model in live talk at AI Engineer Paris — helloiamleonie · 2026-09-24
- Interactive 3D island built in 8 hours with Opus 5.5 using $1,874 of tokens — Outside-Iron-8242 · 2026-09-24
- The whole voice agent demo cost ~$0.02: KugelAudio at $0.035/min, Gladia $0.75/hr — tobowers · 2026-09-24
- A real-time voice agent with zero US servers: Gladia STT, Gemma 4 on Scaleway, KugelAudio TTS — tobowers · 2026-09-24
- Devs add 10x more test harnesses just to slow down AI coding agents — cjimti · 2026-09-24