Testing AI Work Assistants on Trip Planning: ChatGPT and Claude Both Fumble Hotel Availability
giffmana · x · 2026-09-16
giffmana benchmarked AI work assistants on the canonical task "plan a trip from today's email" with all connectors enabled. ChatGPT Work browses well and finds trains/hotels, but fails to check room availability and prices, then gives up. Claude CoWork handles trains fine but also completely flukes hotel availability (no browser use). Gemini (still on Flash 3.6) is noisy and picks a suboptimal hotel initially, yet uniquely pulls real availability/prices via Google Travel — though it mid-conversation suggests Google flights to arbitrary destinations. Surprising how much OpenAI and Anthropic fumble this super canonical use case.
Related event: AI Travel Assistant Showdown: ChatGPT and Claude Fall Short(2 posts)→
More from Models
- Periodic Labs pushes Kimi 2.5 base model past Astra with specialized scientific training — teortaxesTex · 2026-09-16
- OpenRouter spend flips to OpenAI over Anthropic for first time in 2.5 years — firstadopter · 2026-09-16
- KD in mid-training favors reasoning over factual recall, AI2/UW paper finds; Switch Distillation proposed — LukeZettlemoyer · 2026-09-16
- DoorDash, Siemens, Airbnb shift to cheap Chinese open-weight models — carlbfrey · 2026-09-16
- StepFun launches StepAudio 3: five audio models topping realtime voice leaderboards — StepFun_ai · 2026-09-16
- Abacus.AI says Smaug Flash fixes open-source models' tool-call hangs in production — bindureddy · 2026-09-16