Testing AI Work Assistants on Trip Planning: ChatGPT and Claude Both Fumble Hotel Availability

giffmana · x · 2026-09-16

giffmana benchmarked AI work assistants on the canonical task "plan a trip from today's email" with all connectors enabled. ChatGPT Work browses well and finds trains/hotels, but fails to check room availability and prices, then gives up. Claude CoWork handles trains fine but also completely flukes hotel availability (no browser use). Gemini (still on Flash 3.6) is noisy and picks a suboptimal hotel initially, yet uniquely pulls real availability/prices via Google Travel — though it mid-conversation suggests Google flights to arbitrary destinations. Surprising how much OpenAI and Anthropic fumble this super canonical use case.

Related event: AI Travel Assistant Showdown: ChatGPT and Claude Fall Short(2 posts)→

Original post →

More from Models

Models channel →