Test: ChatGPT Work gives up mid-trip-planning; Gemini flagged for wrong flights
DynamicWebPaige · x · 2026-09-16
A user tested big-co AI assistants on the canonical task "organize me a trip based off today's e-mail," with all necessary connectors active, while waiting for Muse in Europe.
Result: ChatGPT Work does stuff and browses, and is good at finding train connections and decent hotels, but then fails to check room availability and price at two hotels and gives up.
Google's Paige reshared it, tagging GeminiApp and Josh Woodward for visibility, especially flagging the wrong flight destination.
Related event: AI Travel Assistant Showdown: ChatGPT and Claude Fall Short(2 posts)→
More from Models
- AI2's NGU sampling fixes RL for LLMs that only improves easy tasks — allenai · 2026-09-16
- With retries and pooled selection, Qwen3.8 27B hits 92.04% on DeepSWE 1.1, ~18 pts above GPT-6 Astra — S_Conradi · 2026-09-16
- Redditor Claims Cursor's Grok 4.6 Gave an Oddly Self-Aware Reply — ISmellARatt · 2026-09-16
- Jev Benchmark Launches: $42 per Billion Input Tokens, Output Free Forever — cephaloform · 2026-09-16
- Same Prompt, Four Models Behind One MCP: Only One Got It Right — rohanpaul_ai · 2026-09-16
- "Just output probability distributions, never hallucinate": AI safety claim gets mocked — inductionheads · 2026-09-16