Peter Yang tests ChatGPT, Claude, Grok and Gemini across 10 use cases
nickbaumann_ · x · 2026-08-27
Peter Yang published a 26-minute tutorial testing ChatGPT, Claude, Grok, and Gemini head-to-head across 10 real use cases — everyday answers, writing, coding, browser use, personal agents, image generation, and video.
- ChatGPT had an incredible year and is the best all-around pick for daily answers, writing, coding, browser, voice, and images; solid for personal agents too, though the Work vs. Codex UX remains confusing.
- Claude is still his favorite for design and planning (including slides and videos, where Fable shines), but he feels Claude regressed in personality and writing — Opus 5 is frustrating to talk to with too many Claude-isms.
- Grok is the most interesting challenger: engineering and solopreneur friends prefer it over GPT Sol for coding, and Grok @bot leads in personal agent innovation with its persistent, dedicated computer.
A few months ago he used Claude for almost everything; his usage mix now looks very different.
More from Models
- Zhipu GLM-5.3-Flash scores 63% on DeepSWE at $0.24 per task — AccBalanced · 2026-08-27
- Gemini 2.0 Flash vs Qwen2.5 Flash: Head-to-Head Comparison — ryanmerket · 2026-08-27
- New Codex reasoning effort "Persistent" spotted in GitHub repo — zephyr_z9 · 2026-08-27
- Nvidia engineer credits DeepSeek for bridging the AI gap — thursdai_pod · 2026-08-27
- Matt Shumer Tests H3 Max: "Really Damn Good," and It Has Thoughts on Instinct — mattshumer_ · 2026-08-27
- Clarification: Ox Alpha is SOTA but not a Flash model — ChrisGPT · 2026-08-27