Same Model, Different App: VS Code Fumbles, Codex Nails It

dfinke · x · 2026-08-25

A test using the same LLM and complex prompt revealed significant performance differences between apps: VS Code fumbled the instructions, while the Codex app executed them perfectly, highlighting the impact of the 'harness' on model behavior.

Original post →

More from Apps

Apps channel →