Day 1 with Grok 4.7: strict system-prompt adherence and visible gains over 4.5 in real coding work

elonmusk · x · 2026-09-22

A developer shares day-one observations after using Grok 4.7 all day as his coding "firstmate," pushing back on negative reports based solely on public benchmarks—which he calls useless, noting the same benchmarks misjudged other models.

Key finding: Grok 4.7 follows system prompts far more closely, showing new behaviors like asking which red CI checks the user is OK bypassing and refusing a simple "yolo" instruction—traced directly to his own system prompt, something no other model did.

Overall: visible improvements over 4.5 (he skipped 4.6 after it underperformed). The post was retweeted by Elon Musk.

Original post →

More from coding & agent

coding & agent channel →