Day 1 with Grok 4.7: strict system-prompt adherence and visible gains over 4.5 in real coding work
elonmusk · x · 2026-09-22
A developer shares day-one observations after using Grok 4.7 all day as his coding "firstmate," pushing back on negative reports based solely on public benchmarks—which he calls useless, noting the same benchmarks misjudged other models.
Key finding: Grok 4.7 follows system prompts far more closely, showing new behaviors like asking which red CI checks the user is OK bypassing and refusing a simple "yolo" instruction—traced directly to his own system prompt, something no other model did.
Overall: visible improvements over 4.5 (he skipped 4.6 after it underperformed). The post was retweeted by Elon Musk.
More from coding & agent
- Prof. Tom Yeh releases printable hand-calc agentic AI cost math problems — ProfTomYeh · 2026-09-22
- apibase unifies 327 tools from 92 providers behind one pay-per-call MCP endpoint — modelcontextprotocol · 2026-09-22
- mcp-server-s3 ships MCP server letting agents browse, upload and share S3 files — modelcontextprotocol · 2026-09-22
- Dev built an AI game-playing plugin but shelved it: vision models too slow and costly — ezshine · 2026-09-22
- Developers debate persistent file storage options for AI agents across runs — OwlZealousideal4779 · 2026-09-22
- Laya: 11.6k-star open-source engine outputs typed decisions in 33ms, no generation — pandeyparul · 2026-09-22