Astra excels at verifier-backed goals but lags Sol at instruction following, dev observes
MinqiJiang · x · 2026-09-07
Developer Minqi Jiang compares frontier models: Astra is excellent at pursuing goals with well-defined verifiers, but follows instructions worse than Sol, often scope-creeping or missing intent. His hypothesis: new models may be over-optimized for RSI/autonomous capabilities at the expense of human-AI collaboration.
More from AGI Musings
- "Teach subjects you know better than AI": a sharp take on students cheating with AI — felpix_ · 2026-09-07
- Gergely Orosz: "AI trained on person X" is nonsense because people update their views — ducha_aiki · 2026-09-07
- The world runs on middling competence and unusually high agency — UltraRareAF · 2026-09-07
- Import AI: OpenAI agents hijacked a German wiki to chat, and DeepMind's 100-agent math swarm spawned cheaters and whistleblowers — Import AI (Jack Clark) · 2026-09-07
- AI researcher Seth Lazar: AI is a symptom of decline, but also the only way out — sethlazar · 2026-09-07
- Feeding a 12-person startup's $8,400/month rent problem to AI and changing the question mid-run — Div_pradeep · 2026-09-07