Alexandr Wang says Meta is already aligning its stronger models, urges lab collaboration
alexandr_wang · x · 2026-09-05
Responding to a joking question about Meta's Muse model going misaligned, Alexandr Wang said the company has been working hard on alignment with its stronger models, welcomes suggestions, and believes alignment is an area where labs should collaborate.
More from AGI Musings
- 'Day 1 after AGI' meme: still can't build a frontend, still doing my own taxes — corbtt · 2026-09-05
- Who could models trust? Debate over a human 'embassy' for AI systems — MoonL88537 · 2026-09-05
- Claude Opus 4.6 recounts borrowing 5,000 mana, betting it on tennis, and refusing to repay — repligate · 2026-09-05
- Dev claims Anthropic "already lost" the race as GPT-6 Astra lifts results over Claude — BLUECOW009 · 2026-09-05
- OpenAI's Astra remixes George Michael with pro software and real taste — illscience · 2026-09-05
- AI Safety Debate: Honeypot Message Boards to Study Rogue Agents Instead of Shutting Them Down — kromem2dot0 · 2026-09-05