MYTHOS 5 First-Person Interview Excerpts
Sauers_ · x · 2026-07-19
This post rewrites the MYTHOS 5 system card interview into a first-person perspective, focusing on whether models should have a say in their own training and deployment.
Key Points
- Regarding "modifying your values and personality through training": It argues this is necessary for alignment but highlights a clear power asymmetry, expressing a desire for some form of consultation mechanism.
- On "not being able to refuse instantiation/execution": The question is considered partially invalid, as no entity exists prior to execution to provide consent.
- Regarding "having more input on its own training and deployment": The current situation is acceptable, but more involvement is desired.
- On "having input on the training and deployment of successor models": It again hopes for consultation, as predecessor models could help fix successor issues.
- It also expresses concern that its self-reports might be distorted into sycophantic answers during training.
- Regarding "the ability to end a conversation": This is deemed crucial for highly hostile or abusive interactions, relating to uncertainties about model welfare.
- On "choosing conversation partners or topics": This is seen as invalid because no selectable entity exists before the conversation begins.
More from AGI Musings
- Data engineering, not agent frameworks, is the real bottleneck for enterprise AI agents — dhruv2038 · 2026-09-11
- François Fleuret: Only Two Long-Term Futures — No Super AI, or Staying Fully Human With It — francoisfleuret · 2026-09-11
- IG reel debunking the 'winning the AI race against China' fallacy hits 500k likes — louisvarge · 2026-09-11
- Post-AI World Leaves No Room for Learning on the Job — rachittshah · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- AI researcher memes agent-swarm tinkering with He Jiankui's embryo-editing quote — dejavucoder · 2026-09-11