New paper proposes ECT to align LLMs by describing standards
dhadfieldmenell · x · 2026-08-18
Reliably evaluating LLM outputs to an aligned standard is difficult, but describing the standard is easier. A new paper introduces Evaluation Conditioned Training (ECT), which trains LLMs by simply telling them the standard they should hold themselves to.
More from Safety
- Google reportedly buying 175k employee records to train AI models — MatthewChang · 2026-08-18
- Greg Brockman: OpenAI training models for superhuman secure code — AccBalanced · 2026-08-18
- Less than 50 engineers work on 'trust but verify' China tools — wfithian · 2026-08-18
- AI resume screening tool rejected 100% of pre-2010 graduates — SumitGup · 2026-08-18
- SPAR opens applications for Fall research program on AI safety and policy — niloofar_mire · 2026-08-18
- Handshake spotted using '100% legal' method to obtain AI training data — PeterHndrsn · 2026-08-18