Preprint proposes ACI to measure how alignment changes LLM triage decisions
zakkohane · x · 2026-07-24
A preprint from Isaac Kohane studies how effectively alignment works for categorical decisions in LLMs.
- The paper proposes the Alignment Compliance Index (ACI) as a simple way to measure whether a model can be aligned to a target preference function or gold standard.
- Using simulated patient triage pairs, it evaluates GPT-4o, Claude 3.5 Sonnet, and Gemini Advanced before and after alignment attempts.
- Results show alignment effectiveness varies widely across models and prompting approaches; in some cases, pre-alignment models did better than post-alignment ones.
- The authors also probe the implicit ethical principles behind model choices and argue that near-term alignment needs practical measurement tools like ACI.
More from AGI Musings
- AI can generate anything to taste, but the post says culture still lives in shared context — matdryhurst · 2026-07-24
- Zvi says you can’t call a model misaligned without knowing its system prompt — TheZvi · 2026-07-24
- OpenAI could become the intelligence layer for dozens of robot brands — VraserX · 2026-07-24
- AI is collapsing coordination costs, but firms may still outlast the market — prasanna_says · 2026-07-24
- Some kids are calling AI creepy, disgusting and even “artificial idiot” — Wired AI · 2026-07-24
- What would an Amazon for AI agents actually sell? — bhakkimlo · 2026-07-24