APO Enables Personalized LLM Alignment with Only 20 Local User Examples
Liyan Yang · hf · 2026-10-07
To address heterogeneous user preferences and scarce per-user feedback in LLM alignment, the authors propose Approximate Pareto Optimality (APO). The method groups users with compatible updates to reduce interference, then combines gradient descent with controlled ascent within each group to coordinate competing objectives and approach preference-specific Pareto front points, producing an initialization effective for few-shot personalization, refined iteratively with local updates. They derive conditional suboptimality bounds for one-local-step collaborative updates and characterize how initialization error affects subsequent adaptation. Experiments on Fed-ChatbotPA and UltraFeedback show consistent gains over prior methods using only 20 local examples.
More from Research
- An excellent overview of AI watermarking and why it can't really be avoided — aronchick · 2026-10-07
- Self-improving agent Steve masters Minecraft unaided, out-progressing 97% of human players — julianweisser · 2026-10-07
- LLM2Vec-Gen: frozen LLMs generate answer embeddings in one forward pass, SOTA self-supervised — sivareddyg · 2026-10-07
- Researchers pitch World Editing: modifying existing worlds instead of generating new ones — yuntiandeng · 2026-10-07
- New paper asks: when agents act for you, whose side are they on? — ZacharyHuang12 · 2026-10-07
- AI's Top 10 research list: Spurious Rewards tops RL-heavy ranking — ShayneRedford · 2026-10-07