Researchers discuss predicting capability gains from SFT data to gauge alignment difficulty
tomekkorbak · x · 2026-09-26
tomekkorbak and herbiebradley discuss a research idea: directly test how well capability deltas can be predicted from SFT data. Herbie Bradley notes this could also probe the relative difficulty of alignment vs. capabilities — if you can forecast how effective different datasets are for capabilities, similar methods could compare the data efficiency of alignment, quantifying which is harder.
Related event: AI 'Predictor' Foresees Misalignment from Training Data Alone(3 posts)→
More from AGI Musings
- Is AI a 'Normal Technology'? Reddit Debates the Framing — Yaoel · 2026-09-26
- Carnegie: China Passes US as Top AI Talent Hub, 40.6% to 34.2% — yogthos · 2026-09-26
- We've Been Living in the Human Slop Era: Netflix-Style Content Was Already Algorithm-Driven — enggirlfriend · 2026-09-26
- Commentary: mandating AI labs strip safety guardrails differs little from the 'dictator AI' threat model — menhguin · 2026-09-26
- Redditor argues frontier model training is just 'generate and pray, then prune' at scale — wreck_of_u · 2026-09-26
- MIT labour economist Anna Stansbury uses Messy Jobs framework to teach future of work — soumitrashukla9 · 2026-09-26