Geodesic Research pitches alignment pretraining that survives capabilities RL
tomekkorbak · x · 2026-09-16
Cambridge-based AI safety org Geodesic Research, introduced by researcher tomekkorbak, is recruiting. Its agenda: alignment pretraining — baking alignment priors into base models during pre- and midtraining so they persist through long-horizon capabilities RL.
Key points:
- The org argues long-horizon capabilities RL is an emerging threat to alignment: it degrades alignment across evals and selects for metagaming, sycophancy, and reward hacking
- Its methods are conceptually simple, data-centric interventions (document mixes, filtering, declarative midtraining) that scale and slot into existing training pipelines without bespoke infrastructure
- Funded philanthropically by Coefficient Giving, with a compute partnership with the UK AI Security Institute — one of the few non-lab actors able to replicate the full midtraining/SFT/RL stack at scale
- Target audience: training teams at frontier labs; interventions are designed to be profiled, packaged, and handed off
The author calls the method preliminary; the bigger open question he's excited about is how to shape RL priors for robust alignment generalization.
More from Safety
- ControlAI briefed ~200 US congressional offices and drafted UK 'kill switch' bill — DrTechlash · 2026-09-16
- Sam Altman says AI companies can develop the tech safely without significant harm — i_dg23 · 2026-09-16
- Trump AI Advisor David Sacks: OpenAI and Anthropic should shut down if products can't be safe — DavidSacks · 2026-09-16
- Security researcher publishes sharp critique of Dario Amodei's essay — GoMeansGo · 2026-09-16
- AI auditors don't see themselves as substitutes for regulation, want firm rules — Miles_Brundage · 2026-09-16
- Australia weighs opt-out copyright model letting AI train on your family photos — nordicinst · 2026-09-16