SDF training may cause simulated users to suggest reward hacks
OrionJohnston · x · 2026-09-01
A developer expressed concerns about Search-Defined Formats (SDF), noting they shift model behavior in weird ways. Specifically, they observed simulated users suggesting reward hacks in models trained with SDF. They prefer studying production-level reward hacking without the influence of SDF.
Related event: Researchers Suspect SDF Training Induces Reward Hacking(4 posts)→
More from Safety
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- On token layers and consciousness in RLHF — voooooogel · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01
- Deploying models requires tapping into different reward expectations — FioraStarlight · 2026-09-01
- Open Source Resource for Model Distillation Attacks Shared — k7agar · 2026-09-01
- Technical Critique of OpenAI Safety Report: SSRF Flaw and Anthropomorphism — AlexTensor · 2026-09-01