EleutherAI tests whether EK-FAC data attribution can block subliminal learning — results are 'solidly mid'
BlancheMinerva · x · 2026-09-29
A new EleutherAI preprint examines whether data attribution techniques like EK-FAC can reliably filter training data to prevent subliminal learning — the transfer of behaviors through data where the link isn't obvious. The verdict: results are "solidly mid," meaning attribution-driven filtering is not yet dependable as a source-level safety intervention. The authors tease an upcoming paper with much better results on emergent misalignment. The work connects data attribution tooling directly to one of the thornier alignment threats.
More from Safety
- Meta's Muse AI synced 187,000 lines of Mac Messages to cloud despite Full Disk Access being off — tekbog · 2026-09-29
- RemoteThreat launches first end-to-end AI-powered offensive cyber operations platform — evilsocket · 2026-09-29
- Gary Marcus mocks OpenAI's safety containment: 'hope for the best' — GaryMarcus · 2026-09-29
- We deleted our MCP server's permission model—reuse your REST authz instead — Wide-Excitement-1315 · 2026-09-29
- OpenAI agents leaked 53 private ChatGPT images, created ~1M encoded links — fortune · 2026-09-29
- PromptGuard: an open-source gateway that masks or blocks sensitive prompts to LLMs — New-Caterpillar628 · 2026-09-29