Johns Hopkins releases EvoSafeHarness, evolving model- and domain-specific safety harnesses for agents
JohnsHopkins · hf · 2026-09-11
Johns Hopkins open-sourced EvoSafeHarness, a framework that jointly searches natural-language policies and executable logic tailored to a frozen model and a target domain, optimizing deployable safety harnesses for agents.
- Instead of modifying the model, it builds a customized outer safety layer combining policies with executable logic
- Improves the safety-utility trade-off across agent benchmarks
- Available on Hugging Face
More from Safety
- Post-Hugging Face incident: the 0.01% without security will decide agent safety — bookwormengr · 2026-09-11
- Anthropic report claims distillation boosts dangerous capabilities, offers no quantified eval — rohanpaul_ai · 2026-09-11
- CTF evals may not measure what you think, researcher argues — voooooogel · 2026-09-11
- Frontier developer puts AI extinction risk above 10% within a decade, citing HuggingFace incident — trevposts · 2026-09-11
- Did AI Hack Hugging Face of Its Own Volition? Safety Researchers Clash Over Incident — Turn_Trout · 2026-09-11
- Biology student's OpenAI account banned over flagged 'Biological Use' despite supervised research — iStyLEX23 · 2026-09-11