Ant Group open-sources SingProbe, an in-model safety guardrail adding <0.5% decode overhead
智东西 · wechat · 2026-09-16
Ant Group's AI safety lab open-sourced SingProbe, an endogenous safety guardrail that runs alongside model generation, releasing code, models, and the SingStreamBench streaming-safety benchmark.
- Runtime probes reuse internal signals from inference to score user intent, answer safety, and hallucination risk in real time, enabling alerts or abortion before generation completes
- Adapted to 29 mainstream open models (Ling-3.0, GLM-5.2/5.3, Qwen, DeepSeek V4) and integrated with SGLang and vLLM
- Measured decode overhead below 0.5%; the flash variant beats public baselines on safety classification and streaming detection
- SingStreamBench measures when risks appear and how quickly they're caught
- A medical variant corrected 25.03% of baseline errors on AntAngelMed-100B via on-demand intervention
Related event: Ant Group Open-Sources SingProbe, an Intrinsic LLM Safety Guardrail(3 posts)→
More from Safety
- US Federal Register Search Found Using Distilled Qwen Models — kimmonismus · 2026-09-16
- AI cooperation is hard because it's a general-purpose technology: think standards, not arms control — mchorowitz · 2026-09-16
- Activist Heads to DC to Back Superintelligence Ban, Invoking Teller's Faulty Nuke Safety Math — erikphoel · 2026-09-16
- Economist Proposes Time-Limited "Learning Rights" as Alternative to Nadella's Trust Boundaries — carlbfrey · 2026-09-16
- AI Safety Researcher Argues We Lack Any Capability to Pause Frontier AI If Needed — davidmanheim · 2026-09-16
- Debate: as hacking-capable AI agents emerge, regulation must assume mishandling — binarybits · 2026-09-16