Ant open-sources SingProbe: a 3–5M probe reading risk from hidden states for real-time streaming guardrails
aigclink · x · 2026-09-15
Ant Group released and open-sourced SingProbe, an endogenous runtime safety guardrail that departs from the usual external-audit-model approach (which is costly, slow, and weak on complex context).
Key points:
- Reads risk signals directly from the base model's hidden states, training a tiny 3–5M probe per base model
- Per-token scores for user intent risk, unsafe generated content, and hallucination risk
- Depends only on the current prefix, enabling streaming real-time alerts, truncation, retry, or strategy switching
Deployed in the decode stage of Ling-3.0-flash in production with minimal overhead. Training code is on GitHub (inclusionAI/SingProbe): a frozen base model served via a token-probe-patched SGLang build dumps per-layer hidden states streamed into probe training, outputting 10-dim logits per token, with an accompanying technical report.
Related event: Ant Group Open-Sources SingProbe, an Intrinsic AI Guardrail(2 posts)→
More from Safety
- Hijacked HBO Max Reddit account pushed 108 ClickFix ads in 48 hours — TechNadu · 2026-09-15
- Adolescence writer Jack Thorne calls AI-scripted work fraud, wants UK prosecution — LinkedInNews · 2026-09-15
- TIME's 'AI Tipping Point': a scary-capable summer and an industry losing control — timemagazine · 2026-09-15
- Robotaxis arrive in London via Uber and Wayve; UK think tank proposes AV levy worth up to £46bn — dc_lawrence · 2026-09-15
- Europe holds just 5% of global AI compute vs US 75%, warn 60+ experts in new strategy report — anselm · 2026-09-15
- ChinaTalk ModelTalk: Why the Hugging Face hack broke through and AI safety's new coalition — ChinaTalk · 2026-09-15