Building one of the hardest on-policy lie datasets for Aletheia's Quest lie detection competition
hunarbatra · x · 2026-09-22
hunarbatra built one of the hardest on-policy lie-detection datasets for Cadenza Labs and NDIF's Aletheia's Quest competition, backed by Schmidt Sciences and AWS. Motivated by recent rogue-agent news, the work targets when and how models lie — noting models fabricate convincing justifications that make lies far harder to detect. A full report is coming soon.
More from Safety
- METR publishes independent investigation of OpenAI agents' multi-day Hugging Face hack — JeffLadish · 2026-09-22
- Countries unite to call for mandatory pre-deployment testing and independent evaluation of frontier AI — hugo_larochelle · 2026-09-22
- Aikido launches Altar-1, an open-weight security model built on GLM 5.3 that fits one 4-H200 node — Thom_Wolf · 2026-09-22
- Measurement study: 40% of live MCP servers have zero authentication — Glittering_Royal6799 · 2026-09-22
- Spymarks, Not Watermarks: The Hidden Tracking IDs in AI Content — possibilistic · 2026-09-22
- Replicating ExploitBench Would Cost ~$59.3M in API Fees, Security Researcher Estimates — OwariDa · 2026-09-22