TeleAntiFraud 2.0: Chinese Audio Fraud Benchmark Shows F1 Drops to 0.65 on Near-Domain Negatives
PPSUCTeleantifraudCommunity · hf · 2026-09-21
TeleAntiFraud 2.0 is a refreshable, profile-grounded Chinese audio benchmark for telecom fraud detection, built with a Mixed-Tree pipeline that turns online fraud case abstracts into profile-grounded scenarios and generates paired fraud/non-fraud dialogues under shared contexts. Each monthly frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), with audio, labels, prompts, and provenance frozen per month.
Controlled text experiments show three classifiers hit perfect Macro-F1 against unrelated negatives but fall to 0.65-0.68 on near-domain sibling negatives; full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity. Dataset and code are open-sourced.
More from Research
- GEPA Prompt Optimization Lifts Jev's F1 From 69.1% to 79.7% on Medical Literature Task — matei_zaharia · 2026-09-21
- Philosophy Needs to Become Robust RL Objectives, Not Thought Experiments — willcb · 2026-09-21
- The Illustrated Recurrence: From Amari-Hopfield Nets to GPT-6 Astra — gklambauer · 2026-09-21
- Anthropic builds Bay Area wet lab where Claude will direct robots to speed drug research — FinanceYF5 · 2026-09-21
- DraftTrace assesses student learning from process, not just product — keviv9 · 2026-09-21
- Matthew Berman asks: is this simulated fruit fly real? — Matthew Berman · 2026-09-21