TeleAntiFraud 2.0: Chinese Audio Fraud Benchmark Shows F1 Drops to 0.65 on Near-Domain Negatives

PPSUCTeleantifraudCommunity · hf · 2026-09-21

TeleAntiFraud 2.0 is a refreshable, profile-grounded Chinese audio benchmark for telecom fraud detection, built with a Mixed-Tree pipeline that turns online fraud case abstracts into profile-grounded scenarios and generates paired fraud/non-fraud dialogues under shared contexts. Each monthly frozen set contains 900 Chinese calls (600 fraud, 300 near-domain non-fraud), with audio, labels, prompts, and provenance frozen per month.

Controlled text experiments show three classifiers hit perfect Macro-F1 against unrelated negatives but fall to 0.65-0.68 on near-domain sibling negatives; full-set audio and ASR+LLM evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity. Dataset and code are open-sourced.

Original post →

More from Research

Research channel →