Lightweight prompt injection detector: MiniLM + logistic regression, F1 just 61.6% on adversarial benchmark

Worldly_Yoghurt8850 · reddit · 2026-09-08

An open-source prompt injection detector using all-MiniLM-L6-v2 embeddings plus logistic regression runs without a GPU. Trained on 1,130 examples, it scores only 53.3% accuracy and 66.4% false positives on a frozen 227-example adversarial benchmark—benign security discussions trip it up. Next steps: hard negatives, contrastive pairs, chunk-aware detection.

Original post →

More from Safety

Safety channel →