Fine-tuned 350M model lifts PII removal from 89.7% to 99.7%
JosephJacks_ · x · 2026-10-05
Developer Gareth Beall benchmarked 7 PII-detection models and picked LiquidAI's 350M LFM2.5-Encoder-350M-PII-Detector, then fine-tuned it.
- Complete identifier removal rose from 89.7% to 99.7% across 5,000 synthetic and real test cases, with zero marked clinical facts lost.
- Median response time is 20ms on DGX Spark; it now runs live in his clinical production system on RTX 6000s.
- He says he built it solo in a day and credits LiquidAI for releasing capable small-model weights.
The model card describes a bidirectional masked-LM encoder for token classification, multilingual (EN/DE/FR/ES/ZH/JA/KO and more), with weights publicly available.
More from coding & agent
- scikit-learn creator on agentic data science: what to delegate to agents and how to verify — hugobowne · 2026-10-05
- NVIDIA open-sources 391 signed agent skills for Claude Code, Codex and Cursor under Apache 2.0 — Arindam_1729 · 2026-10-05
- QuixiAI launches With, a systems language with Rust-level safety and zero-glue C interop — QuixiAI · 2026-10-05
- The With Language Safely Wraps sqlite3, Claiming Rust and Zig Can't — QuixiAI · 2026-10-05
- SuperX MCP in Claude Code grew an X account from 64 to 1,000+ followers since July — tibo_maker · 2026-10-05
- Vanishing Gradients podcast: agents are like coffee — they help you make stupid mistakes faster — hugobowne · 2026-10-05