FineWeb author: annotating pretraining data with a 27B model is wild but pays off at deployment
antoine_chaffin · x · 2026-10-09
FineWeb co-author antoinechaffin defends using a general-purpose 27B model to annotate and even rerank pretraining data: 'wild' but valuable once shrunk, since zero-shot and multi-task learning deliver capabilities hard to get otherwise. He still concedes that for large-scale repetitive tasks, fine-tuned small models like ModernBERT/mmBERT always win on throughput — distillation is the way.
Related event: Hugging Face researcher on small fine-tuned models vs general models(2 posts)→
More from Infra
- Cloudflare Workers can now sit next to your PlanetScale database region, edge FUD debunked — ritakozlov · 2026-10-09
- Phinity Exits Stealth With $5.2M Seed for Autonomous Chip Design, 8-Figure ARR — JeffDean · 2026-10-09
- Chollet: AI capex is growing super-exponentially while progress is only sub-linear — fchollet · 2026-10-09
- Patching MLX to stage quantized weights to FP8 yields +40% prefill on M6 — Brilliant-Hall1387 · 2026-10-09
- a16z: Agents burn 5x the tokens of humans, up 14x in six months as AWS rewires for machine users — a16z · 2026-10-09
- Microsoft's Surface RTX dev box ditches ConnectX-7 NIC, limiting multi-box clustering — Scobleizer · 2026-10-09