Hugging Face Transformers 5.19.0 ships multiple breaking changes; run your evals before upgrading
lmoroney · x · 2026-10-07
Hugging Face Transformers v5.19.0 landed Tuesday with several changes marked as breaking—read the release notes before bumping your pin.
Key changes:
- Adds EmbeddingGemma 2 support;
- Every MoE model whose router computes logits now returns them when outputrouterlogits=True;
- The paged| prefix for SDPA and flash attention is deprecated, as regular sdpa and flashattention2 now handle continuous batching directly;
- OWLv2 image-guided queries now pick the box with the highest objectness score, so detections can shift;
- Expert parallelism gets a token-dispatch mode, now the default for Qwen3 MoE and Mellum, removing the requirement that EP size match TP size.
The author's advice: run your eval set on 5.19 in a fresh environment and diff outputs against your current version before upgrading.
More from coding & agent
- LangChain engineer built an ACP coding agent that replaced Claude Code for 9 months — Hacubu · 2026-10-08
- 30 Real Business Workflow Tests: Keep Agent Evaluation Simple — VibeMarketer_ · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Long-running agents suffer 'constraint amplification': a subtle form of context rot — generativist · 2026-10-08
- Nautilo Ships Text+Vision Model Split, Preps Price/Security-Based Model Routing Gateway — Dan_Jeffries1 · 2026-10-08
- A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy — julsimon · 2026-10-08