Flash Attention Causes Performance Drops in BGE-M3 and Other Models
Incompatibility between packed position IDs and RoBERTa when using flashattention2 caused significant performance drops, notably degrading BAAI/bge-m3's STSB test Spearman score from 0.8485 to 0.7239 before the fix.
2026-07-23 ~ 2026-07-23 · 2 related posts
- Episode 1: Flash Attention Causes Performance Drops in BGE-M3 and Other Models(2026-07-23, 2 posts)
- Episode 2: Sentence Transformers 5.6.1 Fixes Flash Attention Embedding Regression(2026-07-23, 3 posts)
- Sentence Transformers packed position IDs wrong for RoBERTa models under flash attention — tomaarsen · 2026-07-23
- BAAI/bge-m3 lost accuracy under flash attention until the Sentence Transformers patch — tomaarsen · 2026-07-23