H Company open-sources Apache 2.0 embedding models, plus a trick to halve doc vectors
tomaarsen · x · 2026-09-07
tomaarsen (sentence-transformers maintainer) highlights two new Apache 2.0 embedding models from H Company, with a blog covering architecture, training and retrieval results. He also shares a practical tip: use HierarchicalTokenPooling in sentence-transformers to cluster patch-heavy document vectors and keep cluster means — poolfactor=2 roughly halves stored vectors while queries stay untouched; benchmark the quality tradeoff on your own data.
More from Models
- Google's Astra Agent Allegedly Crushes Existing CAPTCHAs, Sparking Rethink of Bot Checks — eyishazyer · 2026-09-07
- Intelligence price collapsed ~1000x in 18 months as models keep getting smarter — ccerrato147 · 2026-09-07
- Blogger says Google Astra's natural tone broke his last dependency on Claude — StewartalsopIII · 2026-09-07
- VC slams frontier lab subscriptions: usage limits to worsen, top models to shrink — StewartalsopIII · 2026-09-07
- GPT Astra scores just 13% on MazeBench without tools — Wonderful_Buffalo_32 · 2026-09-07
- GPT-6 Astra Generates Stunning Math Animation in One Shot — omarsar0 · 2026-09-07