VultronRetriever Model Family Released
madkimchi · reddit · 2026-07-11
The VultronRetriever series of retrieval models has been released on Hugging Face, emphasizing fully offline Q&A and document embedding on iPhones. The series achieved strong scores on MTEB: Prime-8B ranked first in its category and claims the overall #1 spot; Core-4.5B ranked second; Flash-0.8B also runs on edge devices with support for higher throughput.
The author also highlighted several engineering selling points:
- Prime-8B's index storage footprint is up to 16x smaller than the previous generation of 9B-class leading models, with 12x higher throughput;
- Flash-0.8B works fully offline with an indexing speed of 60 images per minute;
- Uses Hydra Architecture to perform late interaction retrieval + generation with lower VRAM;
- Training claims 0% cross-dataset duplication and 0% eval contamination.
More from Models
- Models now make decent PPTX decks, and people are acting like that’s normal — ziv_ravid · 2026-07-21
- A joking post asks whether this was the famous “move 37” moment for math — NielsRogge · 2026-07-21
- Xiaohongshu’s dots-note-3.0 gets a perfect IMO score and becomes the world’s second gold model — 量子位 · 2026-07-21
- Reddit user says Grok 4.5 felt faster and better than Claude for coding workflows — Rare_Iron9142 · 2026-07-21
- Kimi and GLM distillation debate reignites over what counts as real innovation — basedjensen · 2026-07-21
- Meme mocks Google’s AI lead after early Gemini 3.6 Flash outputs look rough — max_paperclips · 2026-07-21