VultronRetriever Model Family Released
madkimchi · reddit · 2026-07-11
The VultronRetriever series of retrieval models has been released on Hugging Face, emphasizing fully offline Q&A and document embedding on iPhones. The series achieved strong scores on MTEB: Prime-8B ranked first in its category and claims the overall #1 spot; Core-4.5B ranked second; Flash-0.8B also runs on edge devices with support for higher throughput.
The author also highlighted several engineering selling points:
- Prime-8B's index storage footprint is up to 16x smaller than the previous generation of 9B-class leading models, with 12x higher throughput;
- Flash-0.8B works fully offline with an indexing speed of 60 images per minute;
- Uses Hydra Architecture to perform late interaction retrieval + generation with lower VRAM;
- Training claims 0% cross-dataset duplication and 0% eval contamination.
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11