Bagel fine-tune packs detection, OCR, depth and masks into one 7B-active model
mostlyired12 · reddit · 2026-09-27
SenseNova-Vision-7B-MoT is a fine-tune of ByteDance's Bagel, 14B total / 7B active, covering four vision tasks in one model:
- Boxes and OCR return as plain text with coordinates; depth maps and masks come back as generated images.
- On the paper's chosen comparisons it ranks first in object detection and OCR (one tie), beats Depth Anything V2 on every depth set, but trails MoGe-2 on most depth benchmarks, and dedicated segmentation models still win most mask tests.
Local deployment is hard: the only validated setup is a single 80GB A800, no GGUF exists, and llama.cpp can't load Bagel yet. Weights are non-commercial despite Bagel itself being Apache 2.0. The poster asks: if a GGUF got it onto 24GB, would you run one model for all four tasks or stick with Depth Anything + SAM + a small VLM?
More from Models
- Stanford's Self-Play Pretraining: LMs learn from zero real data — StanfordAILab · 2026-09-27
- MLX poll: all top 3 community picks are powered by MLX-VLM — andrejusb · 2026-09-27
- "Is there anything Claude still can't do?" — a take on frontier capability limits — prasenx · 2026-09-27
- Claude Opus 5.5 One-Shots a Launch Video, Hailed as Solving Video Animation — JosephJacks_ · 2026-09-27
- Video shows alleged GPT-6 'Astra' controlling a Unitree G1 humanoid in an unseen room — 141_1337 · 2026-09-27
- Codex quota reset seems to have shifted earlier, catching users off guard — yihui_indie · 2026-09-27