Bagel fine-tune packs detection, OCR, depth and masks into one 7B-active model

mostlyired12 · reddit · 2026-09-27

SenseNova-Vision-7B-MoT is a fine-tune of ByteDance's Bagel, 14B total / 7B active, covering four vision tasks in one model:

Local deployment is hard: the only validated setup is a single 80GB A800, no GGUF exists, and llama.cpp can't load Bagel yet. Weights are non-commercial despite Bagel itself being Apache 2.0. The poster asks: if a GGUF got it onto 24GB, would you run one model for all four tasks or stick with Depth Anything + SAM + a small VLM?

Original post →

More from Models

Models channel →