Subnet 44 expands around Satori, a 7B vision-language model for grounding
richdotca · x · 2026-08-04
Subnet 44 is expanding around Satori, a new vision-language model
The quoted post says Webuildscore is incentivizing training for Satori, a VLM that aims to do both reasoning and grounding.
- It should not only answer questions about images, but also point to the evidence in pixels.
- The model is described as capable of scene reasoning, object detection and segmentation, text reading, counting entities, and grounding claims.
- The team contrasts Satori with typical VLMs that are either good at reasoning or good at grounding, saying Satori sits at the intersection.
- They plan to start with a 7B base model.
More from Infra
- Exa says its web index has 80B pages and is on track for Google-scale in 2027 — garrytan · 2026-08-04
- OpenCode Go says it processed 6T tokens in a single day, led by DeepSeek models — ycombinator · 2026-08-04
- Cloud and AI vendors are creating expensive lock-in and uncontrolled future risks — DavidLinthicum · 2026-08-04
- Swarms Cloud adds saved workflows, session persistence, and 1,500+ models — KyeGomezB · 2026-08-04
- SK Hynix prepares to break ground on $3.87B Indiana HBM packaging fab — rwang07 · 2026-08-04
- Kimi K3 reportedly runs on an 8GB CPU setup by streaming experts from SSD — porAssass · 2026-08-04