Google open-sources TIPSv2 image-text encoders, SOTA on all four zero-shot segmentation benchmarks

bdsqlsz · x · 2026-08-21

Google released TIPSv2 in four sizes, each with a general image-text encoder plus a DPT variant for dense vision tasks, under Apache 2.0. It sets SOTA on all four reported zero-shot segmentation benchmarks and top-two results on 5/7 image-text and 7/9 image-only evaluations. At ViT-L scale it beats DINOv3 on 4/6 shared tasks despite its teacher having 6× more parameters and 15× more training images. Technically, iBOT++ lifts ADE150 zero-shot segmentation by 14.1 mIoU, and head-only EMA cuts training parameters by 42%.

Original post →

More from Research

Research channel →