Google open-sources TIPSv2 image-text encoders, SOTA on all four zero-shot segmentation benchmarks
bdsqlsz · x · 2026-08-21
Google released TIPSv2 in four sizes, each with a general image-text encoder plus a DPT variant for dense vision tasks, under Apache 2.0. It sets SOTA on all four reported zero-shot segmentation benchmarks and top-two results on 5/7 image-text and 7/9 image-only evaluations. At ViT-L scale it beats DINOv3 on 4/6 shared tasks despite its teacher having 6× more parameters and 15× more training images. Technically, iBOT++ lifts ADE150 zero-shot segmentation by 14.1 mIoU, and head-only EMA cuts training parameters by 42%.
More from Research
- Two NLP papers accepted to EMNLP 2026: Cross-lingual alignment and bias evaluation — Bollegala · 2026-08-21
- Claude designed de novo protein binders for 14 of 15 targets; now a hosted product — nc_frey · 2026-08-21
- LiteMol-1: Cost-Effective Foundation Model for AI Agents in Drug Design — anshulkundaje · 2026-08-21
- New Open-Source Package Makeshift Integrates NMR Data for Protein Dynamics Models — anshulkundaje · 2026-08-21
- LLM evolves generator programs: Using AI for code gen and reasoning in PCG — Amidos2006 · 2026-08-21
- AgentRadio Research Enables Mid-Task Communication, Boosting Long-Horizon Task Resolution — import_jmr · 2026-08-21