Adapting ViT Representations to Specific Tasks
y_m_asano · x · 2026-07-19
This post shares a research diagram arguing that **pre-trained ViT representations should not be entirely static**. The author suggests that while a standard model's output won't change no matter how many times it runs a forward pass, allowing the model to "know what to look for" can adjust its representation to better suit the attributes of the current task. The diagram contrasts standard DINOv2 attention with SteerViT across different attributes: for the same input, SteerViT can direct its attention to task-relevant areas like a "person" or a "potted plant," rather than fixating on a single salient object like a static representation would.
More from Research
- FloC 2026 AIMACS workshop on AI for math and CS set for July 25 — swarat · 2026-07-21
- Knowledgeless Language Models cut closed-book recall by anonymizing entities during pretraining — gdm3000 · 2026-07-21
- CPU-native LLM pilot passes 4 of 5 gates, but cross-tokenizer distillation still loses — WildPino25 · 2026-07-21
- A GPT 5.6 Sol workflow reportedly generates an infinite family of counterexamples — OwariDa · 2026-07-21
- A research guide v7 surfaces two contradictions instead of smoothing them over — Fantastic_Aside6599 · 2026-07-21
- Agents can remember facts, but still forget how to do the job — No_Advertising2536 · 2026-07-21