Earlier vision injection wins: why all future LLMs will be VLAs
kamalgupta09 · x · 2026-09-07
Arguing against a definition that ties "VLA" to VLM-pretrained bases, the author notes:
- With a fixed vision+text token budget, injecting vision tokens earlier in pretraining improves the final model on both vision and text tasks (shown in K2.5 and other work).
- Nearly all LLMs today already use vision early in pretraining, so defining a VLM by "LLM-pretrained base" would be silly.
- By the same logic, future LLMs will all include action data from early pretraining — meaning all future LLMs will be VLAs.
More from Models
- Ask LLMs to pick a random number 1-30: most all answer 17 — ohnag_eryeah · 2026-09-07
- Polymarket puts 82% odds on Anthropic releasing next Claude Opus this month — Polymarket · 2026-09-07
- Early test: GPT-6 Astra beats GPT-5.6 Sol while using fewer tokens — haider1 · 2026-09-07
- Viral screenshot shows Astra 6 at xhigh responding with 'agi' — viksit · 2026-09-07
- OpenAI Employee Predicts Major Blender x Astra Gains by Mid-2027 — Distinct-Question-16 · 2026-09-07
- GPT-6 Astra builds exploded view of 71,492-atom GPCR simulation in membrane — DeryaTR_ · 2026-09-07