FocusVTC: adaptive-resolution visual text compression hits 87.4 on RULER at 2.9x compression
FangZhi Zhong · hf · 2026-09-30
FocusVTC breaks the compression-quality trade-off in visual text compression (rendering text as images to save tokens) via adaptive resolution.
- Low DPI saves tokens but hurts legibility; high DPI wastes tokens. FocusVTC combines a low-DPI global view with selective region enhancement woven into reasoning.
- Uses 29.4K REL-CoT examples linking reasoning traces to page indices and bounding boxes; REL-SFT teaches localization, GRPO learns when to enhance resolution.
- At 72 DPI scores 87.4 on RULER v1 at 2.9x compression vs Glyph's 57.5 at 3.0x; beats its text backbone on LongBench (56.40 vs 55.86), +13.91 MRCR macro, 2.79x end-to-end speedup.
- General multimodal ability preserved: MMMU 65.12→66.73, MME 2424→2458.
More from Research
- Microsoft's ML System Forecasts Grid Risks from Space Weather for 66,935 US Substations — Microsoft Research · 2026-10-01
- Google Figures Out How to Watermark AI-Designed Proteins for Biosecurity — Ars Technica AI · 2026-09-30
- Amazon Bedrock's Capacity-Pooling Paper Hits SOSP'26 — tianyin_xu · 2026-09-30
- Method: 25 essays run through strict-correction vs rewriting prompts — paulnovosad · 2026-09-30
- Simons Institute working group reaches consensus on how TCS academia should adapt to AI — tallinzen · 2026-09-30
- A 'wisdom gradient' for AI alignment: eliciting diverse human values bottom-up — edelwax · 2026-09-30