FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
FangZhi Zhong, Xuerui Qiu, Yuqi Pan, Ya Liu, Shaowei Gu, Bo Xu, Guoqi Li
cs.CV, cs.AI
2026-09-29
FocusVTC reads long documents as low-DPI pages and zooms only relevant boxes to 144 DPI. At 72 DPI it scores 87.4 on RULER v1 at 2.9× compression, versus Glyph's 57.5 at 3.0×.
Long-context inference pays for attention, latency, and KV cache as the window grows. Visual text compression (VTC) renders documents as page images so a vision-language model can cover the same span with fewer visual tokens. Glyph gets there with dense rendering plus continual pretraining. DeepSeek-OCR studies reconstruction from optical compression.
Fixed resolution ties local legibility to the whole-page token bill. Low DPI saves tokens and blurs characters that matter, including UUIDs and confusable pairs. High DPI makes every paragraph expensive, including the ones the question never touches. Most questions depend on a small subset of the document. Global coverage and precise reading do not need the same resolution.
FocusVTC is from the Institute of Automation, CAS, with UCAS, Shanghai Jiao Tong University, and Zhongguancun Academy. The backbone is Qwen3.5-9B. At inference the model sees low-DPI pages first, then either answers or calls Enhance Region(page, bbox). The tool crops from an aligned 144-DPI page using original-page coordinates in [0, 1000], so the blurry overview and the sharp crop stay registered. No separate continual-pretraining stage.
Font choice is calibrated to the backbone's OCR behavior. Across 15 fonts and eight point sizes on random text, DejaVu Sans and Verdana first hit CER ≤ 5% at 9 pt. DejaVu Sans costs 8,368.4 visual tokens per 32K-token context versus 8,417.8 for Verdana, and it is 0.30 percentage points better on confusable units such as cl, rn, and i. Switching the font lifts LongBench from 54.93 to 56.40. Default rendering is 9 pt with 1 pt extra line spacing. A processed 72-DPI A4 page is 608×832 pixels and 494 visual tokens.
The training set is 29,411 REL-CoT traces that pin each reasoning step to a page index and a box. Sources are mainly ChatQA and ChatQA2, plus TriviaQA, HotpotQA, 2WikiMultihopQA, MuSiQue, and FinQA. DeepSeek V4 Flash drops questions answerable from common knowledge. Gemini 3.5 Flash writes boxed traces. GPT-5 mini checks that the high-DPI crops support the reference answer.
Two training stages follow:
Giving the crop tool to a model that has not learned the policy hurts. After SFT, enabling tools drops LongBench from 37.86 to 25.41. Qwen3.5-9B+tools falls on every aggregate. Skipping SFT and running GRPO alone reaches only 49.30 on LongBench; adding SFT gets 56.40. VTCBench Reasoning is the one exception: 38.15 without SFT versus 36.25 for the full model.
Evaluation uses 72-DPI pages and 144-DPI crops, DejaVu Sans.
| Method | Metric | Result |
| FocusVTC | RULER v1 at 72 DPI | 87.38, 2.9× compression (2,154 prompt + 723 observation tokens vs 8,400 text) |
| Glyph | RULER v1 at 72 DPI | 57.53 at 3.0× |
| Qwen3.5-9B vision, no crops | RULER v1 | 37.40 |
| FocusVTC | LongBench | 56.40 |
| Qwen3.5-9B text | LongBench | 55.86 |
| Glyph | LongBench | 52.34 |
Against the same backbone at fixed 72 DPI, LongBench rises 20.54 points, mostly on single-doc QA, multi-doc QA, and synthetic retrieval. MRCR's macro-average over six length bins and two/four/eight needles moves from 31.65 to 45.56. Per-needle averages are 60.76 / 45.21 / 30.71 versus Glyph at 42.63 / 25.97 / 16.57. In the longest bin, a 197,909-token text reference plus observations compresses about 3.3× / 3.1× / 3.0×. On 300 paired 64K-128K four-needle items, two A100 80GB GPUs, tensor parallel 2, greedy decoding, online end-to-end latency falls from 187.09 s to 67.06 s (2.79×), excluding offline rendering. Generated tokens drop from 5,411.30 to 1,529.77.
VTCBench macro-average is 51.19: Retrieval 91.00, Reasoning 36.25, Memory 26.33. Retrieval and memory lead in the 32K and 64K bins. Reasoning still trails the same backbone's fixed-vision 41.57, beating it only in the two longest bins. Summarization barely moves. Distributed evidence does not fit a crop. Few-shot scores are mixed because demonstrations often span several regions.
A DPI sweep on RULER shows FocusVTC is less sensitive to the starting resolution. 72 DPI minimizes prompt-plus-observation length. Lower DPI inflates observation tokens. Starting at 144 DPI costs about 2.8× the input tokens for modest score gains. Raising the whole page spends capacity on irrelevant text.
General multimodal scores hold: MMMU 65.12 to 66.73, MME 2424.02 to 2457.62, OCRBench 851 to 860, DocVQA 92.38 to 92.43, ChartQA 85.96 to 86.28, InfoVQA 74.76 to 76.20.
For teams running a 9B VLM on long documents, this is a post-training path: 29.4K localization traces, SFT then GRPO, no Glyph-style continual pretraining. At similar compression, RULER v1 jumps from 57.5 to 87.4, and online long-context latency falls. Code is public.
The mechanism matches tasks where the answer sits in a few passages: multi-hop QA, needle retrieval, local grounding. Summarization, document-wide reasoning, and VTCBench-style association should not be expected to move the same way. Font and CER thresholds are tuned to Qwen3.5's OCR. A different backbone needs a rerun.
The appendix is explicit: a 9B backbone and 29.4K REL-CoT examples. Scaling the model, languages, layouts, and reasoning traces is untested. Full interaction cost is not jointly scored with answer quality. The 2.79× latency number subtracts page rendering. Live rendering would shrink that gap.
The experimental table never includes AGAR, SEER, or DeepEyes from the related-work section. Comparisons stay with fixed-resolution VTC and the text backbone. The data pipeline depends on Gemini 3.5 Flash, GPT-5 mini, and DeepSeek V4 Flash, so box quality tracks those teachers. Giving the crop tool to an untrained policy lowers scores. Eight-needle MRCR at 128K-256K is still 15.69.