Fudan's IDSpect Decomposes Chinese Characters into Radicals for Fine-Grained Text-to-Image RL Rewards
Fudan-University · hf · 2026-10-01
- Rendering accurate Chinese text remains hard: OCR-based RL rewards treat each character atomically, so visually different radical-level errors get equally coarse feedback, rewarding glyphs that merely resemble targets.
- Fudan's IDSpect uses Ideographic Description Sequences (spatial operators + components): a trained IDS recognizer transcribes rendered text into IDS tokens, and rewards align crop-level visual IDS predictions with a deterministically decomposed target sequence, plus a whole-character semantic reward.
- It requires no generator changes and adds no inference cost. Applied to GRPO post-training of Qwen-Image, IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.
More from Multimodal
- Seedance 2.5 demo brings anime-level dual-sword choreography into photorealistic cinema — SimplyAnnisa · 2026-10-01
- Midjourney --sref 3896456162 recreates 1970s Kodak film & disco aesthetics — michaelrabone · 2026-10-01
- A finished LoRA run doesn't mean it learned the style: a reproducible SDXL validation workflow — no3us · 2026-10-01
- Editor open-sources open-fusion-mcp: Claude builds editable motion graphics inside DaVinci Resolve — JohnnyLegion · 2026-10-01
- AI-generated clip: Raven Noire writing in her diary while listening to goth music — Street-Pound5762 · 2026-10-01
- Failed MiniMax H3 VAE detail experiment yields a useful 2X detail VAE and workflow — NoMouse9610 · 2026-10-01