Hands-on with Tencent's Hy Image 3.5: strong text rendering, editing and 5-image references
HeyAmit_ · x · 2026-09-23
- References: supports both text-to-image and image-to-image, and accepts up to 5 reference images, giving the model far more visual context than a single prompt.
- Text rendering: generates text across languages, font sizes and layouts while understanding how typography fits the visual — useful for posters, ads, decks and UI concepts.
- Realism: the model pays attention to lighting, materials, camera language, environmental details and textures, making images noticeably more believable.
- Editing: for scene replacement, restyling and transformations it preserves key features of the reference while applying only the requested changes.
- Workflows: the author's loop is prompt/reference → generate → review, extending to product→ad creative, idea→poster, game concept→UI and more.
More from Multimodal
- Opus 5.5 one-shots a full newsletter launch video, creator says "it's over for video guys" — alex_verem · 2026-09-23
- PixVerse R2 hands-on: AI video that becomes a world you walk through with WASD — HeyAmit_ · 2026-09-23
- One-Person AI Music Video Took a Month: 20 Stills and 10 Takes Per Clip — kraussian · 2026-09-23
- MiniMax H3 Ref2VA Freezes at Model Initializing With an 8th Reference Image — itchplease · 2026-09-23
- H3 long-video degradation workaround: a second noise-injection refine stage in latent space — xyzdist · 2026-09-23
- Dev builds a character-swap LoRA dataset end-to-end with Codex and GPT Image 2.5 — ostrisai · 2026-09-23