Qwen's Multimodal Image-Editing Framework Boosts Recommendation CTR by 32%
QwenBusinessUnit · hf · 2026-08-11
The Qwen team addresses follow-up edit recommendations in image-creation conversations with a three-stage multimodal alignment framework.
- Data Foundation: Collected 100,000 real multi-turn image-creation conversations from Qwen App, finding 80.1% are image-dependent.
- Three-Stage Framework: Stage 1 builds a human-reviewed table of editing intents to fine-tune a multimodal policy. Stage 2 uses user click feedback for multi-objective RL optimization. Stage 3 introduces a visual verifier to reduce visual inconsistencies.
- Results: In a live A/B test with millions of users, the framework reduced visual inconsistency from 3.7% to 0.9%, improved recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns by 39.90%.
More from Apps
- Testing Vizard AI: Turning a Single Clip into a Polished Ad Automatically — Aiden_Tech_Ai · 2026-08-11
- Indie hacker's $10k in 100 days challenge: Day 20 revenue $1,966, SEO tool CrawlRaven launches — ayushtweetshere · 2026-08-11
- Rumor: Cursor to Launch 'Origin', a GitHub Alternative, Alongside Grok 4.6 — mark_k · 2026-08-11
- Atomic: Open-Source Tool Turns Markdown Notes into Semantic Knowledge Graph — tom_doerr · 2026-08-11
- Tested 10 Popular ChatGPT Skills: Over Half Are Incompatible, with Workarounds — Nearby_Pair_6483 · 2026-08-11
- Grok Imagine 2.0 Launches with Multi-Reference Fusion and Precise Editing — eyishazyer · 2026-08-11