Oxygen-TryOn is a fashion-native model for multi-item virtual try-on
JD-company · hf · 2026-07-28
JD’s team presents Oxygen-TryOn, a fashion-native foundation model for any-item virtual try-on.
- Unlike general-purpose image editors repurposed with prompts, the model is built specifically for try-on.
- It accepts one or more reference items and a target person image, then synthesizes a photorealistic image of that person wearing the items.
- The system supports a wide range of fashion categories, including clothing, outerwear, accessories, footwear, and bags.
- It handles full-body and half-body views, variable numbers of references, and free multi-item composition.
- The training stack includes a large-scale data engine plus a three-stage recipe: continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL).
- The RL stage uses a hybrid reward combining an in-house try-on reward model and a rubric-guided general-purpose model.
- The authors claim state-of-the-art consistency and realism on public benchmarks and their internal benchmark, including strong results on multi-item try-on versus systems such as Nano Banana Pro, GPT-Image-2, Seedream5 Lite, and FLUX.2.
More from Multimodal
- User shares a new Midjourney style with exact prompt settings — azed_ai · 2026-07-28
- A 5 MB McBess-style LoRA for Krea2 trained on 120 captioned images — Winter_unmuted · 2026-07-28
- Topview launches Film Studio with 3D blocking and micro-expression controls — XFreeze · 2026-07-28
- GaussianGPT uses autoregressive next-token prediction to generate 3D Gaussian scenes — rsasaki0109 · 2026-07-28
- Testing Style LoRAs: How to Isolate Style from Content in Image Generation — Dark_Sytze · 2026-07-28
- FilmBench evaluates cinematic video generation with film-school shot lists and 35 metrics — Shengyi Wang · 2026-07-28