Arena unveils post-training method combining preference and rubric rewards

Arena detailed a new post-training method for text-to-image models that combines a reward model trained on 5 million human preference votes with VLM-generated rubric rewards, addressing reward hacking. Applied to FLUX.2-dev, it gained 69 Elo.

2026-10-02 ~ 2026-10-02 · 2 related posts