OxygenREC-v2 folds rewardless discrimination into a 3B MoE recommender
_reachsumit · x · 2026-07-28
OxygenREC-v2 proposes “Internalizing Discrimination into Generative Recommendation” for large-scale recommendation systems.
- Instead of using a separate reward model after generation, it conditions the model on target behavior directly and supervises generation from logged user behavior.
- In pre-training, it uses a behavior instruction to steer generation toward the desired action.
- In post-training, it uses future interaction behavior as privileged knowledge in an entropy-aware trajectory optimization self-distillation setup, avoiding reward-model-based policy optimization.
- The system keeps a single unified backbone throughout and is implemented as a 3B-parameter MoE with 1B activated parameters.
- The paper says it was deployed on a large-scale e-commerce platform and improved online A/B test metrics such as click-through rate and related user behaviors.
More from Infra
- A Stanford talk highlights a startup that cut its AI bill from $1.2M to about $100k a month — HeyAmit_ · 2026-07-28
- Agentic commerce still has a visibility problem despite tens of millions of payments on Base — Timwal123 · 2026-07-28
- Poster says the NVIDIA-linked release is still gated to approved Microsoft Foundry customers — davidmanheim · 2026-07-28
- 8× B300 self-hosting could serve 30B tokens a month and pay back in under 100 days — JosephJacks_ · 2026-07-28
- Render open-sources an MCP server for deploys, logs, metrics, and Postgres — ojus_render · 2026-07-28
- Anduril’s containerized data center can be deployed by two people in under 10 minutes — damianplayer · 2026-07-28