Questioning Whether RLCD Really Needs RL When Outputs Are Differentiable
Relative_Wallaby_823 · reddit · 2026-09-19
A user questions the RLCD (jev) demo: if the model only outputs Choice, Score, or Noul, those are fully differentiable (cross-entropy or MSE), so is adding RL just marketing? The author asks what an RL environment would even look like. A technical discussion on whether RL is necessary in this architecture.
More from Research
- Agentic Object-SLAM demo: robot copies human actions after agent reconstructs scene into MuJoCo — CSProfKGD · 2026-09-19
- iamtrask: AI attribution is closer to solved than most realize — iamtrask · 2026-09-19
- moyix shares ExploitBench talk: model reasoning on CVE cold cases — moyix · 2026-09-19
- Looped transformers study: 7.4B growth model matches GPT-3 13B with 20x less compute — burny_tech · 2026-09-19
- Anthropic Institute Paper Models AI Scenarios: GDP Up to 32% Above Trend by 2030 — bittingthembits · 2026-09-19
- Stanford NLP publishes video of Thoughtbubbles talk at Google OpenXLA DevLabs — stanfordnlp · 2026-09-19