From RLHF to RLMF: Making Market Payments the AI Reward Signal
curious_vii · x · 2026-07-17
The author proposes an evolutionary path: once the LLM-as-judge loop is compressed into the product UI, the next step is moving from RLHF (Reinforcement Learning from Human Feedback) to RLMF (Reinforcement Learning from Market Feedback). The core idea is "test-driven development, but the test metric is whether someone is willing to pay you"—using real BTC/market payments as the reward signal for DPO.
More from AGI Musings
- Jamie Dimon says bureaucracy, not AI, is the real system crushing intelligence — r0ck3t23 · 2026-07-21
- OpenAI and Anthropic’s internal models are said to be far stronger than today’s public systems — scaling01 · 2026-07-21
- Superintelligence and robot abundance will force a new social contract — Dr_Singularity · 2026-07-21
- The Guardian examines how AI companionship is turning intimacy into an economy — nordicinst · 2026-07-21
- A frustrated user says modern AI keeps hallucinating on real-world repair tasks — doochenutz · 2026-07-21
- A repost argues that AI will make today’s hard tasks trivial within months — OwariDa · 2026-07-21