What's left for hobbyists to post-train on in 2026? One bets on self-play poker
No-Compote-6794 · reddit · 2026-08-21
A LocalLLaMA discussion asks what problems remain for hobbyist RL/agent post-training now that most tasks are saturated or need expert data. The author proposes poker as a candidate: using P&L as the reward signal, similar to running a quant fund, with an edge that never saturates via self-play — and plans to build it from scratch on his Tiny-Qwen open-source project.
More from Models
- Zhipu's SAO: single-rollout async RL trains stably for 1,000 steps, beats GRPO — teortaxesTex · 2026-08-21
- Gemini's Safety Filters Too Strict? Rejects Kissing and Roadside Photos — Dry-Sympathy-3182 · 2026-08-21
- 0.63M Parameter Verifier Matches 7B Models in Specific Tasks — jm_alexia · 2026-08-21
- Pangram v4 Model Claims to Remove AI Watermarks and Mimic Human Writing — Scobleizer · 2026-08-21
- Brundage: something funky going on with ChatGPT inference, likely testing — Miles_Brundage · 2026-08-21
- Does Claude perform better in 'claudish'? Researchers call for empirical measures — voooooogel · 2026-08-21