What's left for hobbyists to post-train on in 2026? One bets on self-play poker

No-Compote-6794 · reddit · 2026-08-21

A LocalLLaMA discussion asks what problems remain for hobbyist RL/agent post-training now that most tasks are saturated or need expert data. The author proposes poker as a candidate: using P&L as the reward signal, similar to running a quant fund, with an edge that never saturates via self-play — and plans to build it from scratch on his Tiny-Qwen open-source project.

Original post →

More from Models

Models channel →