Tutorial: train Qwen3.8-27B via RL with Jev as reward model, no GPU needed, ~$5 total

sophiamyang · x · 2026-09-25

Sophia Yang published fireworks-jev-reward-rl, a tutorial for running RL on Qwen3.8-27B's untrained base on Fireworks with Jev as the reward model, targeting less "AI slop" output.

Result: after 24 RL updates, mean Jev reward rose from 0.583 to 0.759, evidence of less slop by Jev's measure.

Pipeline and cost:

Paid commands require --execute; results are stochastic; a notebook version runs the same commands.

Related event: Tutorial: RL Fine-Tuning Qwen with Jev Reward Model to Cut AI Slop(2 posts)→

Original post →

More from coding & agent

coding & agent channel →