Tutorial: train Qwen3.8-27B via RL with Jev as reward model, no GPU needed, ~$5 total
sophiamyang · x · 2026-09-25
Sophia Yang published fireworks-jev-reward-rl, a tutorial for running RL on Qwen3.8-27B's untrained base on Fireworks with Jev as the reward model, targeting less "AI slop" output.
Result: after 24 RL updates, mean Jev reward rose from 0.583 to 0.759, evidence of less slop by Jev's measure.
Pipeline and cost:
- No GPU needed — just a Fireworks API key (serverless training enabled) and a TypeSafe Jev key
- Smoke test: 24 drafts, 56 Jev calls, $0.10 + under $0.01
- Training: 864 drafts, 896 Jev calls, 24 updates in 72 minutes, $3–5 + $0.06 Jev
- Post-training: audit, blind review, report; optionally promote and download the adapter
Paid commands require --execute; results are stochastic; a notebook version runs the same commands.
Related event: Tutorial: RL Fine-Tuning Qwen with Jev Reward Model to Cut AI Slop(2 posts)→
More from coding & agent
- Never hand agents raw API keys: 75-second video on identity, least privilege, and key gateways — tobowers · 2026-09-25
- LangChain webinar: fast, cheap Jev models make real-time filtering and guardrails viable — LangChain · 2026-09-25
- Peter Yang launches 25-lesson AI workflow course with 40+ copy-paste prompts — petergyang · 2026-09-25
- Claude Code's most-hated flaw: message queuing, 244 issues and still broken after months; Codex does it right — sytelus · 2026-09-25
- Dev runs parallel subagents to build 3D games; Claude writes its own GPU/memory task allocator — BLUECOW009 · 2026-09-25
- Cognition AI says it has surpassed $1 billion in annualized revenue run rate — Polymarket · 2026-09-25