Dodging reward hacking: fine-tuning Qwen3.5-4B with Jev to nail alliterations
AAAzzam · x · 2026-09-21
Responding to the common RL problem of reward hacking, the author shows an alternative to hard-coded rules or non-deterministic judges: using Jev to fine-tune Qwen3.5-4B so it produces genuinely good alliterations. The demo was built with Modal and typesafeai, illustrating a lightweight reward-design approach for small-model RL fine-tuning.
More from Research
- Gwern's 2016 'Tool AIs Want to Be Agent AIs' thesis vindicated, says founder — tszzl · 2026-09-21
- ττ-bench: Best Coding-Agent Setup Passes Just 23.9% of Real Client Simulations — rohanpaul_ai · 2026-09-21
- Prof. Grimmer to speak on "Optimization in the Space of Algorithms" at Brown and Cornell — prof_grimmer · 2026-09-21
- JEPA-Anything proposes one world-modeling framework across vision, biology, weather and more — gekobraa · 2026-09-21
- Mathematician littmath on AI slop: posting PDFs is fine, misrepresenting them is misconduct — littmath · 2026-09-21
- One topological feature boosts noisy speech recognition accuracy from 84.9% to 88.4% on TIMIT — bravo_abad · 2026-09-21