Dodging reward hacking: fine-tuning Qwen3.5-4B with Jev to nail alliterations

AAAzzam · x · 2026-09-21

Responding to the common RL problem of reward hacking, the author shows an alternative to hard-coded rules or non-deterministic judges: using Jev to fine-tune Qwen3.5-4B so it produces genuinely good alliterations. The demo was built with Modal and typesafeai, illustrating a lightweight reward-design approach for small-model RL fine-tuning.

Original post →

More from Research

Research channel →