Fireworks shows two Jev training recipes: GRPO-style reward scoring and offline DPO/SFT filtering
sophiamyang · x · 2026-09-22
Fireworks AI's sophiamyang shared two ways to plug Jev into training pipelines:
- Online: sample k rollouts per prompt, have Jev score each one, map its typed answers to a reward, then train with GRPO-style RFT.
- Offline: use Jev to rank DPO/ORPO preference pairs or filter SFT examples.
Details are in the Fireworks docs, which cover inference and fine-tuning (up to 1T+ params) across 100+ open models.
More from coding & agent
- Xiaomi's CodeMidas turns source code into RL environments, doubling DeepSWE to 21.7% — maier_ak · 2026-09-22
- Alibaba's Qwen team launches RecreationBench to test hybrid computer-use agents by app recreation — TianbaoX · 2026-09-22
- Nautilo launches as 100% open-source multi-user agent platform with browser-controlled Genie — Dan_Jeffries1 · 2026-09-22
- Give your investment research agent paid Substack access via x402 — kleffew94 · 2026-09-22
- Hermes Agent + ComfyUI Autonomously Iterates Overnight to Nail a Character-Swap Video — Teknium · 2026-09-22
- Token Saver lets Codex delegate coding work to cheaper Muse 1.3 to cut token burn — AIandDesign · 2026-09-22