Google Cloud previews reinforcement fine-tuning for Gemini with custom reward functions
rseroter · x · 2026-09-21
Google Cloud has previewed reinforcement learning fine-tuning for Gemini models, letting users fine-tune with their own prompts and self-defined reward functions. The model iteratively generates responses, gets scored by the reward function, and updates parameters via RL algorithms. It applies across text, audio, image, and video, and targets boosting reasoning and output quality for custom use cases and agentic workflows. Docs are already live.
More from coding & agent
- Zhejiang U & SJTU unveil DAS, an agent that writes publication-ready surveys in an hour — jiqizhixin · 2026-09-22
- Dev vibe-codes a living-room wall game with Codex, controlled by hand gestures — tristanbob · 2026-09-22
- A developer built an independent quality-and-popularity index for MCP servers — Prestigious-Web-2968 · 2026-09-22
- GitHub Copilot desktop app to get editable diffs for agent-generated changes — mariorod1 · 2026-09-22
- Dev builds AI Scene Director on JEV: one prompt rewrites Three.js scene — jamestagg · 2026-09-22
- Why your agent keeps resurrecting blockers from tasks fixed weeks ago — Lopsided_Manner_7602 · 2026-09-22