Google Cloud previews reinforcement fine-tuning for Gemini with custom reward functions

rseroter · x · 2026-09-21

Google Cloud has previewed reinforcement learning fine-tuning for Gemini models, letting users fine-tune with their own prompts and self-defined reward functions. The model iteratively generates responses, gets scored by the reward function, and updates parameters via RL algorithms. It applies across text, audio, image, and video, and targets boosting reasoning and output quality for custom use cases and agentic workflows. Docs are already live.

Original post →

More from coding & agent

coding & agent channel →