Baseten Launches Loops Fine-Tuning SDK, Slashing DeepSeek Costs by 98%
baseten · x · 2026-08-06
Baseten has introduced Loops, a training SDK designed for fine-tuning and post-training large language models at long sequence lengths, currently available in early access.
- Core Features: It supports fine-tuning with LoRA and running asynchronous reinforcement learning (RL). The trainer and sampler scale independently, preventing RL rollouts from competing for training compute.
- Model Support: The company claims that fine-tuning DeepSeek V4 Flash using SFT, DPO, PPO, or GRPO allows specialized models to outperform frontier models on specific tasks at an 80-98% lower cost.
- Flexible Deployment: Generated checkpoints can be downloaded directly or deployed to the Baseten Inference Stack via UI, CLI, or API.
More from coding & agent
- Meta Launches Muse Code: Parallel Agents for Real-Time Game Generation — qinzytech · 2026-08-06
- Does AI Summarizing Execution Experience Count as Self-Improvement? — ___Patrice___ · 2026-08-06
- WorkGraph: Turning AI Coding Sessions into Reusable Memory — adnan_hashmi · 2026-08-06
- Claude Agent Hits 41M Views: Self-Grading Loop is the Real Moat — PrajwalTomar_ · 2026-08-06
- Secret to Top Coding Agents: Get Software Engineers to Look at the Data — HanchungLee · 2026-08-06
- agensis Open-Sources Shared Workspace for Humans and AI Agents — jasonkneen · 2026-08-06