Bespoke Labs post-trains Inkling with SFT+GRPO, gaining 57pp on repo-level coding and 40% token efficiency
AlexGDimakis · x · 2026-09-04
Bespoke Labs researchers post-trained the Inkling base model on a specific GitHub repo using SFT on strong-teacher trajectories plus GRPO RL in repo-specialized environments. SFT alone added 52pp on the held-out fontTools eval; RL lifted it to 57pp, with good transfer to Terminal-Bench 2.1 and SWE-Bench Lite and 40% better token efficiency.
More from coding & agent
- Lab lessons from Anthropic MHS: keep fast control out of the model — Empty-Abalone-2952 · 2026-09-04
- MCP veteran launches TDQS, an open spec scoring 15,000+ tool definitions — punkpeye · 2026-09-04
- WebMCP Computer: one URL gives any coding agent a disposable OS in the browser — prd_008 · 2026-09-04
- Grok Bot Hands Out 50 x $200 Codes as User Shares Orchestrator-Bot Workflow — omarsar0 · 2026-09-04
- Dev Spent 5 Months of Claude Max Improving His Open-Source App Store Connect CLI — rudrank · 2026-09-04
- WHALE: a simple recipe to jointly optimize an LLM's weights and harness — lateinteraction · 2026-09-04