QuantCode: domain pretraining + SFT lifts Qwen trading-code pass from 27.8% to 58.2%
Alexey Chernysh · hf · 2026-10-06
A study on specializing LLMs for executable Backtrader trading code, with the 400-task QuantCode-Bench. Continued pretraining lifts Qwen3.6-35B-A3B Judge Pass from 27.8% to 33.0%; adding SFT reaches 58.2% pass and 83.5% backtest success, raising agentic first-turn success from 22.3% to 58.3% and 10-turn success from 47.5% to 79.5%. Continued pretraining alone degrades instruction following (final success 47.5%→32.5%), and domain specialization harms structured tool calling, with recovery SFT restoring format but not repository-level agent performance.
More from Models
- Abacus.AI CEO: Chinese open-source models beat US labs on easy tasks, DeepSeek cheapest for agents — bindureddy · 2026-10-06
- HF researcher disputes GLiDE's Decision Index win over Jev as reasoning-boosted — antoine_chaffin · 2026-10-06
- ChatGPT quietly adds lifetime usage tracking to profile settings — Bpelks · 2026-10-06
- Falcon-Emirati: teaching an LLM the Emirati dialect, culture and nuance — Hugging Face Blog · 2026-10-06
- Reddit user questions the hype around DeepSeek 4.1 Flash and Qwen3-Coder 48B Turbo as coding models — soft_troll · 2026-10-06
- Grok web app to add Team Bots: one shared bot, private chats for each member — nima_owji · 2026-10-06