Bridgewater fine-tuned Qwen3-235B to 84.7%, beating frontier models at 1/14 the cost

anacondainc · x · 2026-09-02

A case study from Bridgewater's AIA Labs (with Thinking Machines Lab) shows frontier models scored under 50% on six investor information-triage tasks, rising to 78.2% with expert-written instructions—still short of their 80% bar. Fine-tuning open-weight Qwen3-235B (44.8% out of the box) on their own expert-labeled data reached 84.7%, 30% fewer errors than the best frontier model at roughly one-fourteenth the inference cost. The moat: trustworthy labeling process, not the model.

Original post →

More from Companies & People

Companies & People channel →