Bridgewater fine-tuned Qwen3-235B to 84.7%, beating frontier models at 1/14 the cost
anacondainc · x · 2026-09-02
A case study from Bridgewater's AIA Labs (with Thinking Machines Lab) shows frontier models scored under 50% on six investor information-triage tasks, rising to 78.2% with expert-written instructions—still short of their 80% bar. Fine-tuning open-weight Qwen3-235B (44.8% out of the box) on their own expert-labeled data reached 84.7%, 30% fewer errors than the best frontier model at roughly one-fourteenth the inference cost. The moat: trustworthy labeling process, not the model.
More from Companies & People
- AI policy figure Matthew Clifford joins Anthropic, based in London for government engagement — SamuelAlbanie · 2026-09-02
- Anthropic reopens Claude Campus Ambassadors with three tracks and $3,600 stipends — claudeai · 2026-09-02
- Gemini Scientific Capabilities Lead Jumps to Anthropic to Push Scientific Discovery — alewkowycz · 2026-09-02
- The Information's Palazzolo says her Google Astra story has more nuance than the headline — steph_palazzolo · 2026-09-02
- Ashley Kramer joins ElevenLabs as Chief Revenue Officer — lukeharries · 2026-09-02
- Berkeley doesn't lack breakthroughs, it lacks storytellers, says The House Fund investor — jfiance · 2026-09-02