PRAXIST open-sourced: 49 golds on MLE-bench at 1/12th the cost of Claude Code

SignalCompetitive582 · reddit · 2026-08-28

Sapient Intelligence (the HRM architecture team) released PRAXIST, a lineage-centered generational system for autonomous R&D agents. It converts reproducible artifacts and evaluator outcomes into a typed evidence graph, separating local artifact construction from cohort-level evidence synthesis so later attempts inherit validated mechanisms, unresolved claims and constraints—stopping long campaigns from relearning the same lessons.

Key results: On the 75-task MLE-bench with official graders, PRAXIST earns 60 medals (80.0%, 49 gold) versus 55 medals (73.3%, 34 gold) for a Claude Code baseline on Claude Opus 4.8—at US$3,054 model spend versus US$38,370, roughly a twelfth of the cost. Four open-ended case studies (quant trading, LiDAR-inertial-visual SLAM, tokamak magnetic control, rocket landing) all beat task-native baselines with the discovery path on record.

Licensing & signal: PRAXIST is open source, but companies above US$1M annual revenue need a commercial license. The poster reads it as proof that harnesses dramatically amplify LLMs—and possibly a hint that HRM itself didn't scale well.

Related event: PRAXIST: Open-Source Multi-Agent Research Framework Scores 49 Gold Medals on MLE-bench(5 posts)→

Original post →

More from coding & agent

coding & agent channel →