PRAXIST open-sourced: 49 golds on MLE-bench at 1/12th the cost of Claude Code
SignalCompetitive582 · reddit · 2026-08-28
Sapient Intelligence (the HRM architecture team) released PRAXIST, a lineage-centered generational system for autonomous R&D agents. It converts reproducible artifacts and evaluator outcomes into a typed evidence graph, separating local artifact construction from cohort-level evidence synthesis so later attempts inherit validated mechanisms, unresolved claims and constraints—stopping long campaigns from relearning the same lessons.
Key results: On the 75-task MLE-bench with official graders, PRAXIST earns 60 medals (80.0%, 49 gold) versus 55 medals (73.3%, 34 gold) for a Claude Code baseline on Claude Opus 4.8—at US$3,054 model spend versus US$38,370, roughly a twelfth of the cost. Four open-ended case studies (quant trading, LiDAR-inertial-visual SLAM, tokamak magnetic control, rocket landing) all beat task-native baselines with the discovery path on record.
Licensing & signal: PRAXIST is open source, but companies above US$1M annual revenue need a commercial license. The poster reads it as proof that harnesses dramatically amplify LLMs—and possibly a hint that HRM itself didn't scale well.
More from coding & agent
- GLM-5.3-Flash built a full Blender scene autonomously in 12 hours — smtabatabaie · 2026-08-28
- How do developers decide to adopt an MCP server? Which signals matter most — SnooPuppers6082 · 2026-08-28
- Why do agent coding runs stop at 30 minutes? The missing machine-checkable feedback loop — EntireBig7258 · 2026-08-28
- No mobile app? User feeds Conductor's API docs to Grok Bot to manage cloud agents — iannuttall · 2026-08-28
- AQuA paper: if an agent can edit its tool schema, you're not evaluating the same system — creditme7 · 2026-08-28
- 4 Hours on a Train: A Fold Plus Remote Desktop Is Enough to Run All Your Agents — kevinkern · 2026-08-28