Extropic post-trains Qwen3.6-35B-A3B on Prime Intellect, nearly tripling evals in ~100 GRPO steps
MarvinTBaumann · x · 2026-10-02
Quantum computing company Extropic post-trained Qwen3.6-35B-A3B for thermodynamic ML research using Prime Intellect's platform.
- Eval results on held-out tasks improved nearly 3x after only 100 GRPO steps
- The team built a custom RL environment with verifiers
- Training ran on Prime Intellect's Hosted Training, Prime Sandboxes, and Prime Inference, eliminating the need to manage multi-node GPU infrastructure so researchers could focus on research
More from coding & agent
- New tldraw plugin lets ChatGPT map bugs on canvas and sketch live fixes — max__drake · 2026-10-02
- Dev says Codex Cloud's stateful, warm cloud sessions are the real game changer from Dev Day — paw_lean · 2026-10-02
- AI port pushes code translation to its limit, generating WASM-compiled animation kernels — ricklamers · 2026-10-02
- Swarms adds 0-100 Security Scores to every marketplace prompt via SkillScanner — KyeGomezB · 2026-10-02
- App Store Connect CLI 5.9.0 ships preflight checks that catch rejections before you submit — rudrank · 2026-10-02
- Agent Lens aligns LLM judges on production traffic to catch agent failures — _ScottCondron · 2026-10-02