OpenAI Details New Training Methods for Non-Verifiable Domains in Astra
morqon · x · 2026-09-05
The quoted post from OpenAI's kevinwng says Astra made significant progress on subjective domains like design aesthetics and knowledge work, enabled by new methods for training on non-verifiable domains — tasks without objective ground truth where standard RL is hard to apply directly. He calls this line of work exciting, teases more to come, and invites users to try Astra.
More from Models
- Astra 3D model floods X, but OpenAI missed the viral moment by delaying launch — bindureddy · 2026-09-05
- Qwen3.8 Max jumps 22% on new RSI-Exam benchmark for recursive self-improvement — cihangxie · 2026-09-05
- Stratechery: Anthropic walks back data retention policy, Nvidia earnings, Meta settles — Stratechery · 2026-09-05
- Claude Suddenly Replied in Russian to a User Who Never Spoke It — roshbakeer · 2026-09-05
- OpenAI's Astra uses 'recurrent depth' reasoning, obscuring its thinking process — JacquesThibs · 2026-09-05
- TAOCP open problems released as a dataset to benchmark frontier models — sytelus · 2026-09-05