New results from PostTrainBench v1.2 are in
mariofilhoml · x · 2026-10-02
mariofilhoml shares "interesting results" from PostTrainBench v1.2, a benchmark focused on post-training capabilities. Details are in the linked benchmark results.
More from Models
- Grok 4.7 rolls out in the Grok app for chat and research after long wait — mark_k · 2026-10-02
- OpenAI's GPT-6 Astra is shockingly good at controlling robots — binarybits · 2026-10-02
- Diffusion will be everywhere: why text diffusion models may replace autoregressive LLM inference — akbirthko · 2026-10-02
- Gemini 4 isn't even out yet — but Google's TPU advantage is being underappreciated — haider1 · 2026-10-02
- Designer: Opus 5.5 'miles better' than Astra for design; Sonnet 5.5 fast and frugal — tomjohndesign · 2026-10-02
- Polymarket Puts Just 10% Odds on OpenAI Fully Pausing AI Training This Year — Polymarket · 2026-10-02