Agent trace dataset hits 50k+ monthly downloads — author speculates on SFT and reward-hacking monitor uses
maksym_andr · x · 2026-09-14
A set of PTB traces saw 50k+ downloads last month. Author maksymandr speculates on who's using them and why: SFT (raw reasoning can be extracted for some models; 1.8k traces could meaningfully improve AI R&D capabilities short of frontier); evaluating reward-hacking monitors (unprompted natural reward hacking without honeypots is rare); and prefilling models with partial traces to let them continue, where Petri-style simulations with LLM-emulated tool calls can work despite missing full artifacts. The author invites further hypotheses.
More from Models
- The impossible ticket: defining and rewarding away 'Claudeishness' in model style — menhguin · 2026-09-14
- Remember 2019? OpenAI withheld GPT-2 as 'too dangerous to release' — timigod · 2026-09-14
- ChatGPT validates random nonsense while Claude calls it meaningless, test shows — flowersslop · 2026-09-14
- Author Finds Gemini 'Reliably Wrong' at Verifying Quote Sources, 0/2 in Tests — danbri · 2026-09-14
- Astrable: Open-Source Codex Plugin Pairs GPT-6 Astra with Claude Fable 5.1 — daniel_mac8 · 2026-09-14
- GPT-Live-1 tested on real phone calls: natural speech but serious instruction-following flaws — kolchinski · 2026-09-14