onPanda: token-level correction tool cuts alignment data annotation time by 52%
stepfun-ai · hf · 2026-09-22
stepfun-ai released onPanda, an open-source tool for efficiently annotating LLM alignment data and agent trajectories.
- Core interaction: token-level correction — annotators locate the first inappropriate token in a response, pick a substitute from the model's candidate tokens or type the fix, and the system truncates and regenerates from the corrected prefix in a locate-correct-continue loop.
- Efficiency: a small controlled study shows a 52% reduction in median annotation time versus manual post-editing.
- Data value: since most final tokens come from the model itself, the data preserves the sampling distribution, making it ideal for on-policy SFT and preference data; corrections also yield fine-grained supervision with precise positions and naturally paired positive-negative samples.
- Extras: connects to external tools/harnesses for interactive trajectory annotation, with the Panda-CVL dataset and a token-level correction benchmark also released.
More from Research
- Physicist pushes back on AI math proof criticism: strategy allowed by Clay criteria — skdh · 2026-09-22
- VR mocap pipeline captures 1,185 human trajectories to teach Unitree G1 to navigate clutter — Scobleizer · 2026-09-22
- AI memory system built with Jev claims 10x speed, 6x lower cost — realferrari · 2026-09-22
- AF3-style confidence heads surprisingly insensitive to diffusion module coordinates — chaitjo · 2026-09-22
- riderless: A Zero-Token Decision API on Gemma 4 Hits 93.1% Accuracy on One RTX 5090 — Pale-Soil-2524 · 2026-09-22
- Shenzhen University and HKUST propose ECA: evidence checks before agent actions — jiqizhixin · 2026-09-22