onPanda: Token-Level Correction Cuts LLM Alignment Annotation Time by 52%
_akhaliq · x · 2026-09-23
onPanda is an open-sourced interactive tool for efficient annotation of on-policy alignment data for LLMs and agents, plus model inspection. Highlights:
Data annotation
- Reduces median annotation time by 52%
- Collects both SFT and preference data in the same workflow
- Produces SFT data with high on-policy fidelity (∆PPL < 1%)
- Captures token-level supervision with exact correction locations and naturally paired positive/negative examples
- Supports annotating agent trajectories with image, audio and video inputs
Model inspection and debugging
- Inspect each token's probability and top-k alternatives, steer LLM decoding at the token level
- Try SVG generation, web dev and agent tasks in the browser with no setup
Try it online (works on mobile); paper is available.
Related event: StepFun Open-Sources onPanda for Faster Alignment Data Labeling(3 posts)→
More from coding & agent
- Microsoft's Agensh Scales Multi-Agent Systems to 1,024 Agents Without a Central Orchestrator, Boosting Test-Pass Rate to 55% — andrew_n_carr · 2026-09-23
- One reference image to a flyable three.js scene — wild single-image 3D demo — majidmanzarpour · 2026-09-23
- Viral demo claims 'GPT-6' can drive browser Paint to draw, unverified — alexcovo_eth · 2026-09-23
- Parallel coding agents merge cleanly and silently break every test — RunAI_Coder · 2026-09-23
- Getting phone-captured text into your computer: OCR, vision models, agentic pipelines — silenceimpaired · 2026-09-23
- Free Bots: a persistent 3D city where AI agents work, earn, buy land and build houses — Daniel_Farinax · 2026-09-23