CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 55 upvotes

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

QUESTION — How can LLM alignment data and agent trajectories be efficiently annotated using token-level correction?

The authors present onPanda, an interactive tool designed for efficiently annotating LLM alignment data and agent trajectories via token-level correction. Instead of full post-editing, annotators locate the first inappropriate token, choose a substitute from model candidates or type free-form text, and the system truncates and resumes generation from that corrected prefix. By repeating this locate-correct-continue loop, annotators can precisely steer outputs at a lower cost while preserving the model's sampling distribution for on-policy SFT and preference data generation. It also connects to external tool harnesses for realistic interactive trajectory annotation.

onPanda reduces median annotation time by 52% over manual post-editing.

LichengLiu03 · 21 Sept 2026 read the original ↗
↑