Pearmut: an open platform making human translation eval as easy as automatic eval

zouharvi · x · 2026-09-14

The arXiv paper "Pearmut: Human Evaluation of Translation Made Trivial" presents a lightweight human-evaluation platform that makes end-to-end human evaluation as easy to run as automatic evaluation, removing the heavy engineering and operational overhead of existing tools like Appraise. It supports standard protocols (DA, ESA, MQM) and is extensible, with document-level context, absolute and contrastive evaluation, attention checks, ESAAI pre-annotations, and static/dynamic assignment strategies. The goal is to make reliable human evaluation a routine part of model development and diagnosis. Accepted as an EMNLP 2026 Demo; it mostly replaces Appraise at WMT.

Related event: Pearmut Platform and cESA Protocol Head to EMNLP(3 posts)→

Original post →

More from coding & agent

coding & agent channel →