Pearmut: an open platform making human translation eval as easy as automatic eval
zouharvi · x · 2026-09-14
The arXiv paper "Pearmut: Human Evaluation of Translation Made Trivial" presents a lightweight human-evaluation platform that makes end-to-end human evaluation as easy to run as automatic evaluation, removing the heavy engineering and operational overhead of existing tools like Appraise. It supports standard protocols (DA, ESA, MQM) and is extensible, with document-level context, absolute and contrastive evaluation, attention checks, ESAAI pre-annotations, and static/dynamic assignment strategies. The goal is to make reliable human evaluation a routine part of model development and diagnosis. Accepted as an EMNLP 2026 Demo; it mostly replaces Appraise at WMT.
Related event: Pearmut Platform and cESA Protocol Head to EMNLP(3 posts)→
More from coding & agent
- Survey of 60+ image generators: retrieval, not capture, is the real pain point — shivam_dewan · 2026-09-14
- LeanDB: Theoric Labs builds a strongly typed Lean frontend for SQL databases — hargup13 · 2026-09-14
- iOS AI dev workflow: AppKit + Figma import + Claude Opus 5 gives best UI fidelity — dotey · 2026-09-14
- Dev seeks blueprint for agent that triages incidents across PagerDuty, Datadog, GitLab and Slack — cruelcaricature · 2026-09-14
- Redpen CLI checks whether your coding agent's 'done' claim matches repo, test and build evidence — Scobleizer · 2026-09-14
- Infinite Bookshelf: open-source app generates a whole book from one prompt using Llama on Groq — Roger_M_Taylor · 2026-09-14