A staged diagnosis finds most short-text generation loss comes from the codec
ITMO · hf · 2026-07-29
This paper argues that failures in compressed short-text generation should be diagnosed in stages: first the codec, then the latent generator.
- In a controlled 64→16 TinyStories case study with a hierarchical VQ-VAE-2 codec and masked discrete diffusion generator (MDLM), the authors separate codec reconstruction, latent generation, and auxiliary diagnostics under one GPT-2-based scorer.
- Reconstructing through the codec alone raises median external PPL from 15.17 to 27.36 and p95 from 25.10 to 98.91, implying most quality loss happens before latent generation.
- Code-space MDLM still beats token-space diffusion, cutting mean/median/p95 PPL by 32.9% / 30.9% / 36.6%.
- Geometry-aware regularization improves latent diagnostics but not decoded text quality, so the practical takeaway is: audit the codec first.
More from Embodied
- Four Major Bottlenecks for AI Hardware Companies: Marketing Data Silos and Shifting Needs — 创业邦 · 2026-07-29
- Screenless health wearables move from workout tracking to always-on recovery — 创业邦 · 2026-07-29
- ALICE5 made its RoboCup 2026 debut two weeks ago — CyberRobooo · 2026-07-29
- AI agent phones are turning the smartphone into a permissions layer — APPSO · 2026-07-29
- AeiROBOT’s ALICE v4 humanoid serves drinks at a cocktail party with near-zero-delay teleop — CyberRobooo · 2026-07-29
- Daeduck Electronics pilots wheeled humanoid ALICE M1 on a PCB line for high-mix production — CyberRobooo · 2026-07-29