Georgia Tech's TextReg Fixes Prompt Distributional Overfitting, Gains Up to +11.8% OOD
GeorgiaTech · hf · 2026-10-06
Georgia Tech researchers study "prompt distributional overfitting": iteratively optimized prompts (e.g., via TextGrad) grow longer, accumulate sample-specific rules, and generalize poorly out of distribution. They formalize this via a dual-factor "representational inefficiency" measure combining capacity cost and scope narrowness.
TextReg implements a soft-penalty objective through regularized textual gradients, combining Dual-Evidence Gradient Purification, Semantic Edit Regularization, and Regularization-Guided Prompt Update.
Across reasoning benchmarks, TextReg substantially improves OOD generalization: up to +11.8% over TextGrad and +16.5% over REVOLVE.
More from Research
- Sander Dieleman on why continuous diffusion language models are making a comeback — LucaAmb · 2026-10-06
- Reading the source of 7 LLM eval tools uncovered 13 scoring bugs, 6 fixes merged — maverick_man1111 · 2026-10-06
- Terminal-Bench Pro: 400 tasks across 8 domains with zero contamination risk — thisguyknowsai · 2026-10-06
- ROME's training pipeline: 500B-token CPT, error-masked SFT, chunk-level RL — thisguyknowsai · 2026-10-06
- ROME's IPA assigns RL credit at chunk level, not token level, for tool-use agents — thisguyknowsai · 2026-10-06
- ROME team open-sourced full agent infra: ROLL, ROCK and iFlow CLI before training the model — thisguyknowsai · 2026-10-06