Randomized YaRN: training on short context with sampled positions boosts 128K reasoning

gregd_nlp · x · 2026-09-05

A study slated for EMNLP 2026 Findings finds that YaRN alone isn't enough for LLM reasoning at 128K context.

The paper proposes Randomized YaRN (RYaRN): training on short-context data while randomly sampling positions from a longer length range. Results show RYaRN improves out-of-distribution reasoning accuracy, particularly at 128K context—a context-extension trick with essentially no extra data cost.

Original post →

More from Research

Research channel →