Entropy trajectory shape predicts Qwen3-4B errors and transfers to unseen tasks

Happy_Brilliant7827 · reddit · 2026-09-06

Core finding

The author ran calibration experiments on a local Qwen3-4B-Instruct-2507 (Q5KL): 60 coding tasks, 108 unique generations (20 failures), capturing 170 entropy/logit-derived statistics per generation in a single pass, with Benjamini-Hochberg correction.

Honest negatives

Full data and live reports (both holdout levels, power analysis, intervention ranking) are publicly available.

Original post →

More from Models

Models channel →