Encode Bench finds Base64 output quality tracks AI intelligence scores at 0.91

Valuable-Repeat-7347 · reddit · 2026-07-22

The author introduces Encode Bench, an open benchmark that asks models to solve tasks and return the answer as a Base64 payload.

The surprising result is that the benchmark’s pass rate correlates strongly with broader AI scores: Pearson correlation is 0.91 with the Artificial Analysis Intelligence Index and 0.94 with its Agentic Index, based on the current nine-model snapshot. The author stresses that this is a small observational sample and does not prove Base64 measures intelligence or causation.

Key findings:

The author argues the benchmark likely mixes reasoning, exactness, tokenizer behavior, post-training, endpoint reliability, and inference limits. They say the most important missing control is a plain-text version of the same battery without the Base64 requirement, and they also want to test hexadecimal and random strings.

Original post →

More from Models

Models channel →