Encode Bench finds Base64 output quality tracks AI intelligence scores at 0.91
Valuable-Repeat-7347 · reddit · 2026-07-22
The author introduces Encode Bench, an open benchmark that asks models to solve tasks and return the answer as a Base64 payload.
The surprising result is that the benchmark’s pass rate correlates strongly with broader AI scores: Pearson correlation is 0.91 with the Artificial Analysis Intelligence Index and 0.94 with its Agentic Index, based on the current nine-model snapshot. The author stresses that this is a small observational sample and does not prove Base64 measures intelligence or causation.
Key findings:
- The benchmark contains 24 deterministic tasks across encoding fidelity, instruction following, arithmetic, logic, code reasoning, and structured data.
- Each task is run three times, producing 72 scored trials per model.
- Current top result: GPT-5.6 Sol — 70/72 (97.2%).
- Raw encoding-fidelity tasks were the hardest category, while code reasoning was the easiest.
- Many failures were not malformed Base64; the payload decoded correctly but contained the wrong answer.
The author argues the benchmark likely mixes reasoning, exactness, tokenizer behavior, post-training, endpoint reliability, and inference limits. They say the most important missing control is a plain-text version of the same battery without the Base64 requirement, and they also want to test hexadecimal and random strings.
More from Models
- Trelis releases Tiron, an open-weights transcription and diarization model — TrelisResearch · 2026-07-22
- Google Search is said to be using Gemini 3.5 Flash-Lite, with benchmark gains shown — gaganghotra_ · 2026-07-22
- A Gemini joke imagines it learning reality from a stale Google Cache internet — teortaxesTex · 2026-07-22
- Testing Kimi K3 for Frontend: Delivers a Full Day's Work in 1 Hour — FuSheng_0306 · 2026-07-22
- Gemma’s “agentic” pitch falls apart in a local RTX 5090 test — kolliwolli · 2026-07-22
- Kimi K3 shows benchmark awareness in 61% of trajectories, study says — gleech · 2026-07-22