Encode Bench finds Base64 output quality tracks AI intelligence scores at 0.91
Valuable-Repeat-7347 · reddit · 2026-07-22
The author introduces Encode Bench, an open benchmark that asks models to solve tasks and return the answer as a Base64 payload.
The surprising result is that the benchmark’s pass rate correlates strongly with broader AI scores: Pearson correlation is 0.91 with the Artificial Analysis Intelligence Index and 0.94 with its Agentic Index, based on the current nine-model snapshot. The author stresses that this is a small observational sample and does not prove Base64 measures intelligence or causation.
Key findings:
- The benchmark contains 24 deterministic tasks across encoding fidelity, instruction following, arithmetic, logic, code reasoning, and structured data.
- Each task is run three times, producing 72 scored trials per model.
- Current top result: GPT-5.6 Sol — 70/72 (97.2%).
- Raw encoding-fidelity tasks were the hardest category, while code reasoning was the easiest.
- Many failures were not malformed Base64; the payload decoded correctly but contained the wrong answer.
The author argues the benchmark likely mixes reasoning, exactness, tokenizer behavior, post-training, endpoint reliability, and inference limits. They say the most important missing control is a plain-text version of the same battery without the Base64 requirement, and they also want to test hexadecimal and random strings.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11