Llama-3-8B accuracy jumps to 94% on planning benchmarks
mdancho84 · x · 2026-08-27
Benchmark results show that Llama-3-8B's accuracy on planning benchmarks surged from 28% to 94%. This significant leap indicates the emergence of a completely new capability rather than just incremental improvement.
Related event: MIT's PDDL-INSTRUCT Method Boosts LLM Planning Accuracy from 28% to 94%(5 posts)→
More from Models
- microduck training has started, HuggingFace CEO confirms — huggingface · 2026-08-27
- Qwen3.8-Flash-Next Benchmarked Locally: First Model to Break 94% on 128GB Mac — tolitius · 2026-08-27
- OpenAI adds Luna Reserve fallback for Codex usage limits — kimmonismus · 2026-08-27
- Open Source Models Surpass SOTA: Kimi, GLM, Qwen Outscore GPT-5.5 — solyarisoftware · 2026-08-27
- Sora locks everyone out: all accounts logged out, re-login fails — Shrapnel_FEH · 2026-08-27
- Community open-sources theoretical reconstruction of Claude Mythos — Shruti_0810 · 2026-08-27