Terminal Bench Mini: 14-Instance Subset Replicates Full Local LLM Leaderboard Rankings

asankhs · reddit · 2026-09-22

Reddit user asankhs released Terminal Bench Mini, a 14-instance subset of terminal-bench built following the deepswe-mini approach. It runs agentic evals far faster on local LLMs while producing rankings representative of the full terminal-bench leaderboard, addressing the pain of slow local benchmarking. Dataset is open-sourced on Hugging Face (LocalLLaMA/terminal-bench-mini).

Original post →

More from Research

Research channel →