Terminal Bench Mini: 14-Instance Subset Replicates Full Local LLM Leaderboard Rankings
asankhs · reddit · 2026-09-22
Reddit user asankhs released Terminal Bench Mini, a 14-instance subset of terminal-bench built following the deepswe-mini approach. It runs agentic evals far faster on local LLMs while producing rankings representative of the full terminal-bench leaderboard, addressing the pain of slow local benchmarking. Dataset is open-sourced on Hugging Face (LocalLLaMA/terminal-bench-mini).
More from Research
- HWREBench: AI researcher hacks Amazon smart devices daily to benchmark hardware reverse engineering — johnowhitaker · 2026-09-22
- TMLR submissions quadruple on AI-generated influx; ICLR 2027 caps single authors — petitegeek · 2026-09-22
- Stanford's VirtualBiotech puts tens of thousands of AI scientist agents in Science, NYT reports — StanfordAILab · 2026-09-22
- Irit Dinur wins Gödel Prize for her landmark 2005 proof of the PCP theorem — willcb · 2026-09-22
- Researcher proposes using RL to teach AI when to give up — sqcai · 2026-09-22
- Nature publishes Delphy: scalable near-real-time Bayesian phylogenetics for outbreak tracking — burny_tech · 2026-09-22