Papers with Code Adds Model Size Filtering for Benchmarks

NielsRogge · x · 2026-08-24

Papers with Code has updated to allow filtering benchmarks based on model size. For example, Terminal Bench 2.1, a verified iteration of the agent evaluation benchmark, is now available with 26 updated tasks for reliable testing in containerized environments. The listing includes entries for various frontier models like GPT-5.6, Kimi K3, and GLM-5.3 (listed with future dates).

Related event: Papers with Code Adds Model Size Filtering to Leaderboards(2 posts)→

Original post →

More from coding & agent

coding & agent channel →