Papers with Code Adds Model Size Filtering for Benchmarks
NielsRogge · x · 2026-08-24
Papers with Code has updated to allow filtering benchmarks based on model size. For example, Terminal Bench 2.1, a verified iteration of the agent evaluation benchmark, is now available with 26 updated tasks for reliable testing in containerized environments. The listing includes entries for various frontier models like GPT-5.6, Kimi K3, and GLM-5.3 (listed with future dates).
Related event: Papers with Code Adds Model Size Filtering to Leaderboards(2 posts)→
More from coding & agent
- Google launches Developer Knowledge MCP integrated into gcloud CLI — rseroter · 2026-08-24
- Built a Cloudflare Agent to diagnose and fix 'slop' in web design — craigsdennis · 2026-08-24
- Obsidian's Smart Chat saves AI thread links and status back into your notes — AINewsletter · 2026-08-24
- Vibe coding feels faster until your project grows and you can't understand its history — Warm-Reaction-456 · 2026-08-24
- We gave agents real email addresses and broke deliverability, threading, and privacy — saltexx · 2026-08-24
- Same model, different coding agent: harness choice swings scores from 10/10 to 0/10 — oliver-zehentleitner · 2026-08-24