Tinfield 1 open-weight model claims to beat Claude Opus 4.8 on Terminal-Bench
victormustar · x · 2026-09-22
Badtheorylabs released Tinfield 1, an open-weight model targeting terminal work and long-horizon software engineering:
- Benchmarks: 33.0 on Terminal-Bench 4.0 and 62.0 on DeepSWE v1.1, claimed to exceed Claude Opus 4.8's published scores on both
- Architecture: 177B total parameters, 6.6B active per token, 256K context
- Availability: three builds today — Base, Compact and Mini
User victormustar describes it as a Qwen3.8 Flash Next finetune and calls it his favorite local model. Scores are vendor-claimed, not independently verified.
More from coding & agent
- Cua Hits 25K GitHub Stars, Puts Claude Code Inside Legacy Desktop Apps — mhdfaran · 2026-09-22
- Cua Puts Claude Code to Work Inside Legacy Desktop Apps With No API — mhdfaran · 2026-09-22
- User Claims Codex Sol Hit by Shrinkflation as Quotas Quietly Shrink — StewartalsopIII · 2026-09-22
- Microsoft's BI-Bench: Frontier LLMs Score Under 50% on End-to-End Business Intelligence — MicrosoftResearch · 2026-09-22
- Dev packages hand-built MCP client into a Spring Boot 4 starter in one dependency — therealdanvega · 2026-09-22
- One-Prompt Call of Duty Clone Hits 20M Views: Inside the Gauntlet Loop Agentic Workflow — mattshumer_ · 2026-09-22