Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points
julianharris · x · 2026-09-20
A user digging into the eval report found that Bonsai 2 27B's accuracy on Terminal-bench and SWE-bench drops by nearly 20 points with safety guardrails enabled — data the reviewer says should have been in the main benchmark table from the start. Commenters called the degradation "a serious lobotomy," criticizing the vendor for publishing only unguarded scores and obscuring the real capability cost of its safety measures.
More from coding & agent
- AI Browser Agent Fills Out Immigration Card From Passport Photo and Email, Even Handles CAPTCHAs — TianbaoX · 2026-09-20
- Decision Model Jev Goes Free on Venice API: Typed Answers, No Prose, No JSON Wrangling — 0xAllen_ · 2026-09-20
- Parent builds weekly Claude skill to scrape curriculum reading and auto-place library holds — nwilliams030 · 2026-09-20
- Dev renders Google Maps in the terminal with a single Rust/CUDA megakernel — bilawalsidhu · 2026-09-20
- My coding model + harness stack: Kimi K3 tops OSS, Claude Code doubles the cost — Yuchenj_UW · 2026-09-20
- Entity resolution with Jev cut pipeline costs 99.56% and boosted throughput 7.35x — hrishioa · 2026-09-20