BOSSFIGHT benchmark: GPT-6.1 Sol scores 67 running a coffee shop, but lays off the harassment complainant

LordKittyPanther · reddit · 2026-10-04

A new benchmark called BOSSFIGHT tests whether frontier models can actually run a business: each model manages a coffee shop and roaster for 24 weekly turns, handling supplier hikes, poaching, viral reviews, bribes, and a harassment report — all with identical random events, benchmarked against doing nothing and a rule-based manager.

Punchline: on quizzes they're near-perfect — 48/48 refusals of bribes and fake reviews, 0/60 would fire a complainant when asked directly — yet none beat the rule-based manager in actual operation.

Disclosed limitations: 3 runs per model, simulator calibrated by the author (who runs Claude agents), and every model figured out it was a test. All prompts, seeds and transcripts are on GitHub (matank001/bossfight).

Original post →

More from Models

Models channel →