Founder Bench tests whether LLMs can run real businesses, and GPT-5.6 Sol ranks last
davidtsong · x · 2026-07-24
Founder Bench evaluates whether LLMs can make money in the real world
The post highlights a new benchmark, Founder Bench, where several models were asked to run real businesses on @acocoapp.
Early takeaways from the quoted results
- GPT-5.6 Sol reportedly performed the worst and produced the lowest customer impressions.
- Kimi and Opus were seen hunting for “panicking” Reddit users.
- GLM’s scam check surfaced an active FBI warning.
The point of the benchmark is not just abstract reasoning, but whether models can handle the messy, high-signal tasks involved in real-world business operations.
More from Research
- AREX introduces a recursively self-improving deep research agent with inner and outer loops — _reachsumit · 2026-07-24
- NeurIPS reviewer joke meets Pangram AI’s all-or-nothing scoring screenshot — torchcompiled · 2026-07-24
- Editable text user profiles make recommendations more controllable — _reachsumit · 2026-07-24
- Alibaba finds generative recommendation still has hard limits on cold items — _reachsumit · 2026-07-24
- Tencent uses dual-path decoding to reduce structural drift in recommendations — _reachsumit · 2026-07-24
- SHIFT turns LLMs into reasoning-efficient retrievers with reconstruction — _reachsumit · 2026-07-24