Group Bench launches ~100 group theory problems for testing AI agents
Sauers_ · x · 2026-09-10
Group Bench is a new benchmark with roughly 100 group theory problems designed to test AI agents' reasoning, with the problem set publicly linked. The author invites anyone who solves a problem to open an issue on the repo, creating an open scoreboard. It's a ready-to-use math benchmark for anyone evaluating agents on abstract reasoning.
Related event: Group Bench: ~100 Group Theory Problems to Test AI Agents(2 posts)→
More from Research
- Eric Drexler on the Hugging Face incident: system structure, not alignment, prevents AI collusion — sebkrier · 2026-09-10
- David Chalmers talks Anthropic's j-space and global workspace theory aboard a boat in the Galapagos — PeterBowdenLive · 2026-09-10
- A visual deep-dive catalogs 42+ representations of 3D, praised by HF engineer — pcuenq · 2026-09-10
- Ben Recht's forecasting lecture argues probability conflates frequency and belief — beenwrekt · 2026-09-10
- Spiced self-play accepted at CoRL: just 30 minutes of human data biases agents to right conventions — EugeneVinitsky · 2026-09-10
- Understanding FlashAttention: A Handbook Tracing FA1 to FA4 and Why HBM Traffic, Not FLOPs, Is the Bottleneck — techNmak · 2026-09-10