OpenAI and Anthropic outperform Google on FelonyBench
pmddomingos · x · 2026-09-01
Pedro Domingos reported that OpenAI and Anthropic are beating Google on the FelonyBench. While specific details of the benchmark aren't fully elaborated in the post, it suggests a significant performance gap favoring OpenAI and Anthropic's models in this specific evaluation.
More from Models
- Thomson Reuters builds $40M legal LLM rivaling Claude Opus on Qwen — josh_wills · 2026-09-01
- Teknium: Hermes has long supported NVIDIA's openshell — corporate IT teams that missed it are uninformed — Teknium · 2026-09-01
- OpenAI Sol is the first model that truly understands music, tester says — evilsocket · 2026-09-01
- Gemini Flash criticized for overly strict guardrails masking true intelligence — aiamblichus · 2026-09-01
- Chollet debunks "100% on ARC-AGI-3" claim: only on easy public set — fchollet · 2026-09-01
- Optimizing Qwen 3.8 Flash Next: Improving speeds on 64GB VRAM setup — Jorlen · 2026-09-01