New Model Scores 55 on Rescaled AA Benchmark, Competitive but Disappointing
teortaxesTex · x · 2026-08-13
Based on self-published evals, Sol estimates that a new model scores 55 on the newly rescaled AA benchmark. The author notes that while this score is competitive, it is still somewhat disappointing.
More from Models
- DeepSeek V4 API Fingerprint Changes, Hinting at New Checkpoints — teortaxesTex · 2026-08-13
- Users Report Severe Model Degradation Across Google's APIs — EthanBeMe · 2026-08-13
- xAI Releases Grok 4.6: AI Now Autonomously Optimizes Inference Code — rayhotate · 2026-08-13
- Inside Grok 4.6: AI Explores 297 Optimizations, Boosting Inference Throughput — rayhotate · 2026-08-13
- Qwen3.8-27B Model Surfaces on ModelScope Ahead of Hype — Ok-Shower7286 · 2026-08-13
- Leaked Grok 4.6 Leads in Agent Tasks, Beats GPT-5.6 in Coding — rohanpaul_ai · 2026-08-13