8 frontier LLMs benchmarked on 50 tasks across 10 dimensions over two days

lxfater · x · 2026-08-17

Developer servasyyai spent two days and burned a large number of tokens benchmarking 8 frontier LLMs across 10 dimensions (D1–D10) and 50 sub-items. Each item is scored 0–5 and aggregated into a 100-point total.

The article keeps two original leaderboards (Round 1 for one-shot complex tasks and Round 2) plus one combined reference board to reduce single-round randomness. Reposter lxfater joked the piece was too long and asked for just the conclusions; full rankings are at the linked article.

Original post →

More from Models

Models channel →