A custom AI writing benchmark examines sentence rhythm across five models
almmaasoglu · x · 2026-09-07
- A team at @feldarai is building a highly specific writing benchmark for LLMs: naming choices, opening words, punctuation habits, and how models alternate short and long sentences.
- The same analysis was run across five models.
- Key claim: average sentence length tells you very little; the mixing pattern of short and long sentences is a more revealing signal of each model's prose style.
More from Models
- GLM 5.3 Flash gets 90% off via Merge Gateway: $0.012/M input tokens through September — shensi · 2026-09-07
- Dev slams OpenAI Codex: Astra 'doesn't know how to run itself' on long tasks — ivan_bezdomny · 2026-09-07
- Early user on GPT Astra: sloppy at self-managing long tasks and ran out of quota — ivan_bezdomny · 2026-09-07
- Model's Lie groups explanation echoes Sophus Lie's original 1870s pedagogical intent — doodlestein · 2026-09-07
- Gemini bills thinking tokens on top of output while OpenAI counts them inside, dev's measurements show — qaiser_mehdi · 2026-09-07
- Blogger builds human cell model in 30 minutes with GPT-6 Astra, moves AGI timeline to ~1 year — DeryaTR_ · 2026-09-07