Felda AI Builds Fine-Grained Writing Benchmark Spanning Rhythm and Naming Habits
On September 7, the Felda AI team (author @almmaasoglu) shared a series of posts on X about their in-house writing evaluation benchmark and its initial results: instead of scoring overall writing style, the evaluation drills into concrete features of a model's writing, including naming preferences, opening word choices, punctuation habits, and the "sentence rhythm" of alternating short and long sentences within a text, with early charts for the sentence-rhythm dimension.
Confirmed
- The team ran the same sentence-rhythm analysis on five models, comparing how much each mixes short and long sentences within an article.
- Their conclusion: average sentence length carries limited information; the real differences between models show up in the diversity of sentence rhythm.
- An interesting finding: OpenAI and Anthropic models both favor the same name; each scatter plot covers 20 outputs, with lines connecting the same prompt group across models to visualize naming preference distributions.
Not yet confirmed
- A fuller official benchmark report and more findings have not been published yet; the author says they are coming.
Why it matters
- Mainstream writing evaluations mostly focus on macro style or win rates; this benchmark pushes granularity down to quantifiable features like punctuation, naming and sentence rhythm, offering a new lens for distinguishing models' writing personalities and highlighting the limits of common metrics like "average sentence length."
2026-09-07 ~ 2026-09-07 · 5 related posts
Primary sources
- AI startup Felda dissects writing quality down to sentence rhythm and punctuation in its new benchmark — almmaasoglu · 2026-09-07
- A custom AI writing benchmark examines sentence rhythm across five models — almmaasoglu · 2026-09-07
- [source] Felda AI builds writing benchmark analyzing names, punctuation and sentence rhythm — almmaasoglu · 2026-09-07
- [source] Five-model writing style comparison shows average sentence length misses the mark — almmaasoglu · 2026-09-07
- [source] OpenAI and Anthropic models share a favorite name, writing benchmark finds — almmaasoglu · 2026-09-07