Most AI video model "comparisons" are rigged — here's how to test them fairly
kimmonismus · x · 2026-09-04
kimmonismus argues that most "comparisons" of AI video models aren't real comparisons at all: prompts differ, settings get tweaked, bad clips get edited out — and people still decide which model is better based on that.
So he tested HappyHorse 1.1 and Kling 3.0 under identical conditions, the same way he'd test any tool, and recommends following @HappyHorseATH for the results.
Related event: Blogger Slams Unfair AI Video Model Comparisons, Advocates Rigorous Testing(2 posts)→
More from Models
- ChatGPT's Reddit-flavored reply to a sexual assault victim sparks training-data backlash — airkatakana · 2026-09-04
- Astra Shows Near-Ideal Test-Time Scaling Gains on LifeSciBench Benchmark — soleio · 2026-09-04
- Astra reportedly hits 97% on ARC-AGI-3 without chain-of-thought — FeeAvailable3770 · 2026-09-04
- GPT-6 'Astra' Smashes ARC-AGI Records: 62.7% on Standard Harness, 99.9% on ARC-AGI-3 — repligate · 2026-09-04
- Browser QA harness: GLM 5.3 Flash beats DeepSeek V4 Flash Vision on screenshots — Certain_Pension6305 · 2026-09-04
- Reddit user flags Artificial Analysis as unreliable: same model shows conflicting scores — metigue · 2026-09-04