Testing LLMs on Relationship Understanding with 600k Tokens

Recently, an evaluation of large language models' "relationship understanding" capabilities caught the community's attention. Reviewer @cisfemale fed approximately 600,000 tokens of real chat logs with a friend over the past year into multiple models, asking them to comment on the relationship based on specific questions, which were then scored by Fable 5. By introducing cost data to compare the cost-effectiveness of each model, this evaluation provided a new perspective on assessing model performance in real-world emotional and relational scenarios.

Evaluation Mechanism and Rankings

The results were scored by Fable 5. @cisfemale noted that while there was some fluctuation in the scoring, the margin was small, and the model rankings remained largely consistent. In the cost-effectiveness leaderboard, Claude Fable 5 led with a score of 4.30, followed closely by Claude Opus 4.8. Furthermore, the cost comparison per task covered major products including Claude, Mistral, Qwen, GPT, Gemini, GLM, Kimi, and DeepSeek, with prices ranging from $2.75 to $0.02.

Cost-Effectiveness Insights and Data Corrections

In the horizontal comparison, Meta Muse Spark 1 performed near the best and clearly outperformed Grok-4.20 in terms of cost-effectiveness. @cisfemale pointed out that the industry previously considered these two models as cheap alternatives with similar performance, but actual tests showed Grok-4.20 performed significantly worse. Additionally, @cisfemale later issued a correction, reminding readers that the cost figures in some charts were incorrect and that the accurate data should be referenced from another chart.

2026-07-20 ~ 2026-07-22 · 6 related posts

Primary sources