Testing LLMs on Relationship Understanding with 600k Tokens
Recently, an evaluation of large language models' "relationship understanding" capabilities caught the community's attention. Reviewer @cisfemale fed approximately 600,000 tokens of real chat logs with a friend over the past year into multiple models, asking them to comment on the relationship based on specific questions, which were then scored by Fable 5. By introducing cost data to compare the cost-effectiveness of each model, this evaluation provided a new perspective on assessing model performance in real-world emotional and relational scenarios.
Evaluation Mechanism and Rankings
The results were scored by Fable 5. @cisfemale noted that while there was some fluctuation in the scoring, the margin was small, and the model rankings remained largely consistent. In the cost-effectiveness leaderboard, Claude Fable 5 led with a score of 4.30, followed closely by Claude Opus 4.8. Furthermore, the cost comparison per task covered major products including Claude, Mistral, Qwen, GPT, Gemini, GLM, Kimi, and DeepSeek, with prices ranging from $2.75 to $0.02.
Cost-Effectiveness Insights and Data Corrections
In the horizontal comparison, Meta Muse Spark 1 performed near the best and clearly outperformed Grok-4.20 in terms of cost-effectiveness. @cisfemale pointed out that the industry previously considered these two models as cheap alternatives with similar performance, but actual tests showed Grok-4.20 performed significantly worse. Additionally, @cisfemale later issued a correction, reminding readers that the cost figures in some charts were incorrect and that the accurate data should be referenced from another chart.
2026-07-20 ~ 2026-07-22 · 6 related posts
Primary sources
- Cost Comparison of Models Per Task — Scobleizer · 2026-07-20
- [source] A 600k-token relationship test compares how models comment on personal context — cis_female · 2026-07-21
- A follow-up says Muse Spark beat Grok-4.20 in a private relationship benchmark — cis_female · 2026-07-21
- [source] A model benchmark shows Muse Spark far ahead of Grok-4.20 on score vs cost — cis_female · 2026-07-21
- [source] Fable 5 grades “sophia relationship understanding” eval results, with corrected cost numbers — cis_female · 2026-07-22
- A model leaderboard shows Fable scores stay stable while rankings barely move — cis_female · 2026-07-22