A 600k-token relationship test compares how models comment on personal context

cis_female · x · 2026-07-21

A user shared evaluation results for a relationship-understanding prompt, graded by Fable 5.

The post says the test used about 600k tokens of messages with a friend over the past year, then asked multiple models to comment on the relationship with a few specific prompts.

The actual score table is shown in the attached image, but the post itself is mainly a pointer to the eval setup and results.

Related event: Benchmark Compares LLMs on Cost-Efficiency and Deep Context Understanding(4 posts)→

Original post →

More from Models

Models channel →