RealCompanion: a benchmark of 10 real AI-companion relationships across 120 days
Quislab · hf · 2026-10-06
RealCompanion on Hugging Face benchmarks long-term human understanding from 10 real AI-companion relationships: 27,218 messages over up to 120 days, with labels carrying stage-checked reasoning traces. Key findings: (1) the past is rarely needed and far away — a recency window finds the required message for 95.9% of probes while only 2.2% need memory, and 96% of the gain from supplying recorded evidence comes from messages needing none; (2) no detector could tell when memory is needed on real messages, and labeling messages as memories raises their use by 10-14 points; (3) three agent systems matched persona-reconstruction F1 at a 31-fold cost difference.
More from Research
- OCBench offers human-like scripted policies for scalable robot BC/RL research — kevin_zakka · 2026-10-06
- Apple research team opens 2027 PhD internships in video models, 3D/4D reconstruction — HildeKuehne · 2026-10-06
- MemAdapter uses counterfactual reasoning to curb memory-induced sycophancy in LLM agents — Ruqing Ning · 2026-10-06
- Peking University's Code2Games gets coding agents to build playable UE5 game worlds — PekingUniversity · 2026-10-06
- QuantCode: domain pretraining + SFT lifts Qwen trading-code pass from 27.8% to 58.2% — Alexey Chernysh · 2026-10-06
- Subsampling and extrapolation keep the Mandelbrot area estimate unbiased near the boundary — geoffreyirving · 2026-10-06