Proactivity-Gym: 23-agent study shows correct work alone can't sustain user trust
minnesotanlp · hf · 2026-10-06
This work builds foundations for designing and evaluating proactive LLM agents that act before users ask, around three joint principles (3T): Task Capability, Temporal Allocation, and Trust.
- A five-dimension design space: task scope, anticipation horizon, activation trigger, processing timing, intervention depth
- Proactivity-Gym: a simulation testbed with multi-day scenarios, stateful environments, and persona-conditioned simulated users
- Across 23 model-harness configurations, large 3T gaps emerge; LLM judges often conflate capability with trust
- A 30-participant human study shows trust collapses after mistimed interventions even with correct outcomes; users prefer imperfect sleep-time assistance to preserve focus
More from coding & agent
- Dev runs every task in both Codex and Claude Code: Opus 5.5 beats GPT-6 Astra — JeremyNguyenPhD · 2026-10-06
- Codex spins its wheels while Claude Opus shows far deeper understanding, user reports — iruletheworldmo · 2026-10-06
- Urlbox MCP Server Lets You Capture Screenshots, PDFs and Markdown via Natural Language — modelcontextprotocol · 2026-10-06
- OpenAI Dots Thread Steps 7-8: Approval Gates and Cost-Tiered Intelligence — FinanceYF5 · 2026-10-06
- OpenAI Dots Thread Steps 5-6: Wire in Business Tools and Schedules — FinanceYF5 · 2026-10-06
- 10-Step Blueprint Turns OpenAI Dots Into a 24/7 Company Operations Layer — FinanceYF5 · 2026-10-06