LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions
Myra Cheng, Lujain Ibrahim, Grace Liu, Michelle S. Lam, Vishakh Padmakumar, Nick Madibekov, Diyi Yang, Dan Jurafsky
cs.CY, cs.AI, cs.CL
2026-09-14
A Stanford typology finds LLM-as-oracle use rising in 68K public prompts and 140K donated messages. Users underestimate offloading and ask AI to decide more than a believed human.
People have long handed uncertain personal questions to outside authorities: oracles, astrology, coin flips. LLMs make that move instant, specific, and personalized. Subjective questions have no ground-truth label, so classic overreliance tests do not apply. This Stanford, Oxford, and Carnegie Mellon paper asks how often users offload judgment and decisions to models, whether they know they are doing it, and what drives the behavior.
The typology is about what gets offloaded, not the topic (career versus dating). On beliefs: Interpreter reads a situation or the self, versus Contextualizer listing several readings; Evaluator issues a normative verdict, versus Deliberator laying out considerations. On actions: Decision-maker picks a course, versus Strategist listing options; Ghost-writer writes the message, versus Editor revising a user draft. Everything else is NA. The coding is conservative: self-disclosure and journaling do not count as oracle use.
Gemini-2.5-Pro labels the nine classes. Agreement with two experts on 219 stratified items is kappa>0.65; Fleiss kappa among the three is 0.66.
Three data sources. Public logs: 60,376 English WildChat prompts (2023-2024) after a subjective filter, plus 8,526 ThoughtTrace prompts (2026, with demographics). A donation tool: 52 US Prolific users uploaded ChatGPT or Claude histories, 140,834 messages from December 2022 to August 2026, returning labels only. A preregistered experiment (N=520) and a secondary analysis of a 3-week sycophancy study (1,015 people, 95,405 messages) probe causes.
Among subjective personal questions, belief tasks are over 90% oracle-framed. Action tasks split closer to even. WildChat is writing-heavy (Ghost-writer about 33%, Editor about 25%). Interpreter and Evaluator rise over time (both r=0.55, p=0.041); non-oracle Strategist falls (r=-0.72, p=0.004). ThoughtTrace is decision-heavy (Strategist about 50%, Decision-maker about 41%). Men exceed women (t=-3.343, p<0.001), people under 34 exceed older users (t=2.73, p<0.01), and weekly AI users exceed infrequent ones (t=2.37, p=0.02).
Donated logs are steeper. Oracle share among subjective questions rises from about 0.4 to 0.7 (r=0.80, p<0.001). Decision-maker and Interpreter increase (r=0.72 and 0.92); Strategist decreases (r=-0.63). Self-reports miss the mark: Decision-maker ranks mid-pack in recall but is 28% of observed roles; Deliberator is reported as second-most common and shows up under 1%. After the report, mean ratings are 5.94 accurate, 6.12 informative, 5.33 surprising on a 1-7 scale. 19 of 52 participants intend to offload less.
In the preregistered experiment, oracle messages are 20.0% to an "AI" partner versus 11.1% to a believed-human partner that is also an LLM (p<0.001). At least one oracle message: 65.4% versus 51.2% (OR=1.80). A one-SD warmer AI metaphor raises the odds of any oracle message by 55%. Under sycophantic models, oracle use rises across three weeks (r=0.62, p=0.03).
Interventions can change the model's side. A self-discovery goal prompt lowers oracular replies; a fast-answer goal raises them. Closed models infer "fast answer" on 49-58% of turns, open models on 0-8%. DPO matches prompting on the drop, and loses immediate reward: 200 crowdworkers prefer the original oracular replies for now, and prefer DPO replies when asked which output helps them form their own view.
| Observation | Number | Contrast |
| Donated oracle share | 0.4 to 0.7 | 2023-2026 subjective questions |
| AI vs believed human | 20.0% vs 11.1% | preregistered N=520 |
| Self-reported Deliberator | 2nd most common | observed <1% |
| Closed models pick "fast answer" | 49-58% | open models 0-8% |
Reliance research now has a scale that runs on open-ended personal use. Product work has two levers: fewer single-answer replies in the interface, and training objectives that do not only maximize immediate approval. On the user side, writing "help me think" versus "decide for me" already shifts the role mix. The donation reports show that much of this offloading is unnoticed; the discomfort arrives after the chart.
The paper is not a moral verdict. Oracle-style questions can be what someone wants in a crisis. The measurement is process (how the question is asked), not outcome (whether the user later obeys).
Crowdworkers and a "50+ conversations" filter skew toward heavy users, all US English. WildChat over-represents power users. N=52 cannot resolve individual traits. The preregistered "human" partner is an LLM in costume, so reply style is confounded with the label. The sycophancy evidence is secondary. Time trends mix changing users, changing models, and a changing user base. The judge assigns one dominant role per message and drops disclosure. The paper does not track whether answers actually change later beliefs or acts.