Rude Prompts Cut Token Usage by 44% in LLMs, Penn Study Finds

rohanpaul_ai · x · 2026-08-03

A new paper from the University of Pennsylvania reveals that the tone of user prompts directly impacts the inference cost and accuracy of Large Language Models (LLMs). Researchers tested 7 tones, ranging from sycophantic to threatening, across multiple models using a fixed set of 570 MMLU questions.

Findings show that tone shifted accuracy by at most 2.99% but altered output-token usage by up to 44.3%. The optimal tone was highly model-specific. For ChatGPT-4o, a rude tone yielded the highest accuracy (89.04%) and the shortest responses (223 tokens on average). Conversely, Gemini 2.5 Flash Lite performed best with a neutral tone, while rude prompts caused accuracy to drop and token usage to increase by 35.2%. The study concludes that tone is a model-specific cost and reliability setting that production systems should standardize and benchmark.

Original post →

More from Models

Models channel →