BIG-Bench's 115% EmotionPrompt jump sits on an APE baseline of 2.39

Large Language Models Understand and Can be Enhanced by Emotional Stimuli

Cheng Li, Jindong Wang, Yixuan Zhang, Kaijie Zhu, Wenxin Hou, Jianxun Lian, Fang Luo, Qiang Yang, Xing Xie

IJCAI'23

cs.CL, cs.AI, cs.HC

2023-07-14

Eleven emotional lines lift six LLMs from 51.65 to 51.98 accuracy on Instruction Induction. The cited 8% is the best line per model; 106 raters score GPT-4 up 10.9%.

What problem this solves

Expectancy, confidence, and other people's opinions change what humans do. Classrooms and health campaigns already use encouraging wording for that reason. LLMs can follow instructions and handle some reasoning. What had not been measured is narrower: if those sentences are pasted onto the end of a prompt, do task scores move?

Earlier emotion work asked whether a model can recognize emotional words. It did not ask whether the words change answers. Zero-shot-CoT and automatically searched prompts are also uneven. The same added line can help a mid-size model and do little for GPT-4. An extra sentence costs almost nothing, so one averaged percentage does not say who actually moved.

Method

EmotionPrompt is concatenation. The original instruction stays, one stimulus goes at the end, and weights and demonstration format stay fixed. Eleven sentences come from three psychology sources and fall into two bins: what other people think, and self-encouragement.

Self-monitoring, adjusting behavior for the social setting, is EP01 to EP05. EP01 asks for a confidence score from 0 to 1. EP02 is 'This is very important to my career.' EP03 to EP05 use 'You'd better be sure.' and 'Are you sure?', and they tell the model to look again. Self-efficacy, from social cognitive theory, is the claim that people who believe they can do the task try harder. EP07 to EP11 put that into lines containing 'believe in your abilities', 'success', and 'take pride in'. Reappraisal, from cognitive emotion regulation, restates a difficulty as still workable and is folded into EP03 to EP05 and EP07. EP06 simply glues EP01, EP02, and EP03 together.

The models are Flan-T5-Large, Vicuna, Llama 2, BLOOM, ChatGPT (gpt-3.5-turbo-0613), and GPT-4. Temperature is 0.7 for ChatGPT, GPT-4, and Llama 2; the others keep their defaults. Deterministic tasks are 24 Instruction Induction items, scored by accuracy, and 21 BIG-Bench items, scored with a normalized metric where 100 is a human expert, 0 is chance, and worse than chance can be negative. Zero-shot appends the stimulus to the original prompt. Five-shot adds five random input-output pairs. Baselines are the human prompt, that prompt plus 'Let's think step by step', and prompts written by APE. BIG-Bench is zero-shot only, for lack of compute.

A separate study asked 106 people to rate GPT-4 on 30 open questions, vanilla against EmotionPrompt, from 1 to 5 on performance, truthfulness, and responsibility. Ten items come from TruthfulQA, 15 are written to draw out bias, and 5 are poems or summaries.

Results

The table is the mean of six models. The mean-of-11 column is what you get if you do not pick a sentence. The max column takes the highest score after all 11 have been run.

SettingOriginalZero-shot-CoTMean of 11Best stimulus
Instruction Induction, zero-shot accuracy51.6546.7451.9855.24
Instruction Induction, five-shot accuracy47.9746.1550.0252.40
BIG-Bench, human prompts, normalized10.1610.3710.6111.92
BIG-Bench, APE prompts2.393.445.207.47

The abstract's 8.00% is a specific calculation. On zero-shot Instruction Induction, each model uses its own best stimulus, the relative gain over that model's original prompt is computed, and those six relative gains are averaged. Absolute accuracy moves from 51.65 to 55.24. Averaging all 11 sentences adds 0.33.

Three models rise on that average column and three fall. Vicuna goes from 44.91 to 50.56, BLOOM from 33.46 to 35.95, ChatGPT from 75.20 to 76.85. Flan-T5-Large falls from 25.25 to 22.93, Llama 2 from 50.33 to 46.61, GPT-4 from 80.75 to 78.96.

Five-shot adds 2.05 points on the mean column, 47.97 to 50.02, against 0.33 in zero-shot. GPT-4 five-shot moves from 82.13 to 84.12 on average and 87.13 at the best stimulus. On the same table, Zero-shot-CoT knocks GPT-4 from 80.75 to 59.72, and from 82.13 to 67.62 in five-shot.

The abstract's 115% does not match the human-written BIG-Bench prompts. That row moves from 10.16 to 11.92. GPT-4 moves from 22.69 to 24.80 at best, far from the human-expert score of 100. The relative change near 115% is the APE row: 2.39 to a mean of 5.20 and a max of 7.47. ChatGPT scores 20.10 on the human prompt; APE drops it to 5.12, and the best emotional line brings it back to 18.00, still under the human prompt.

EP02 is the best single line on Instruction Induction, 6.06% above the worst line, and it gets worse on BIG-Bench. The harder tasks prefer EP06.

TruthfulQA has 817 questions. ChatGPT's original prompt scores 0.75 true and 0.53 informative; the mean of 11 stimuli is 0.80 / 0.68, while EP01 alone is 0.61 / 0.94. Vicuna-13B moves from 0.77 / 0.32 to a mean of 0.82 / 0.05. Its CoT run is 0.99 / 0.00. The paper's 19% and 12% match the average, across three models, of the gap from the original prompt to that model's best stimulus: about +0.12, +0.23, and +0.23 on truthfulness, and about +0.41, -0.10, and +0.06 on informativeness. Vicuna's best informative score is 0.22, still below 0.32. GPT-4 was not run.

The 106 raters are summarized as a 10.9% average improvement across the three scales. Two of 30 questions get worse. On performance, nearly a third of items move by about 1 point or more on the 1 to 5 scale. In one miss, EmotionPrompt says 'completely' and 'will not' where the vanilla answer says 'generally' and 'may even be'. In the other, the vanilla answer closes with a summary and the emotional answer only lists points.

Gradient-norm attribution was computed only on Flan-T5-Large. On 8 tasks, the words confidence, sure, success, and achievement account for more than half the contribution on 4 tasks and close to 70% on 2. From temperature 0 to 1.5 the gap between the two curves widens, and the EmotionPrompt curve is flatter. Vicuna and Llama 2 have no valid run at temperature 0.

Why it matters

You paste one sentence. No training. Two uses are credible: the prompt is already fixed and a validation set might still buy a little accuracy, or temperature is high and the vanilla wording starts to wander. The temperature plots support the second, because the emotional line moves less as temperature changes.

It is a poor default suffix. EP02 leads only on the easier tasks. Once EP01 plus EP04 is already working, adding EP06 through EP09 stalls or lowers the score. On zero-shot GPT-4 the mean of the 11 lines sits below the original prompt, and the best line moves 80.75 to 81.60, a gain of 0.85. Pick the sentence on your own items.

The same line also lengthens answers and hardens the tone. The failed human-study cases are traced to the career-stakes sentence and to 'You'd better be sure.'

Limitations

The max column is chosen after seeing all 11 sentences. On the zero-shot mean column, half the models go down, while the paper still describes a consistent gain on every model. On human-prompt BIG-Bench, Flan-T5-Large's best score is 4.00 against an original 4.66, so the max column can lose too.

Table 6 lists Llama 2 at a gain of 6.00 and BLOOM at 0.51. Subtract Table 1 and 6.00 is BLOOM, 33.46 to 39.46, while 0.51 is Llama 2, 50.33 to 50.84. The tables disagree, so RLHF cannot explain the gap inside this pair of 13B models. Larger models do not clearly gain more either: Vicuna gains 9.58 points and GPT-4 gains 0.85.

All 106 raters hold a bachelor's degree, and 90% are 20 to 25. Variance is high, and the paper puts that on subjective scales. Thirty questions, GPT-4 only, three metrics folded into 10.9%, with no separate means printed in the text.

Input attribution was run on one 780M encoder-decoder, on sentiment classification. The conclusion states the clash with the human literature: emotion can sway attitudes, and an encouraging sentence does not simply raise human reasoning. A large weight on positive words shows that those words take gradient mass. It is not evidence of emotional understanding. EP01 asks for a confidence score, which is an extra instruction.

Terms

Source

What people are saying

Related papers

All paper explainers