Revisiting APE (2022): the paper where LLMs became human-level prompt engineers
cephaloform · x · 2026-09-27
The author revisits the pre-chatbot paper 'Large Language Models Are Human-Level Prompt Engineers' (arXiv:2211.01910), calling it an inspiration resembling today's synthetic data generation factories built on base models.
Key points (Automatic Prompt Engineer, APE):
- Treats the instruction as a 'program': an LLM proposes candidate instructions, optimized via search against a score function
- Instruction quality is evaluated by another LLM's zero-shot performance following it
- Across 24 NLP tasks, auto-generated instructions beat the prior LLM baseline by a large margin and match or exceed human-written ones on 19/24 tasks
- APE prompts also improve truthfulness/informativeness and few-shot in-context learning.
More from Research
- Colored Noise Sampling: dyeing injected noise to boost diffusion sample quality — serrjoa · 2026-09-27
- Researcher: within months, submitting proofs without AI verification will be malpractice — RexDouglass · 2026-09-27
- JevBench evaluator: most models lose to fixable settings and calibration — airesearch12 · 2026-09-27
- "Mathematics is Effectively Dead": essay extends Daniel Litt on AI and the future of math — burny_tech · 2026-09-27
- ForecastingCo's first paper shows how to train transformers on real temporal data — fpedregosa · 2026-09-27
- LLMs crack AES keys from just 12 power traces in first systematic side-channel study — chaumian · 2026-09-27