Benchmarking LLMs by Their Ability to Prompt-Engineer GPT-2

pokeuser61 · reddit · 2026-08-09

The author proposes a novel proxy for evaluating model intelligence: having various LLMs write a single prompt template for GPT-2. GPT-2, combined with the generated template, is then scored across 395 examples of a basic task (generating a short report and deciding the correct action for a farm).

While admittedly limited in usefulness as a traditional benchmark, the author notes that this mechanism reveals interesting points and differences in model capabilities regarding prompt engineering and instruction following.

Original post →

More from Research

Research channel →