Zhou et al. (2023) — Large Language Models Are Human-Level Prompt Engineers
Zhou, Y., Muresanu, A.I., Han, Z., Paster, K., Pitis, S., Chan, H. & Ba, J. (2023). Large Language Models Are Human-Level Prompt Engineers. ICLR 2023. https://arxiv.org/abs/2211.01910
What this source contributes
The paper proposes Automatic Prompt Engineer (APE): the instruction given to a model is treated as a program, the model itself proposes candidate instructions, and a search selects the one that scores best on the task. Tested on 24 natural-language tasks, the automatically generated instructions outperformed the earlier LLM baseline and matched or beat instructions written by human annotators on 19 of the 24. Prepended to ordinary prompts, APE-found instructions also improved truthfulness and few-shot performance.
The claim in the title is the finding: for the tasks measured, a model writes prompts as well as a person does. What was, in 2022, the craft the term prompt engineer named — and the skill the AI-ninja identity is built on — was shown to be automatable by the system it is applied to.
Analytical function in AI-Ninja
This is the evidence under the entry’s central claim, that the ninja’s mastery consists of a skill the platform is being engineered to require less of. Without it the claim is a forecast; with it, it is a measurement made the year the term appeared. The later, weaker Khan (2025) preprint in the secondary bundle extends it to newer model generations.
About the author(s)
University of Toronto and the Vector Institute; Jimmy Ba is the senior author. Peer-reviewed as a conference paper at ICLR 2023.
Related entries
- AI-Ninja — the skill at the core of the identity, shown to be automatable
- Prompt Engineer — the profession named after that skill
- Deskilling — the mechanism, applied to AI proficiency itself