Prompt engineering
Prompt engineering is writing and testing the instructions and examples you give a model so its answer meets a goal you can check.
Prompt engineering is a loop. You change the words a model reads, run it, and check the result against a goal you wrote down first. So you need three things before you start: a clear picture of success, a way to test for it, and a first draft to improve.
Why can text alone change what a model does? The GPT-3 paper showed that a large model can pick up a task from a plain description, or from a handful of worked examples placed in its input. Nothing inside the model is retrained. With zero examples it gets only the instruction. With a few, it sees demonstrations at the moment it answers. Some tests stayed hard even so.
A few habits help. Say exactly what you want instead of hoping the model guesses. Show a few examples when format and tone matter. Try the colleague test: if a smart person with no background would be confused by your prompt, the model probably will be too. With Claude, XML tags help keep instructions, context and examples apart.
A survey of the field sums up the job as finding the prompt that best lets a model solve the task. The other route is to change the model itself. The InstructGPT authors noted that a bigger model does not automatically follow intent better. So they fine-tuned GPT-3 on demonstrations written by human labelers, then trained it again with reinforcement learning from human feedback. Prompt engineering skips all that training and changes only what you send.
The wording of the request is part of the task, and a vague request leaves the result underspecified.
Change the text, then score it against a result you defined.
- 1 · defineWrite what a good answer is, and how you will test answers against that description.
- 2 · instructState the output format, the constraints and any ordered steps in the prompt itself.
- 3 · showAdd a few examples that match the real task, including edge cases.
- 4 · checkScore the outputs and revise the wording when they miss the criteria.
Examples and instructions change the input. They do not update the model's weights.
| Who | What they ask | What it works with |
|---|---|---|
| Support editor | “Answer in our voice, and refuse requests we do not cover.” | Role, policy steps, three sample replies |
| Data analyst | “Return the comparison as a table with these columns.” | Format rules and one finished table |
| Teacher | “Explain this to a beginner, then give one practice question.” | A worked explanation used as the example |
| Engineer | “Which prompt wording passes the eval set?” | Success criteria and scored outputs |
- A model can attempt a new task from text alone, with no gradient update and no fine-tuning.
- A few examples can steer format, tone and structure toward a pattern you show.
- A role sentence in the system prompt focuses behaviour and tone for that job.
- You can test a wording change against criteria you wrote, without training a new model.
- Few-shot results still struggle on some tasks. The GPT-3 paper names ANLI, RACE and QuAC.
- Prompting does not change weights. The few-shot setting allows no weight updates.
- Not every failed evaluation is best fixed by editing the prompt.
- Examples that are too alike can teach a pattern you did not mean to teach.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsClaude prompting best practices, Anthropic · read 27 Sept 2026
- docsPrompt engineering overview, Anthropic · read 27 Sept 2026
- paperLanguage Models are Few-Shot Learners, Brown et al. · read 27 Sept 2026
- paperPre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Liu et al. · read 27 Sept 2026
- paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 27 Sept 2026