Building with AI

Prompt engineering

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Prompt engineering is writing and testing the instructions and examples you give a model so its answer meets a goal you can check.

1 · What it is

Prompt engineering is a loop. You change the words a model reads, run it, and check the result against a goal you wrote down first. So you need three things before you start: a clear picture of success, a way to test for it, and a first draft to improve.

Why can text alone change what a model does? The GPT-3 paper showed that a large model can pick up a task from a plain description, or from a handful of worked examples placed in its input. Nothing inside the model is retrained. With zero examples it gets only the instruction. With a few, it sees demonstrations at the moment it answers. Some tests stayed hard even so.

A few habits help. Say exactly what you want instead of hoping the model guesses. Show a few examples when format and tone matter. Try the colleague test: if a smart person with no background would be confused by your prompt, the model probably will be too. With Claude, XML tags help keep instructions, context and examples apart.

A survey of the field sums up the job as finding the prompt that best lets a model solve the task. The other route is to change the model itself. The InstructGPT authors noted that a bigger model does not automatically follow intent better. So they fine-tuned GPT-3 on demonstrations written by human labelers, then trained it again with reinforcement learning from human feedback. Prompt engineering skips all that training and changes only what you send.

2 · Why it exists

The wording of the request is part of the task, and a vague request leaves the result underspecified.

Unstated standardsIf you want a specific result, you have to ask for it. A short prompt does not make the model infer your standard.
Wandering formatFormat, tone and structure move around unless examples show the pattern you want.
Wrong repairA weak result is not always a prompt problem. Latency and cost can be easier to improve by choosing a different model.
3 · How it works

Change the text, then score it against a result you defined.

The examples sit in the prompt. They do not update the model's weights.
  1. 1 · defineWrite what a good answer is, and how you will test answers against that description.
  2. 2 · instructState the output format, the constraints and any ordered steps in the prompt itself.
  3. 3 · showAdd a few examples that match the real task, including edge cases.
  4. 4 · checkScore the outputs and revise the wording when they miss the criteria.

Examples and instructions change the input. They do not update the model's weights.

InstructGPT came from two training rounds: demonstrations, then RL from human feedback.
4 · Where it's used
WhoWhat they askWhat it works with
Support editor“Answer in our voice, and refuse requests we do not cover.”Role, policy steps, three sample replies
Data analyst“Return the comparison as a table with these columns.”Format rules and one finished table
Teacher“Explain this to a beginner, then give one practice question.”A worked explanation used as the example
Engineer“Which prompt wording passes the eval set?”Success criteria and scored outputs
5 · What it solves, and what it doesn't
solves
  • A model can attempt a new task from text alone, with no gradient update and no fine-tuning.
  • A few examples can steer format, tone and structure toward a pattern you show.
  • A role sentence in the system prompt focuses behaviour and tone for that job.
  • You can test a wording change against criteria you wrote, without training a new model.
doesn't solve
  • Few-shot results still struggle on some tasks. The GPT-3 paper names ANLI, RACE and QuAC.
  • Prompting does not change weights. The few-shot setting allows no weight updates.
  • Not every failed evaluation is best fixed by editing the prompt.
  • Examples that are too alike can teach a pattern you did not mean to teach.