Image editing models
An image editing model changes a picture you already have.
InstructPix2Pix takes an input image and a written instruction, then follows that instruction to edit the image. The authors train this conditional diffusion model on generated examples. At inference it generalizes to real images and user-written instructions. They base the editor on Stable Diffusion. The diffusion runs in the latent space of a pretrained variational autoencoder. That autoencoder has an encoder and a decoder.
Extra channels on the first convolutional layer concatenate the noisy latent with the encoded source. Weights on those new channels start at zero. The instruction reuses the text path that was built for captions. Two guidance scales trade off a match to the input image against a match to the instruction. A finetuned GPT-3 writes the instructions and the edited captions. The authors fine-tuned GPT-3 Davinci for one epoch with its default parameters. Prompt-to-Prompt keeps the paired generations similar. The resulting set has over 450,000 examples. The edit does not need per-example fine-tuning or inversion. It edits an image in a matter of seconds.
Imagen Editor is given an image, a masked area, and a text prompt, and it fills that masked area. It concatenates that image and mask with the diffusion latents. SDEdit adds noise to a guide drawn in RGB pixels, then denoises through a stochastic differential equation prior so the picture looks more realistic. It does not need task-specific training or inversion. Prompt-to-Prompt injects cross-attention maps during diffusion. Those maps tie the picture layout to each word in the prompt. Editing a real photo this way uses an inversion process. Emu Edit trains one model on sixteen editing tasks. A learned task embedding steers the model toward the right kind of edit. One model trained on all of the tasks beat a separate expert for each task.
A small change to the prompt often replaces the whole picture.
Follow one instruction through an InstructPix2Pix edit.
- 1 · encodeEncode the source picture in the latent space of a pretrained variational autoencoder.
- 2 · joinExtra channels on the first layer concatenate that source latent with the noisy latent.
- 3 · guideTwo guidance scales then trade off a match to the source image against a match to the instruction.
- 4 · returnThe edit does not need per-example fine-tuning or inversion.
Those scales can be adjusted so the sample matches the source, the instruction, or a mix of both.
| Who | What they ask | What it works with |
|---|---|---|
| Photo editor | “Can you change the jacket to leather and leave the face alone?” | The source photo plus a written instruction |
| Product designer | “Can you replace only the background behind the shoe?” | The product photo, a binary mask, and a prompt |
| Illustrator | “Can this painted guide become a realistic scene?” | A stroke painting and a chosen noise level |
| Benchmark team | “Did the edit keep the untouched parts while following the instruction?” | The image pair before and after the edit |
- A written instruction can edit the image without a full description of the output picture.
- The edit does not need per-example fine-tuning or inversion.
- A masked editor can fill a chosen region from the image, a binary mask, and a text prompt.
- SDEdit does not need a new training run or an inversion for each guide.
- InstructPix2Pix struggles with counting objects and with spatial reasoning.
- Edited images may inherit biases from the data and the pretrained models, or introduce other biases.
- In SDEdit, more noise and a longer denoising run make the image more realistic and less faithful to the guide.
- As a group, these inpainting models render objects better than text, and material, color, and size better than count or shape.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperInstructPix2Pix: Learning to Follow Image Editing Instructions, Brooks, Holynski, and Efros · read 28 Sept 2026
- paperSDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Meng and colleagues · read 28 Sept 2026
- paperImagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting, Wang and colleagues · read 28 Sept 2026
- paperPrompt-to-Prompt Image Editing with Cross Attention Control, Hertz and colleagues · read 28 Sept 2026
- paperEmu Edit: Precise Image Editing via Recognition and Generation Tasks, Sheynin and colleagues · read 28 Sept 2026