DSPy (Declarative Self-improving Python)
DSPy is a Python framework where you describe an AI job by naming its inputs and outputs, and it tunes the prompts for you.
Many AI apps are built on long prompts that people tweak by trial and error. DSPy, short for Declarative Self-improving Python, takes a different route. You describe the task in code, and DSPy turns that description into prompts. A part called an adapter writes the chat messages. It also reads the model’s reply back into the fields you asked for.
The task description is called a signature. It names the inputs, the outputs and a short instruction. Picture a help desk that must sort tickets. The signature says a ticket goes in, and an urgency of low or high comes out, plus a team name. A module, a ready-made building block, then decides how the model works on it. One module, ChainOfThought, adds step-by-step reasoning first.
To improve the program, you give DSPy some examples and a metric, a function that scores each answer. An optimizer runs the program many times and keeps the version that scores best. Depending on the optimizer, it rewrites the instructions, picks worked examples to show the model, or fine-tunes the model, meaning it adjusts the model’s own internal numbers (its weights). The tuned program is saved as a file you can reuse.
For engineers: DSPy began as a project of the Stanford NLP group and is released under the MIT license. The 2023 paper describes a compiler, a tool that tunes any DSPy pipeline (a chain of model calls) so it scores as high as possible on a chosen metric. Prompt-only optimizers leave the model untouched, so they also work with closed models, ones run by outside companies that you cannot change.
Hand-written prompts are hard to build on.
Follow one ticket-sorting task from code to a tuned program.
- 1 · declareYou write a signature that names the inputs, the outputs and a short instruction.
- 2 · chooseYou pick a module, such as ChainOfThought, which adds step-by-step reasoning.
- 3 · scoreYou supply example inputs and a metric, a function that scores each output.
- 4 · optimizeAn optimizer runs the program many times and keeps the version that scores best.
- 5 · saveThe tuned program is saved as a file you can load and run again.
In DSPy you describe the task and the score, and the prompts are tuned for you.
| Who | What they ask | What it works with |
|---|---|---|
| Support team | “Is this ticket urgent, and which team should take it?” | A signature with a ticket in and urgency and team out |
| Office assistant app | “What event is this email about, and when is it?” | A signature with an email in and an event name and date out |
| Newsroom | “Which claims does this article make, and does it back each one up?” | A module that lists claims, then gives a verdict on each |
| Machine learning team | “Can a smaller model do this job as well?” | Prompts moved from a larger model to a smaller one |
- You list what goes in and what comes out, rather than writing a long prompt string.
- The same task can switch strategy, such as adding reasoning or tools, without being rewritten.
- Optimizers improve the program against a score you define.
- In the 2023 paper, programs tuned by DSPy did better than prompts with a few hand-picked examples on the tasks tested.
- It cannot tell what good means on its own; you must write the metric.
- It does not pick an optimizer for you.
- A big optimizer run can cost hundreds of dollars in model calls.
- Examples chosen from a small training set may not help with very different inputs.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialDSPy, DSPy · read 28 Sept 2026
- repostanfordnlp/dspy: DSPy: The framework for programming, not prompting, language models, Stanford NLP on GitHub · read 28 Sept 2026
- paperDSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines, Khattab et al., arXiv · read 28 Sept 2026
- docsSignatures in depth, DSPy · read 28 Sept 2026
- docsOptimizers: choosing one, DSPy · read 28 Sept 2026
- docsBuilding metrics for evaluation and optimization, DSPy · read 28 Sept 2026
- docsAdapters: how signatures become prompts, DSPy · read 28 Sept 2026