Most prompt frameworks hand you a string and wish you luck. DSPy, from Stanford NLP, hands you a compiler mindset instead: you declare what goes in and what comes out with typed Signatures, compose those calls into Python Modules, and then let optimizers rewrite the prompts and few-shot examples behind your back until a metric says the program got better. The name once stood for Demonstrate-Search-Predict, the research line that produced a string of papers from the original DSP in 2022 through the DSPy compilation paper at ICLR 2024, the MIPRO work on optimizing instructions and demonstrations, and GEPA in 2025, whose reflective prompt evolution claim is right there in the title: it can outperform reinforcement learning. The repository at stanfordnlp/dspy, MIT-licensed and currently at version 3.4.0, is the living code behind that research program, and its README summarizes the philosophy in one line: programming rather than prompting language models.
What makes the source worth a slow read is that every big idea in the papers has a small, findable file. The package is organized by concept rather than by layer: signatures, adapters, clients, primitives, predict, teleprompt, evaluate, retrieve, datasets, streaming, propose. There is even a dsp subpackage, the ancestral Demonstrate-Search-Predict layer, still riding along with its global Settings and its ColBERTv2 retriever client like an heirloom in the attic. Python 3.10 through 3.14 is supported, installation is one pip install away, and as always in this series what follows is an educational tour of published source code.
DSPy at a glance: Signatures and Modules define programs, the Predict family executes them through adapters and LiteLLM-backed client engines with caching and cost tracking, teleprompter optimizers from BootstrapFewShot to MIPROv2 and GEPA rewrite the programs against an evaluation harness, and streamify wires token streaming through the same path.
Reading the overview from left to right:
- The typed contract is the Signature class in dspy/signatures/signature.py, a Pydantic model with a custom metaclass that turns annotated fields into input and output schema.
- Programs subclass Module in dspy/primitives/module.py, whose metaclass intercepts instantiation so that predictors register themselves automatically.
- The workhorse is Predict in dspy/predict/predict.py, with ChainOfThought, ProgramOfThought, ReAct, BestOfN, and friends in the same folder.
- Translation between your Signature and the model’s chat format happens in the adapters, led by dspy/adapters/chat_adapter.py and the JSON variant in dspy/adapters/json_adapter.py.
- The LM layer centers on dspy.LM in dspy/clients/lm.py, which routes requests through the LiteLLM glue in dspy/clients/_litellm.py into the engine in dspy/clients/engines/litellm_engine.py, with caching in dspy/clients/cache.py and token accounting in dspy/clients/costs.py.
- The optimizers share the Teleprompter base in dspy/teleprompt/teleprompt.py, with MIPROv2 in dspy/teleprompt/mipro_optimizer_v2.py and GEPA in dspy/teleprompt/gepa/gepa.py.
- Scoring and streaming bookend the loop: the Evaluate harness in dspy/evaluate/evaluate.py and the streamify wrapper in dspy/streaming/streamify.py.
Why You Need This
The first reason is that DSPy solves the correct abstraction problem: prompts are not code, but they behave like compiled artifacts, and treating them that way changes how you iterate. A Signature is a class whose fields carry InputField or OutputField annotations from dspy/signatures/field.py; the signature metaclass reads those annotations, and the adapters render them into either a structured chat transcript or a JSON schema, with the JSONAdapter asking for strict structured output and the ChatAdapter parsing fenced sections. Because the schema lives in the type system rather than in prose, swapping the model behind a program does not rewrite your code, and changing what a step produces is a one-line annotation edit. The adapter folder goes further with typed multimodal fields under dspy/adapters/types, covering images, audio, documents, tool calls, and history, so a modern multimodal payload has a first-class place in a signature.
The second reason is the Predictor architecture, which is dependency injection for reasoning strategies. Predict itself is both a Module and a Parameter, so it appears in named_predictors listings that optimizers walk; its forward method preprocesses your arguments against the signature, calls the configured LM, and postprocesses completions back into a Prediction, which is itself an Example. Everything else in the predict folder is a variation on that loop: ChainOfThought adds a reasoning output field, ProgramOfThought generates Python and executes it through the sandboxed interpreter in dspy/primitives/python_interpreter.py, ReAct interleaves thought, action, and observation until a finish action, BestOfN samples N candidates with a reward function, and the aggregation module supplies majority-vote style merging. There are exotic corners too, like the Avatar module that runs a debate between agents, and the RLM module for recursive loops that lean on the interpreter. The lesson is that a reasoning strategy is just a Module with a different forward, which makes every strategy composable, subclassable, and optimizable by the same machinery.
The third reason is the optimizer stack, the part that justifies the word compiling in the paper title. Teleprompters take a program, a metric, and a trainset, and return a better program. BootstrapFewShot traces your program’s own executions to harvest demonstration examples, MIPROv2 combines bootstrapped demos with instruction proposals from the grounded_proposer module and Bayesian-style search over candidates, SIMBA does hill-climbing over bootstrap batches, and GEPA performs reflective prompt evolution in which an LM critiques failure modes and mutates instructions, the technique from the 2025 paper. Under the optimizers sits the client engine: requests flow through LiteLLM for provider breadth, responses are cached by dspy/clients/cache.py so repeated optimization trials do not re-bill, and costs.py accumulates token usage so you can watch an optimization run spend itself in real time. The evaluation harness runs the metric across dev sets in parallel, which closes the loop that makes the whole thing self-improving.
The detailed view: signatures and modules with their adapters, the full prediction module family, the LM client stack with engines, cache and costs, the teleprompter hierarchy with its proposers, and the evaluation and streaming support layer.
The detailed diagram rewards a slow pass. Notice how the predict group hangs off a single Predict node: ChainOfThought, ProgramOfThought, ReAct, BestOfN, and Avatar all extend or wrap it, while aggregation sits beside it as the combiner that BestOfN-style modules call. On the runtime side, dspy.LM and dspy.OpenAI both extend BaseLM from dspy/clients/base_lm.py, with the LiteLLM glue translating provider names like openai/gpt-4o-mini into actual calls, the engine handling retries and streaming, and the cache sitting before the network so identical requests never leave the machine twice. The optimize group is a clean hierarchy: BootstrapFewShot, MIPROv2, GEPA, and SIMBA all extend the Teleprompter base, MIPROv2 consumes instruction proposals from GroundedProposer, and both MIPROv2 and GEPA score candidates through the Evaluate harness. That symmetry is the architecture: one contract for programs, one for language models, one for optimizers, and metrics gluing them together.
The support layer rounds out the picture. streamify wraps any module so its output tokens stream to a listener as they are generated, which matters when a long ReAct trace would otherwise feel frozen. The retrievers folder ships retrieval model clients including ColBERTv2 and Weaviate, the datasets folder carries loaders for GSM8K, HotpotQA, and Math so the classic benchmarks are one import away, and the utils folder hides gems like asyncify for turning sync modules asynchronous, an MCP tool adapter, and a LangChain tool bridge. Even the legacy dsp layer with its global Settings still functions, a deliberate compatibility window for older programs.
From Install to an Optimized Program
The README path is short. Install, configure a model, define a signature, and predict:
pip install dspy
import dspy
lm = dspy.LM("openai/gpt-4o-mini")
dspy.configure(lm=lm)
class Summarize(dspy.Signature):
"""Summarize a document in one sentence."""
document: str = dspy.InputField()
summary: str = dspy.OutputField()
summarizer = dspy.Predict(Summarize)
result = summarizer(document="...long text...")
print(result.summary)
Upgrade the step to reasoning by changing one word, from dspy.Predict(Summarize) to dspy.ChainOfThought(Summarize), and the signature gains a rationale field automatically. Compose modules into a bigger program by subclassing dspy.Module, calling predictors inside forward, and returning them; the metaclass registers every predictor for you. Then optimize it with a metric and a small trainset, for example bootstrap few-shot examples with dspy.BootstrapFewShot, or run the heavier search with MIPROv2, which will propose instructions, bootstrap demos, evaluate candidates through your metric, and hand you back a compiled program whose prompts you never had to write. Every trial’s prompts and completions remain inspectable through module.inspect_history, which prints the exact messages the framework sent, a debugging affordance that most prompt frameworks promise and few actually deliver.
The honest closing is that DSPy is opinionated, and that is the product. If your application is one call to one model, the abstraction will feel heavy. The moment your system is five steps, two models, and a metric you care about, the compiler mindset pays for itself: you stop hand-tuning strings and start optimizing programs, with the papers’ tricks one import away and the whole research lineage readable in the tree. For teams that treat LLM systems as engineering rather than incantation, this repository is the reference implementation of the idea, and it compiles your prompts like it means it. Enjoyed this post? Never miss out on future posts by following us