Prompt Flow is Microsoft’s answer to a question every team eventually asks: what would it look like if LLM application development had the same engineering discipline as ordinary software? The README defines it as “a suite of development tools designed to streamline the end-to-end development cycle of LLM-based AI applications, from ideation, prototyping, testing, evaluation to production deployment and monitoring.” The repository at microsoft/promptflow is MIT-licensed and organized as a family of ten packages under src, from an umbrella promptflow metapackage at version 1.18.5 down to promptflow-core, promptflow-devkit, promptflow-tools, promptflow-tracing, and cloud-oriented siblings for Azure, evaluations, and RAG. This is not an agent framework and it does not want to be; it is the connective tissue that turns prompts, Python functions, and model calls into testable, traceable, deployable units.

The design idea that carries the codebase is the flow as a first-class artifact. A flow is a folder with a YAML definition, Python tools with typed inputs and outputs, and connections that hold secrets outside your code. When you run it, the FlowExecutor validates the definition against typed contracts, the DAG manager resolves node dependencies into a graph, node schedulers execute that graph with concurrency where the graph allows it, and every step emits tracing spans you can inspect locally or send to a collector. Batch runs apply the same flow to datasets through process pools, and evaluation is just another flow. Once you see that architecture, you understand why the project calls its unit a flow rather than an agent: the value is in deterministic structure, not autonomy.

As always in this series, this is an educational tour of published source code. The safety-relevant parts are also the most educational ones: connections exist precisely so API keys never live in flow code, the validator refuses malformed graphs before anything executes, and tracing makes every prompt and completion inspectable after the fact, which is what you want when a model misbehaves in production. Borrow all three habits even if you never install the framework.

Prompt Flow overview architecture diagram

Prompt Flow at a glance: the pf CLI drives the DevKit SDK, the promptflow-core runtime executes flows through FlowExecutor with a DAG manager and node schedulers, typed contracts and connections keep definitions honest, tracing spans every step, storage persists runs, and the Azure extension carries the suite into the cloud.

Reading the overview from left to right:

Why You Need This

The first reason is that Prompt Flow is the best open study of DAG-based LLM orchestration done rigorously. The executor package reads like a textbook: the validator checks the flow before execution, the DAG manager turns node references into a dependency graph, the flow nodes scheduler runs independent nodes concurrently and waits where it must, and the async variant handles coroutine-based nodes while the line execution process pool scales batch runs across processes. None of this is glamorous, and all of it is exactly what breaks when teams hand-roll orchestration. If your project has more than three chained LLM calls, reading these files will save you from writing your own worse version of them.

The second reason is evaluation as a native concept, not an afterthought. The same machinery that runs a flow once runs it across a dataset, with run info captured per line, and the dedicated promptflow-evals package builds evaluators that grade outputs against expectations. The README explicitly pushes integrating testing and evaluation into CI/CD, and the package structure makes that practical because an evaluation run is just a flow run with a different entry point. This is the discipline gap between demo and production, and Prompt Flow closes it with typed artifacts like run_info.py rather than dashboards alone.

The third reason is tracing that treats LLM debugging as a systems problem. The promptflow-tracing package is deliberately standalone so any Python function can carry spans, with a tracer that manages context, OpenAI-specific utilities that enrich spans with token usage and model details, and span enrichers for the rest. Combined with the serving layer under core/_serving, which wraps flows as REST endpoints, you get the full lifecycle the README promises: trace during development, evaluate across datasets, serve in production, and compare runs from storage afterward. Very few open projects cover the whole arc, and fewer do it with package boundaries this clean.

Prompt Flow detailed architecture diagram

The detailed view: entry points from CLI to orchestrator and proxy, the core runtime with FlowExecutor and its validator, graph execution through DAG, schedulers, and process pools, typed contracts for flows, tools, and runs, data and serving with connections, storage, and integrations, the tracing package with its tracer and OpenAI utilities, tool support including Prompty, and the cloud packages for Azure, evals, and RAG.

The detailed diagram rewards a slow pass. In the entry group, the orchestrator in src/promptflow-devkit/promptflow/_orchestrator drives batch jobs while the proxy layer connects local runs to external surfaces. In the graph group, src/promptflow-core/promptflow/executor/_async_nodes_scheduler.py handles coroutine nodes and src/promptflow-core/promptflow/executor/_line_execution_process_pool.py parallelizes dataset lines. In the tracing group, src/promptflow-tracing/promptflow/tracing/_tracer.py manages span context and src/promptflow-tracing/promptflow/tracing/_openai_utils.py packs OpenAI responses into traceable objects. The tools group includes a Prompty executor at src/promptflow-core/promptflow/executor/_prompty_executor.py for the .prompty format and an assistant tool invoker at src/promptflow-core/promptflow/executor/_assistant_tool_invoker.py for function-calling loops.

From Flow Folder to Production

The README quickstart is worth reproducing because it shows the tool-first philosophy:

pip install promptflow promptflow-tools
pf flow init --flow ./my_chatbot --type chat

and to iterate:

pf flow test --flow ./my_chatbot --interactive

What you get is a flow folder with a YAML definition and Python tools, which you edit like ordinary code and debug with an interactive session that shows each node’s inputs and outputs with tracing attached. The same flow then scales to batch runs over datasets, gets evaluated by evaluator flows, and deploys through the serving layer or the Azure extension, with the cloud version of Prompt flow in Azure AI for teams that want managed collaboration. The flow-recording package rounds the loop out by enabling recorded, reproducible runs against mocked model responses for tests.

Prompt Flow earns its place in this series because it represents the engineering-first pole of the LLM tooling spectrum: no role-playing metaphors, just typed contracts, a real scheduler, tracing, evaluation, and deployment, all split across packages with textbook boundaries. If agent frameworks are the exciting half of the field, this is the responsible half, and its executor package should be mandatory reading for anyone building LLM infrastructure.

Next up in this series: Flowise, the low-code builder that makes LLM orchestration visual. Until then, read the seams, not the slogans.

Watch PyShine on YouTube

Contents