Frameworks for LLM applications tend to fail in one of two directions: too much magic, so nothing can be debugged, or too little structure, so every project reinvents plumbing. Haystack, the Apache-2.0 orchestration framework from deepset, has spent years threading the needle, and version 3 — 3.4.0 at the time of writing, installable as pip install haystack-ai — makes the pitch explicit: design modular graphs and agent workflows with explicit control over retrieval, routing, memory, and generation. The README backs the claim with substance: agents built for production with lifecycle hooks like before_llm, before_tool, and on_exit, token usage and step counts tracked out of the box, one graph object that runs synchronously or asynchronously and streams token by token, and a model-agnostic component library spanning OpenAI, Mistral, Anthropic, Cohere, Hugging Face, Azure, Bedrock, and local models. The roster of organizations using it — Apple, Meta, Netflix, Airbus, the European Commission among them — suggests the production claim is not decorative.

Under the hood the repository is one Python package with a clean center of gravity: the core directory holds the graph engine and the component contract, components holds the ready-made library from converters to generators, tools holds a function-calling system complete enough that agents, graphs, and components can all be exposed as tools, and dataclasses plus tracing plus evaluation carry the runtime payloads and observability. An earlier post in this series looked at Haystack from a distance; this tour goes into the machinery. As always in this series, what follows is an educational tour of published source code.

Haystack overview architecture diagram

Haystack at a glance: the component contract feeds the graph engine, which runs both static workflows and the autonomous Agent; component families cover retrieval, generation, conversion, and writing; documents flow into pluggable stores, dataclasses define the runtime payloads, and tools plus YAML serialization round out the developer surface.

Reading the overview from left to right:

Why You Need This

The first reason is the graph engine itself, which is where Haystack’s years of iteration show. Components are plain Python classes annotated with the decorator from haystack/core/component/component.py; their inputs and outputs become typed sockets defined in haystack/core/component/sockets.py, and connecting them builds a directed graph that the engine source validates at connect time — type mismatches and missing inputs are errors before anything runs. The same graph object executes synchronously, asynchronously, and with per-component streaming, and the topological run logic lives in the shared run core. Breakpoint support lets you pause a run mid-graph for debugging, and the visualization helper renders the graph as Mermaid or ASCII. Explicitness here is a feature: you can always see the graph you built.

The second reason is the Agent, which turns the framework from a static DAG into something autonomous without abandoning the type discipline. The class at haystack/components/agents/agent.py runs a tool-calling loop over any chat generator, parsing responses through the strategies in haystack/components/agents/tool_calling.py and carrying working memory in the state object from haystack/components/agents/state/state.py. The README’s production features map directly onto source: lifecycle hooks fire around LLM calls and tool executions, step counts and token usage are tracked via the counters in haystack/token_counters, and the budget machinery in haystack/hooks/budget enforces limits so an agent cannot run away. Progressive skill discovery — skills described only when needed — comes from the skills system under haystack/tools/skills backed by the filesystem loader in haystack/skill_stores/file_system.

The third reason is a component and tool ecosystem designed for composition rather than vendor capture. The retrieval families in haystack/components/retrievers — including the BM25 retriever at haystack/components/retrievers/in_memory/bm25_retriever.py — query any store behind the protocol in haystack/document_stores/in_memory/document_store.py, while deeper integrations with Qdrant, Weaviate, OpenSearch, and friends live in the separate haystack-core-integrations repository, keeping the core dependency-light. The tool system is symmetrical in a way few frameworks manage: functions become tools via haystack/tools/from_function.py, whole graphs become tools through the graph-as-tool adapter, components through haystack/tools/component_tool.py, and other agents through haystack/tools/agent_tool.py — so an agent’s toolbox can contain the same abstractions you compose by hand.

Haystack detail architecture diagram

The detail view: the core graph engine with sockets, breakpoints, drawing, and super components at the top, the Agent and its tool-calling and state machinery with retrieval and generation components in the middle, the document store below, the tool and skills column on the right, and dataclasses, serialization, evaluation, and tracing at the bottom.

Walking the detail diagram through the engine layer, everything begins with registration. The decorator in haystack/core/component/component.py captures a component’s input and output sockets from its run method’s signature, the graph source checks type compatibility on connect, and the SuperComponent wrapper in haystack/core/super_component/super_component.py lets an entire graph masquerade as a single component — the abstraction behind ready-made patterns like the Agent Pack the README links. Serialization is a first-class concern: the YAML marshaller in haystack/marshal/yaml.py round-trips graphs to files, with callable serialization handled in haystack/utils/callable_serialization.py.

The component tier shows the framework’s economics in practice. A working RAG indexing branch is three nodes: the PDF converter at haystack/components/converters/pypdf.py turns bytes into Documents, an embedder turns them into vectors, and the writer at haystack/components/writers/document_writer.py persists them into the in-memory store at haystack/document_stores/in_memory/document_store.py, which also powers the BM25 retriever at haystack/components/retrievers/in_memory/bm25_retriever.py for keyword search. On the generation side, the chat generator at haystack/components/generators/chat/openai.py consumes ChatMessage objects and produces streaming chunks defined in haystack/dataclasses/streaming_chunk.py, while routers like haystack/components/routers/conditional_router.py and joiners like haystack/components/joiners/document_joiner.py provide the branching and merging vocabulary.

The Agent tier and the observability floor complete the picture. The loop in haystack/components/agents/agent.py alternates generation and tool execution, dispatching through the toolset abstraction at haystack/tools/toolset.py and validating schemas from haystack/tools/tool.py, while every step is traced through the hooks in haystack/tracing/tracer.py — pluggable to OpenTelemetry — and evaluation results accumulate in haystack/evaluation/eval_run_result.py. The core dataclass for content is the Document in haystack/dataclasses/document.py, and conversation history is the ChatMessage of haystack/dataclasses/chat_message.py. Serving the result as an API or MCP server is deliberately out of scope here — that is the companion Hayhooks project the README points to, which also exposes OpenAI-compatible endpoints that chat UIs can consume.

From Install to a Running Graph

The install is one line — pip install haystack-ai — and the shape of a program is: build components, connect them with » and »= into a graph, call run with your question, and stream the answer. The same graph can be saved to YAML and reloaded elsewhere, which is how workflows travel between notebooks, services, and the Hayhooks deployment story. Moving from static RAG to agency is additive: pass a toolset to an Agent, optionally set budgets and lifecycle hooks, and the agent plans its own tool calls against the same components. Evaluation is run-based rather than vibes-based, metrics are attached to the components you already use, and the tracing hooks export spans to the observability stack you already have. Nightly pre-releases arrive through the –pre flag for anyone tracking the fast-moving edge.

Honest limits: the core is deliberately spare, so real deployments almost always pull in the companion integrations repository for vector stores and providers, which means two repos and their release cadences to track. The typed-socket discipline, while the framework’s greatest strength, takes real adjustment coming from prompt-string frameworks, and graph YAML is powerful enough that mistakes in it surface at connect time — strict, but strict like a compiler. The in-memory document store is a reference implementation rather than a production database. But as a study in how to build an orchestration framework that stays debuggable — typed sockets, explicit graphs, tools that compose, agents that are just components — deepset-ai/haystack remains one of the most instructive codebases in this series.

Watch PyShine on YouTube

Contents