Most agent frameworks stop at the SDK and leave you to figure out serving, storage, and access control. Agno, Apache-2.0-licensed and currently at version 3.1.2, describes itself as the programming language for agentic software, and the README makes the scope explicit: build your agent platform using the Agno SDK, run it using the AgentOS runtime, and manage it through a web UI. In practice that means one repository contains the Agent and Team classes you write against, a FastAPI application that serves them with over fifty REST endpoints plus SSE and websockets, JWT-based role-based access control with multi-tenant isolation, a dozen database backends for sessions and memory, and the plumbing for cron scheduling, OpenTelemetry tracing, and channels like Slack, Telegram, WhatsApp, and A2A. Python 3.9 through 3.13 is supported, and the suggested onboarding is unusual for 2026: you paste a prompt into your coding agent, it clones a starter template, and Docker brings up a REST API, a Postgres database, an MCP server, and a control plane.
The repository is a monorepo with a disciplined split: the SDK lives in libs/agno, the command-line tooling in libs/agnoctl, infrastructure automation in libs/agno_infra, and hundreds of runnable examples in cookbook. Inside the SDK package, the top-level directory listing reads like a map of everything an agent platform needs: agent, team, workflow, models, tools, db, memory, knowledge, vectordb, run, scheduler, eval, guardrails, tracing, and more. The internal modules of the agent package use a single underscore prefix to mark privacy, which keeps the public surface small while the machinery stays inspectable. As always in this series, what follows is an educational tour of published source code.
Agno at a glance: the Agent, Team, and Workflow classes form the SDK core over model adapters and toolkits, the run layer carries events into the AgentOS FastAPI runtime, which guards access with JWT RBAC, exposes channels through interface adapters, and persists sessions, memory, and knowledge into pluggable storage.
Reading the overview from left to right:
- The primary abstraction is the Agent class at libs/agno/agno/agent/agent.py, thousands of lines with synchronous and asynchronous run entry points.
- Model calls flow through the abstract base at libs/agno/agno/models/base.py, with adapters for over fifty providers under libs/agno/agno/models.
- Actions come from the toolkit library in libs/agno/agno/tools/toolkit.py and roughly one hundred thirty toolkit modules in libs/agno/agno/tools.
- Multi-agent work happens in the Team class at libs/agno/agno/team/team.py and the workflow engine in libs/agno/agno/workflow/workflow.py.
- The server side is AgentOS itself, defined at libs/agno/agno/os/app.py as a FastAPI application.
- Authorization is enforced by the engine package under libs/agno/agno/os/authz, and channel adapters live in libs/agno/agno/os/interfaces.
- State lands in the storage backends of libs/agno/agno/db, memory in libs/agno/agno/memory/manager.py, and document grounding in libs/agno/agno/knowledge/knowledge.py.
Why You Need This
The first reason is the runtime, because demos are easy and serving is not. The AgentOS class at libs/agno/agno/os/app.py wraps your agents in a FastAPI application with lifespan hooks that initialize the database layer and the scheduler before the first request, manage MCP connection lifecycles, and mirror authorization state into the MCP app. The run layer in libs/agno/agno/run defines the event vocabulary for agent, team, and workflow runs, including background execution, cancellation, concurrency limits, and status persistence, so a long-running job survives a process restart. Streaming is first-class: event streams have in-memory and Redis backends under libs/agno/agno/os/event_streams/redis.py, and the same run events feed REST SSE, websockets, and the control-plane UI. This is the difference between an agent you can demo and an agent you can operate.
The second reason is security that starts at the front door instead of being bolted on. The authz package under libs/agno/agno/os/authz implements JWT-based role and scope enforcement with an engine, a native evaluation module, fine-grained authorization helpers, an admin router, audit logging, a user directory, and pluggable scope providers. Human approval is a runtime primitive, not an afterthought: libs/agno/agno/run/approval.py lets a run pause for user confirmation and lets you block tools that require admin sign-off. Combine that with multi-user, multi-tenant session isolation in the database layer and you get the pattern most teams end up reinventing badly when they move an agent prototype to production.
The third reason is the breadth of the integration surface, which is simply unmatched in a single library. The models directory ships over fifty provider adapters, from OpenAI, Anthropic, Gemini, and Ollama to vLLM, LiteLLM, Bedrock-style AWS setups, and dozens of regional providers, all conforming to the invoke and streaming interface of the Model base class in libs/agno/agno/models/base.py. The tools directory holds around two hundred Python files covering GitHub, Slack, SQL, Postgres, Pandas, DuckDB, shell execution, browser automation services, search APIs like Tavily and Exa, firecrawl scraping, E2B sandboxes, Twilio, email, and a native MCP bridge in libs/agno/agno/tools/mcp that turns any MCP server into callable tools. Storage flexibility matches it: eleven backends under libs/agno/agno/db including Postgres, MongoDB, MySQL, Redis, DynamoDB, Firestore, ClickHouse, and plain JSON.
The detail view: the Agent with its session, message, tool, and storage modules on the left, model adapters and the toolkit base beside it, Team and Workflow orchestration in the middle, the AgentOS runtime with router, MCP, interfaces, and RBAC on top, and storage, memory, knowledge, and observability below.
Walking the detail diagram, the Agent’s private modules show how a run actually proceeds. Session state is loaded and saved through libs/agno/agno/agent/_session.py, the prompt is assembled in libs/agno/agno/agent/_messages.py, model tool calls are resolved against registered toolkits in libs/agno/agno/agent/_tools.py, and persistence wiring lives in libs/agno/agno/agent/_storage.py. The memory manager at libs/agno/agno/memory/manager.py extracts and stores user memories across sessions, while the knowledge package supports filesystem sources and remote knowledge bases, backed by the vector store abstractions in libs/agno/agno/vectordb.
The orchestration tier is where Agno differentiates multi-agent design. Team supports coordinated member delegation with modes defined alongside the class in libs/agno/agno/team/team.py. The workflow engine at libs/agno/agno/workflow/workflow.py composes agents, teams, and functions from primitives in the same folder: steps in libs/agno/agno/workflow/step.py, parallel fan-out, loops, routers, and conditionals that branch on the previous step’s output, plus a CEL expression module for declarative conditions. Because every primitive accepts the same actor types, a workflow that starts as a linear chain can grow branches and parallel arms without changing how state flows.
The runtime tier ties it together, and the observability story completes the loop. The router module defines the fifty-plus REST routes AgentOS mounts, the MCP module serves the same agents over Model Context Protocol, and the interface adapters cover Slack, Telegram, WhatsApp, A2A, and AG-UI. Tracing is OpenTelemetry-based under libs/agno/agno/tracing, and the eval package runs structured evaluations against models, agents, and teams so behavior is measured rather than vibes-checked. Telemetry from the library itself is one event per run and switches off with an environment variable, a reasonable trade the README documents openly.
From Install to a Served Agent
The twenty-line path is real: pip or uv install agno, subclass nothing, instantiate Agent with a model and tools, call print_response or run, and the framework handles sessions, tool schemas, and streaming. The production path adds AgentOS around the same agents: point the runtime at your agent module, choose a database from the eleven backends, and deploy the container with the starter templates for Docker, Railway, AWS, GCP, Azure, Fly, Render, Modal, or Helm. The same deployment then serves REST clients, MCP clients, and chat channels from one process.
Honest limits: the SDK surface is enormous, and the framework does a lot of typing for you, which can feel heavy when you only wanted a single scripted agent loop; the monorepo’s cookbook helps but the docs are the real entry point. Some advanced runtime features assume Postgres for full fidelity, and the JWT RBAC layer, while thorough, is one more system to configure correctly. But if your destination is a multi-agent platform with real users, real data, and real access control, starting from Agno beats assembling the same parts from five libraries. Enjoyed this post? Never miss out on future posts by following us