Run one container and you get a ChatGPT-shaped interface pointed at models you control: that is the promise that made Open WebUI one of the most starred projects on GitHub. The README calls it a home for AI β€” a self-hosted, extensible, feature-rich platform that runs entirely offline, with first-class support for Ollama alongside any OpenAI-compatible API, whether that is LMStudio, GroqCloud, Mistral, OpenRouter, or vLLM. Install with pip and serve on localhost:8080, or use the Docker images tagged for Ollama and CUDA. The feature list reads like a product roadmap for an entire company: granular RBAC with user groups, plugins of five kinds connected through MCP and OpenAPI tool servers, model presets that become agents, agentic terminal execution, notes and channels, persistent memory, calendars and automations, voice and video calls, local RAG across nine vector databases with hybrid BM25 and vector search, web search from dozens of providers, image generation, usage analytics with ELO-based model evaluation, and enterprise identity through LDAP, OAuth SSO, and SCIM 2.0 provisioning.

The source backs that breadth with a disciplined split: a SvelteKit frontend under src and a FastAPI backend under backend, where the package open_webui organizes routers, peewee data models, the retrieval stack, socket machinery, and a plugin loader, all assembled in one FastAPI application. One earlier post in this series covered this project from a deployment angle; this tour goes into the source itself. As always in this series, what follows is an educational tour of published source code.

Open WebUI overview architecture diagram

Open WebUI at a glance: the SvelteKit app and chat component drive client stores, calls land on the FastAPI app which verifies sessions and mounts the Ollama and OpenAI provider routers, the chat middleware orchestrates plugins and RAG, and data models, the WebSocket hub, and pluggable file storage sit underneath.

Reading the overview from left to right:

Why You Need This

The first reason is provider freedom with one interface. The two routers in backend/open_webui/routers/ollama.py and backend/open_webui/routers/openai.py proxy local Ollama instances and any number of OpenAI-compatible endpoints, so switching from a local Llama model to a cloud frontier model is a settings change, not a migration. The middleware at backend/open_webui/utils/middleware.py normalizes what happens around the model call β€” system prompts from model presets, persistent memories retrieved from backend/open_webui/models/memories.py, RAG context, tool calls, filters β€” so the model behind the request can change without the surrounding behavior changing. Multi-model conversations, where several models answer side by side, fall out of the same orchestration.

The second reason is extensibility as a core primitive rather than an SDK you must adopt. The loader at backend/open_webui/utils/plugin.py dynamically imports user-authored Python β€” filters that edit requests and responses, pipes that become models themselves, tools that extend function calling, and actions that add chat buttons β€” while the MCP client at backend/open_webui/utils/mcp/client.py connects external tool servers described in the README alongside MCPO and OpenAPI servers. Permissions for all of it flow through the access control module in backend/open_webui/utils/access_control, which evaluates group-based rules behind the roles the admin router manages. For teams, the Open Terminal router at backend/open_webui/routers/terminals.py exposes the agentic execution environment where models run real commands on real files.

The third reason is the RAG stack, which is unusually complete for a project whose headline is the chat box. Document ingestion runs through the loaders in backend/open_webui/retrieval/loaders with extraction engines including Tika and Docling per the README; vector storage is abstracted by backend/open_webui/retrieval/vector/main.py over nine backends in backend/open_webui/retrieval/vector/dbs β€” ChromaDB, PGVector, Qdrant, Milvus, Elasticsearch, OpenSearch, Pinecone, S3Vector, and Oracle 23ai; hybrid search combines BM25 with vector similarity and optional reranking. Web search adds dozens of external providers through the registry utilities in backend/open_webui/retrieval/web/utils.py. Knowledge bases built on this stack are managed at backend/open_webui/routers/knowledge.py and attached to models or referenced inline in chat.

Open WebUI detail architecture diagram

The detail view: the SvelteKit page, chat component, stores, and API client on the left, the FastAPI assembly with its routers and socket hub, the chat and models tier of middleware, provider proxies, plugin loader, and access control, the persistence layer of peewee models and storage provider, and the RAG and tools column on the right.

Walking the detail diagram from the frontend, the chat component at src/lib/components/chat/Chat.svelte is the most complex file in the UI: it manages sending, streaming, message queues, tool call rendering, and files, coordinating through the stores in src/lib/stores/index.ts and the typed fetchers of src/lib/apis/chats/index.ts. History round-trips through the chat CRUD endpoints on backend/open_webui/main.py, and the WebSocket connection registered by backend/open_webui/socket/main.py carries live events β€” streaming updates, typing state, usage notifications β€” which is also what enables Redis-backed scaling to multiple workers per the README.

The backend assembly in backend/open_webui/main.py mounts dozens of routers, and the tour only highlights a few: authentication and identity in backend/open_webui/routers/auths.py, model presets that wrap base models into agents at backend/open_webui/routers/models.py, knowledge base CRUD at backend/open_webui/routers/knowledge.py, speech in backend/open_webui/routers/audio with the realtime call handling, image generation at backend/open_webui/routers/images.py covering ComfyUI and AUTOMATIC1111 locally plus cloud engines, and terminals as above. Configuration is centralized in backend/open_webui/config.py, which reads environment variables and persists admin changes back to the database.

The middleware is where a chat completion becomes an orchestrated flow. It loads the model preset, merges system prompts and memories, consults access control in backend/open_webui/utils/access_control, runs inlet filters from the plugin machinery in backend/open_webui/utils/plugin.py, builds RAG context through backend/open_webui/retrieval/utils.py, dispatches the completion to the Ollama or OpenAI proxy, then processes tool calls β€” including those served by the MCP client at backend/open_webui/utils/mcp/client.py or the sandboxed interpreter at backend/open_webui/utils/code_interpreter.py β€” and finally persists the chat via the peewee models in backend/open_webui/models with file attachments handled by the storage provider at backend/open_webui/storage/provider.py. Migrations in backend/open_webui/migrations evolve the schema across releases for SQLite and PostgreSQL alike.

From Install to a Multi-Provider Chat

The whole platform is two commands deep: pip install open-webui, then open-webui serve, and the interface is on localhost:8080 with the first registered account becoming an administrator. Connect Ollama or point an OpenAI-compatible connection at any provider, upload documents to build a knowledge base, and attach that knowledge to a model preset together with custom instructions and tools to create an agent. The chat itself supports markdown, LaTeX, voice input and calls, and the hash command for pulling documents or URLs into context. For teams, switch the database to PostgreSQL, add Redis for session sharing, put several workers behind a load balancer, and manage users through groups, LDAP, or SCIM. The same deployment then serves the desktop app and companion projects described in the README’s ecosystem section.

Honest limits: the license is the project’s own, a BSD-style text in LICENSE with a branding clause that prohibits removing Open WebUI identity above fifty users without an enterprise arrangement β€” history lives in LICENSE_HISTORY. The codebase is large and feature-dense, with a frontend chat component that is genuinely hard to modify casually, and the feature surface means more configuration than a minimal chat UI demands. But if the goal is a self-hosted, provider-agnostic AI workspace that a whole organization can actually use β€” with RAG, tools, identity, and governance included β€” the source here is the most complete open reference available.

Watch PyShine on YouTube

Contents