AnythingLLM calls itself the all-in-one AI app you were looking for: chat with your docs, use AI agents, hyper-configurable, multi-user, and no frustrating setup required. The pitch at Mintplex-Labs/anything-llm, MIT-licensed by Mintplex Labs and currently at version 1.17.0, is a private, fully-featured ChatGPT that runs locally by default, connects to any of forty supported language model providers, ingests your documents into one of ten vector databases, and adds agents, memories, scheduled tasks, and an embeddable chat widget on top. The project ships a desktop app for Mac, Windows, and Linux, a Docker deployment with true multi-user support, a mobile app, and even a browser extension, all speaking to the same open-source core. Node 18 or newer is all the runtime the server and collector need, with a React 18 and Vite frontend.
The repository layout maps cleanly onto that story: server holds the Express API and the agent runtime, collector is a separate ingestion microservice, frontend is the chat UI, embed builds the website widget, and cloud-deployments plus docker carry the operations story. Two details in the server entry point betray production maturity before you read a line of chat code: the boot sequence patches SDK timeouts, warms a model-pricing cache, and supports HTTPS out of the box on port 3001. As always in this series, what follows is an educational tour of published source code.
AnythingLLM at a glance: the React frontend and embed widget hit the Express endpoint registry, which routes chats through the model router to forty provider adapters, retrieves context from ten vector databases, runs agents through aibitat, and persists everything in a thirty-model Prisma schema while the collector service turns documents into chunks.
Reading the overview from left to right:
- The front end is the React application under frontend/src, with a separate embeddable widget in embed.
- The API core is server/index.js, which mounts the endpoint modules in server/endpoints.
- Chat requests flow through server/endpoints/chat.js, with workspaces managed by server/endpoints/workspaces.js.
- Model routing rules live in server/endpoints/modelRouter.js and the engine in server/utils/AiProviders/modelRouter.
- The agent runtime is aibitat at server/utils/agents/aibitat/index.js, extended by plugins in server/utils/agents/aibitat/plugins.
- Retrieval backends are the ten providers under server/utils/vectorDbProviders, defaulting to LanceDB.
- State lives in the thirty-model Prisma schema at server/prisma/schema.prisma, and documents are parsed by the collector at collector/index.js.
Why You Need This
The first reason is the retrieval stack, because a private ChatGPT is only as good as what it remembers from your documents. The collector at collector/index.js is a standalone Express service with endpoints for processing single files, links, and raw text, which means heavy parsing never blocks the chat API. Inside it, collector/processSingleFile handles PDF, DOCX, TXT, EPUB, audio, and more, with an OCR loader in collector/utils/OCRLoader for scanned pages, and a link crawler in collector/processLink that turns websites into documents. Connector extensions in collector/extensions pull from external systems so ingestion is a workflow, not a drag-and-drop only affair. On the serving side, the text splitter in server/utils/TextSplitter, embedding engines in server/utils/EmbeddingEngines, and rerankers in server/utils/EmbeddingRerankers form a complete retrieval chain you can reconfigure per workspace.
The second reason is provider and vector freedom, which is the difference between a demo and a deployable product. The provider directory at server/utils/AiProviders holds forty adapters, from OpenAI, Anthropic, and Gemini to Ollama, LM Studio, LocalAI, vLLM-style local stacks, Bedrock, and a long tail of regional providers, each conforming to a common chat interface so the rest of the code never learns vendor specifics. The vector side mirrors it: ten backends under server/utils/vectorDbProviders including LanceDB for zero-setup local use, pgvector when Postgres is already in your stack, and Pinecone, Qdrant, Weaviate, Milvus, Chroma, Astra, and Zilliz for cloud scale. The model router in server/utils/AiProviders/modelRouter adds a layer few competitors have: rules that send each conversation to the best provider and model, with cooldowns handled by a dedicated plugin.
The third reason is the agent runtime, aibitat, which predates most agent frameworks and shows in its polish. The core loop at server/utils/agents/aibitat/index.js orchestrates multi-turn conversations between the model, the user, and callable skills, while the plugins folder at server/utils/agents/aibitat/plugins ships web browsing, web scraping, image generation, chart rendering, summarization, chat history, memory, and a classifier that selects skills intelligently to cut token usage on tool-heavy chats. MCP compatibility comes through the registry at server/endpoints/mcpServers.js, so any MCP server becomes another source of tools. Agent flows add no-code automation on top, and scheduled jobs let prompts run on cron with full agent capabilities.
The detail view: frontends on top, the Express API with middleware and endpoint modules below them, the aibitat runtime with plugins and MCP bridging in the middle, provider adapters and the model router on the right, the retrieval and Prisma storage layers beneath, and the collector flow at the bottom.
Walking the detail diagram, the API layer reveals how multi-user works. Middleware in server/middleware enforces authentication and role checks before endpoints like server/endpoints/admin.js manage users, invites, and API keys. The Prisma schema at server/prisma/schema.prisma is the data spine: users, workspaces, workspace_users, workspace_chats, workspace_threads, document_vectors, embed_configs, event_logs, scheduled_jobs, prompt_history, and more, thirty tables that make permissions, audit, and the embed widget first-class rather than bolted on. Chats are streamed over websockets for the UI and stored through the same schema, which is what allows the memory service in server/utils/memories to recall user and workspace facts across sessions.
The collector-to-vector flow is worth tracing end to end. A file upload lands on the collector, collector/processSingleFile extracts text and metadata, the OCR loader rescues scanned content, and the chunked output returns to the server, which embeds it via the configured engine and writes vectors and document records. At query time the chat endpoint retrieves top matches, optionally reranks them, and assembles the prompt with citations back to source documents. Background workers under server/jobs handle document sync queues so re-synced files stay fresh without user action.
From Install to a Private ChatGPT
The zero-config path is the desktop installer: pick Ollama or any cloud key, drag in a PDF, and start chatting with citations. The self-hosted path is Docker Compose, which brings up server, collector, frontend, and a LanceDB-backed data volume, and unlocks multi-user mode with role-based access. Developers who want the source path run the setup script, which wires environment files, generates the Prisma client, and starts server, collector, and frontend concurrently. Once running, the full developer API lets you script workspaces, documents, and chats from anything that can speak HTTP.
Honest limits: the all-in-one breadth means the codebase is large and mostly JavaScript with pragmatic rather than exhaustive typing, so tracing a feature across server, collector, and frontend takes patience. Multi-user and the embed widget are Docker-version features, desktop users trade them for simplicity. And because the project wraps so many providers, individual adapter behavior varies with each vendor’s quirks. But as a reference for what a complete, private, production-shaped LLM application looks like, this repository is the benchmark the all-in-one category is measured by. Enjoyed this post? Never miss out on future posts by following us