Ollama is the tool that made running open models a one-line command, and it is now a standard piece of developer infrastructure: install it with curl -fsSL https://ollama.com/install.sh | sh on macOS or Linux, irm https://ollama.com/install.ps1 | iex on Windows, or pull the official ollama/ollama Docker image, and you have a local model server on localhost. The repository is MIT-licensed Go, small enough to read in a few sittings and organized with unusual clarity for something that spans a CLI, an HTTP API, subprocess management, and CUDA-level GPU discovery. This tour walks through how a ollama run becomes tokens on your screen.
What makes the codebase worth studying is how much real engineering hides behind the friendly surface. There is a scheduler that decides which model stays loaded in VRAM and which gets evicted, a runner manager that spawns and babysits llama.cpp server processes, a model format built on the same manifest and layer ideas as container registries, and compatibility layers that let tools written for the OpenAI or Anthropic APIs talk to your local models without changes. That last part matters right now: the README documents launching coding agents such as Claude Code, Codex, Copilot CLI, and OpenCode against local models with ollama launch claude.
As always in this series, this is an educational tour of published source code. Ollama runs powerful models locally and exposes them over a network API, and its design assumes you will use those powers deliberately: models pull from the ollama.com registry with verified signatures, the server binds to localhost by default, and tool calls execute with the same care you would want from any agent runtime. Study the guardrails alongside the features, because local inference is only as safe as the server that hosts it.
Ollama at a glance: CLI, desktop app, and API clients meet one gin HTTP server that exposes native, OpenAI-compatible, and Anthropic-compatible endpoints, backed by a GPU-aware scheduler, llama.cpp runner processes, and a manifest-based model store.
Reading the overview from left to right:
- The binary starts at main.go, with the command set registered in cmd/cmd.go.
- The macOS desktop application lives under app, embedding the same server the CLI uses.
- Every client speaks to the HTTP server at server/routes.go, built on the gin framework.
- Compatibility endpoints for OpenAI-style clients are implemented in openai/openai.go, and Anthropic-style clients are served through anthropic/anthropic.go.
- Inference is queued through the scheduler at server/sched.go, which manages model lifecycle on your hardware.
- The scheduler spawns runner processes managed by llm/llama_server.go, which drive the llama.cpp-based engine under llama/server.
- Hardware selection flows through GPU discovery at discover/gpu.go.
- Models are stored as manifests like manifest/manifest.go and fetched from the registry by the transfer logic in transfer/download.go.
Why You Need This
The first reason is that Ollama is the reference implementation of a problem you will eventually face: serving large models on heterogeneous hardware. The scheduler in server/sched.go decides how many models fit, when to load, when to keep one warm, and when to evict, while the discovery package interrogates CUDA, Metal, ROCm, and CPU fallbacks to build a picture of what the machine can actually do. Reading this code teaches memory management for LLM serving in a form far more approachable than production inference stacks, and every claim in this paragraph maps to a file you can open.
The second reason is API compatibility as a design pattern. Ollama does not ask the ecosystem to rewrite tools for a proprietary interface: the server exposes the native API at /api/chat and /api/generate, an OpenAI-compatible surface at /v1/chat/completions and /v1/responses, and an Anthropic-compatible surface at /v1/messages. That triple compatibility is why coding agents designed for cloud providers can run against local models today, and the implementation shows how to translate between streaming protocols without losing features such as tool calls and reasoning output. If you are building any developer product that speaks to models, this is the compatibility layer to imitate.
The third reason is the packaging model. Models are distributed as manifests referencing content-addressed layers, the same conceptual design as container images, implemented in manifest/ and transfer/ with verified downloads and resumable transfers. The Modelfile format in parser/ lets anyone describe a model from a base plus prompt templates and parameters, and template/ ships the chat templates for the popular model families so new users get correct prompts automatically. It is a complete lesson in how to make a distribution format that people can trust, fork, and extend.
How It Works
Inside Ollama: the CLI and its chat loop, the route table with OpenAI and Anthropic compatibility handlers, the scheduler and runner subprocesses down to llama.cpp, the model behavior layer of templates, tools, and thinking output, and the manifest-driven storage stack.
Understanding the Architecture
A CLI that talks to one server. The cobra command set in cmd/cmd.go covers run, serve, pull, push, create, show, stop, list, ps, copy, and remove, with cmd/start.go responsible for getting the server running on each platform and cmd/interactive.go providing the chat loop you land in after ollama run. The client half lives in api/client.go, a typed Go client that every command uses to call the local server, which means the CLI and external integrations exercise exactly the same code paths.
One route table, three dialects. server/routes.go registers the native API, including /api/chat, /api/generate, /api/embed, /api/pull, /api/push, /api/tags, /api/show, and /api/ps, plus the version and status endpoints fed from version/version.go. The OpenAI compatibility handlers in openai/openai.go serve /v1/chat/completions, /v1/completions, /v1/embeddings, and /v1/models, with openai/responses.go adding the stateful Responses API, and anthropic/anthropic.go serves /v1/messages so Anthropic-protocol clients work unmodified. All of these funnel into the same underlying chat and generate handlers, which is why behavior stays consistent across dialects.
A scheduler with real opinions. The scheduler in server/sched.go tracks loaded models against available memory, queues pending requests, and decides when to load, keep, or evict a model, using server/model.go as the per-model record of layers and options. GPU selection starts with discover/gpu.go probing the machine and continues through the device abstraction in ml/device.go, so a laptop with an Apple GPU, a workstation with CUDA cards, and a CPU-only box all run the same logic with different hardware answers.
Runner processes, not libraries. Inference runs in separate llama.cpp-based server processes rather than in-process: llm/llama_server.go manages the subprocess lifecycle, llm/llama_binary.go locates the right engine build, and the runner itself under llama/server speaks back to the Go side over a local endpoint. Constrained decoding is available through the GBNF grammar support in llm/gbnf.go, letting you force valid JSON or other formats from the model.
Model behavior as data. The behavior layer is deliberately declarative: template/template.go selects and renders per-family chat templates, tools/tools.go handles tool schemas and tool-call parsing for agentic use, thinking/parser.go separates reasoning output from final answers, and harmony/harmonyparser.go parses the channel-based format used by newer open models. The Modelfile parser in parser/parser.go and the creation logic in server/create.go turn a description file into stored layers, with server/images.go handling the conversion from GGUF and safetensors sources in fs/gguf and fs/safetensors.
A registry you can trust. Pulls flow from server/download.go into the transfer package at transfer/download.go, which fetches blobs with resumable, verified requests and writes them into the content-addressed store, while transfer/upload.go handles pushes. Everything above the blob layer is described by manifest/manifest.go, so a model name resolves to a manifest, a manifest resolves to layers, and layers are addressed by digest. It is the container-registry mental model, rebuilt for weights and prompts.
Advantages
- One-command local inference. A single install script and a
ollama run gemma4gets you chatting, with no Python environment or GPU driver ceremony for the common path. - Triple API compatibility. Native, OpenAI-compatible, and Anthropic-compatible endpoints on one server mean existing tools and coding agents work against local models unchanged.
- Honest hardware handling. GPU discovery and the scheduler support CUDA, Metal, ROCm, and CPU fallbacks with real memory accounting rather than hope.
- Container-style distribution. Manifest and layer design gives you verified, resumable pulls, easy forking of models, and a format engineers already understand.
- Extensible behavior. Modelfiles, prompt templates, tool calling, thinking output, and grammar constraints cover the features agent builders actually need.
- Readable Go. The codebase is compact, dependency-light, and cleanly split, so tracing a request from HTTP handler to token output is a realistic afternoon.
Benefits
- Privacy by default. Models and data stay on your machine, with the server bound to localhost unless you deliberately expose it.
- A reference for serving design. The scheduler, runner management, and discovery packages are study material for anyone building inference infrastructure.
- Ecosystem leverage. First-party Python and JavaScript clients and the compatibility surfaces connect you to the existing LLM tooling world immediately.
- Agent-friendly. Tool calling, structured output via GBNF, and the documented coding-agent integrations make it a practical backend for local agents.
- Multi-platform discipline. macOS, Windows, Linux, and Docker are all first-class, with platform code isolated rather than scattered.
- License and transparency. MIT-licensed with the full serving stack in the open, from HTTP routing down to the inference engine builds.
Usage
Install, then start the server or just run a model, which starts it automatically:
curl -fsSL https://ollama.com/install.sh | sh
ollama run gemma4
Pull a model without entering chat, then manage what is stored and running:
ollama pull llama3.2
ollama list
ollama ps
ollama stop llama3.2
ollama rm codellama
Create a custom model from a Modelfile, then inspect it:
ollama create my-assistant -f Modelfile
ollama show my-assistant
Call the REST API directly, in the native or OpenAI dialect:
curl http://localhost:11434/api/chat -d '{
"model": "gemma4",
"messages": [{"role": "user", "content": "Why is the sky blue?"}],
"stream": false
}'
curl http://localhost:11434/v1/chat/completions -d '{
"model": "gemma4",
"messages": [{"role": "user", "content": "Write a haiku about caches"}]
}'
Launch a coding agent against your local models, and sign in for registry features:
ollama launch claude
ollama signin
Conclusion
Ollama earns its place as default local-model infrastructure by refusing to cut corners anywhere: the HTTP layer is a real API with real compatibility, the scheduler treats your memory as a budget it must manage, inference runs in supervised processes, and model distribution uses the same trust model as container registries. Read it to learn how to serve models on arbitrary hardware, how to design a compatibility layer that opens an ecosystem instead of closing one, or how to package weights so people can verify what they download. Then run it, because the fastest way to understand local inference is to watch a model load, serve, and evict on your own machine.
Links:
Enjoyed this post? Never miss out on future posts by following us