When an LLM application misbehaves in production, the question is never whether something went wrong but where, and that question needs infrastructure to answer. Langfuse, the MIT-licensed open source LLM engineering platform at langfuse/langfuse, is that infrastructure: tracing, evaluation, prompt management, and dashboards that teams can self-host in minutes, with SDKs for Python and JavaScript and an OpenTelemetry-compatible ingestion path. The README is direct about the stack and the ambition, describing a platform to develop, monitor, evaluate, and debug AI applications, proudly made with ClickHouse, and it carries a development the observability world felt in January 2026: Langfuse is now part of ClickHouse, the company behind the database that already powers its analytics core. This post tours the published source of that monorepo, which is one of the most instructive TypeScript codebases in the AI infrastructure category.
The layout is a pnpm and Turborepo workspace with clear seams. web is the Next.js application users live in. worker is a separate Node service that consumes background jobs. packages/shared holds the code both sides import, including the Prisma schema for Postgres and the ClickHouse client. ee carries enterprise features that stay out of the permissive core path, fern holds the API specification from which SDKs are generated, and ai-gateway rounds out the top level. As always in this series, what follows is an educational tour of published source code.
Langfuse at a glance: Fern-specified public API routes and an OpenTelemetry endpoint feed BullMQ ingestion queues, the worker consumes more than twenty queues that write traces and scores into ClickHouse and metadata into Postgres, and the Next.js web app renders the feature modules that read it all back.
Reading the overview from left to right:
- The API contract is generated from the Fern specification in fern, served through the public routes under web/src/pages/api/public.
- OpenTelemetry spans arrive at the server code in web/src/server/otel, which feeds the queue in worker/src/queues/otelIngestionQueue.ts.
- The native ingestion path lands in worker/src/queues/ingestionQueue.ts, one of the many queues the BullMQ worker in worker/src/queues consumes.
- Span-to-row mapping happens in the OtelIngestionProcessor in packages/shared/src/server/otel/OtelIngestionProcessor.ts, which writes through the ClickHouse client in packages/shared/src/server/clickhouse/client.ts.
- Relational state lives in Postgres via the Prisma schema in packages/shared/prisma, from organizations and projects to prompts and API keys.
- The user surface is the Next.js app in web/src/app, composed from the feature folders beginning with web/src/features/traces.
Why You Need This
The first reason is the ingestion architecture, which is a masterclass in swallowing bursts without losing data. The public API does almost no heavy lifting: it authenticates, validates, and enqueues, pushing ingest events into a Redis-backed BullMQ queue consumed by the separate worker service. The same is true for OpenTelemetry; the OTel endpoint accepts protocol-native spans, and the OtelIngestionProcessor in the shared package maps the generic span tree onto Langfuse’s domain model of traces, observations, and scores, with a dedicated media processor extracting binary attachments into object storage. Batching logic in the worker’s traceBatching feature groups rows for bulk inserts into ClickHouse, and the tokenisation feature computes model-specific token counts. The design lesson is architectural honesty about write patterns: LLM telemetry is spiky, high-volume, and append-heavy, so the front door is thin, the queue absorbs the spikes, and the columnar database gets tidy batches.
The second reason is the queue catalog, which reads like a map of everything an LLM platform must eventually do. The worker/src/queues folder contains ingestion, OTel ingestion, evalQueue, experimentQueue, batchActionQueue, batchExportQueue, dataRetentionQueue, datasetDelete, projectDelete, scoreDelete, entityChangeQueue, eventPropagationQueue, notificationQueue, monitorQueue, topicsQueue for conversation topic mining, topicsEmbeddingQueue, blobStorageIntegrationQueue, coreDataS3ExportQueue, and cloud-specific metering queues. Each has a matching processing module under worker/src/features, from the evaluation runner to retention cleaners that purge aged data from both databases and object storage. The eval queue deserves special attention: it executes LLM-as-judge and template-based evaluations over production traffic, writing scores back that then appear next to the traces they judge. If you have ever wondered what features hide behind an observability product’s pricing page, this folder is the itemized receipt.
The third reason is the dual-database storage design, which answers a question every platform team eventually faces: where should relational state end and analytics state begin? Postgres via Prisma owns the things that need transactions and foreign keys, organizations, users, projects, API keys, prompts with their versioned templates, and dataset definitions. ClickHouse owns the firehose: traces, observations, scores, and the event tables that dashboards aggregate, served through a typed client in packages/shared/src/server/clickhouse with schema helpers, query tagging, and migration tooling. Table mapping modules translate rows into the shapes the UI consumes. The README’s pride in ClickHouse predates the acquisition, and the source makes the reason obvious: aggregating millions of spans for a latency histogram is exactly the workload a columnar engine loves, and Langfuse bet its core on that early.
The detailed view: the API surface with auth and media handling, the ingestion path from queues through processors and batchers, the queue catalog on the worker, the storage layer with ClickHouse mappings, Prisma and the PromptService, and the web application's feature modules with the enterprise seam.
The detailed diagram rewards a slow pass. On the surface group, both the public API routes and the OTel endpoint verify credentials through the shared auth module, which carries multiple single sign-on providers alongside classic API keys, and the media processor routes binary payloads into the S3-backed StorageService rather than into database rows. In the queue group, the eight highlighted queues are only part of the catalog; each one pairs with an interface module that owns retries, failure logs, and dead-letter handling. In the storage group, the ClickHouse client sits behind table mapping modules such as mapTracesTable and mapScoresTable, while the PromptService caches versioned prompt templates that the playground and the SDKs fetch. The web app group shows the breadth of the feature tree, traces, evals, datasets, prompts, and a playground that calls models through the ai-gateway package, plus newer additions like the MCP feature that lets coding agents query the platform, an in-app agent runner with its own queue, and RBAC and entitlements modules where the enterprise edition plugs in.
One more group matters more than its size: natural-language-filters and the topics features show where the platform is heading, letting users interrogate trace data in prose while an embedding queue clusters conversations into topics in the background. The v4 migration folders signal a generational API revision in progress. Both are worth watching because they change what users touch, not just what the code contains.
From Install to First Trace
The self-hosting path is docker compose. From a clone of the repository:
docker compose up -d
That brings up the web UI, the worker, Postgres, ClickHouse, Redis, and MinIO-backed object storage, wired by the environment files shown in the .env.prod.example at the root. Sign in, create an organization and project, and generate an API key pair. Then instrument from Python, using the same package the platform itself lists in its worker dependencies:
pip install langfuse
from langfuse import get_client
langfuse = get_client()
with langfuse.start_span(name="answer-question") as span:
span.update(input="Why is the sky blue?")
# ... call your model here ...
span.update(output="Rayleigh scattering.")
langfuse.flush()
The OpenTelemetry route works with any OTLP exporter already in your stack, so existing instrumented services can adopt Langfuse without new agents. Once traces land, the workflow is the product: inspect a trace tree of observations, score outputs manually or let the eval queue run LLM-as-judge templates, promote a good prompt into versioned management, then build datasets from production traces and run experiments against them. Batch exports push everything to your own S3 bucket when you want the raw rows.
The honest closing is about the acquisition banner. Platforms in the observability tier rarely publish this much engineering in the open, and the ClickHouse acquisition both validates the database bet and raises fair questions about where the open core ends and the enterprise begins, a boundary the ee folder makes explicit. For a team choosing instrumentation today, the source offers something rare: you can read exactly what the hosted product does, run the whole stack yourself, and understand every queue that touches your data before you send a single trace. That transparency is the feature no dashboard can show, and it is the reason this repository belongs on the short list. Enjoyed this post? Never miss out on future posts by following us