Langfuse: observe your local LLMs (traces, prompts, evals)
Langfuse is an observability platform for LLM applications, open source under the MIT license (main repository, excluding “ee” directories), that can be self-hosted with Docker Compose in a few minutes and has over 35,000 GitHub stars. It records the exact prompt sent to the model, its response, the duration of each step, and, if a price is defined, its cost—enough to trace a bad response back to its cause.
An application built on an LLM fails differently from traditional software: nothing crashes; the response is simply worse than yesterday's. Without a trace of what was sent to the model and what it returned, you debug blindly. Langfuse is an observability platform designed for this, open source and self-hostable, making it a good fit for a local installation: your conversation traces stay with you.
#Why monitor a local LLM
Langfuse is an observability platform for LLM applications, licensed under the MIT license for its main repository, that you can self-host so your conversation traces stay with you. For each request, it records the prompt actually sent to the model, the response, the duration of each step, and, if a price is defined, the cost. It helps you understand why a response is bad: an irrelevant retrieved passage, a prompt different from what you thought, or a model failure. It becomes useful as soon as a second person depends on the result or the pipeline has multiple steps; for solo exploratory use, a log file is enough. Expect a stack of several containers (web server, worker, PostgreSQL, ClickHouse, Redis, object storage) and, for production, Kubernetes rather than Docker Compose.
When an answer disappoints, there are three possible causes, and only one is visible without instrumentation: the final prompt sent to the model was not the one you thought, the passages retrieved by document search were irrelevant, or the model itself failed. Without logging, you change the prompt at random until things improve, never knowing why.
The argument is the same locally as online, with one difference: the bill is no longer the warning signal. No one gets a statement when an agent chain loops on your own graphics card; latency and heat are what tell you. Observability replaces that price signal with measurements: token count, duration, failure rate, and quality.
#What Langfuse records
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The project describes itself as an open-source agent evaluation and observability platform: trace, evaluate, and improve LLM applications on a single open platform. Its Python SDK (version 4) and JS/TS SDK (version 5) are built on OpenTelemetry, and other languages can send their traces through this protocol. The main repository has more than 35,000 stars on GitHub. Its README states that Langfuse has been part of ClickHouse since January 2026, and ClickHouse is also the database that stores the traces.
- Traces
- The complete flow of a user request: every step, in order, with its duration. This is the basic unit.
- Observations
- Inside a trace: model calls with their exact prompt and response, document retrieval steps, and tool calls.
- Sessions and users
- Grouping traces by conversation and by person, to track a journey rather than an isolated call.
- Scores
- A note attached to a trace: a user's thumbs-up, an automated evaluation result, or a manual annotation.
- Prompts
- Versioned prompt templates retrieved by the application at runtime rather than hard-coded into the code.
- Costs
- Langfuse calculates a cost from a model definition that assigns a price to each usage type. It provides definitions for OpenAI, Anthropic, and Google models. For a local model, no price exists until you add one: without a definition, do not look for a cost in the interface, or define a fictitious price to reflect electricity and depreciation.
#Install it at home
Langfuse self-hosts with Docker Compose. The README says startup takes five minutes; the deployment documentation estimates two to three minutes before the web container displays “Ready.” The stack is more than one container: two application containers (the web server, which serves the interface and API, and a worker that processes events asynchronously) and four storage components: PostgreSQL for transactional data, ClickHouse for traces, observations, and scores, Redis or Valkey for the queue and cache, and S3-compatible object storage that retains all incoming events. This is heavier than a simple dashboard, and that follows from what it is asked to do: ingest many events without losing a trace, even when the database is temporarily unavailable.
Two limitations to know before installing it alongside a model. First, the documentation says Docker Compose is for trying it out: this configuration provides no high availability, scaling, or backups, and for production it recommends Kubernetes with Helm. Second, for a virtual machine, it recommends at least 4 cores, 16 GiB of memory, and 100 GiB of storage. On a workstation already serving a model, these figures matter: an observability stack that consumes part of the model’s memory does so at the worst possible time. Finally, change the secrets in docker-compose.yml, marked CHANGEME: there is no default account, and the first user is created with the Sign up button on http://localhost:3000.
| Solution | What it offers | Its limitations |
|---|---|---|
| Custom log file | Nothing to install, starts immediately | No structure: impossible to link a tool call to the final response without rereading the entire file |
| Self-hosted Langfuse | Structured traces, versioned prompts, scores—everything stays local | A stack of several containers to run and back up |
| Online observability platform | Zero installation, automatic updates | Prompts and responses pass through a third party, contrary to the spirit of a local installation |
#Connect it to Ollama or vLLM
Langfuse is not an inference server and does not sit in the traffic path: your application sends it the traces. The project documents three concrete paths depending on how you call your model, plus a dedicated page for Ollama, whose example uses the OpenAI SDK: Ollama exposes an OpenAI-compatible API at http://localhost:11434/v1, and Langfuse provides a direct replacement for that SDK, requiring only an import change. There is also a page for vLLM.
- 01Through the SDK, in your codeYou decorate the functions that call the model. This is the most explicit and faithful approach, since you choose what gets traced.
- 02Through framework integrationAutomated instrumentation is available for a direct replacement of the OpenAI SDK, for a callback handler connected to a LangChain application, and for LlamaIndex's callback system: traces are collected without changing the business logic.
- 03Through an API gatewayIf your calls already go through an OpenAI-compatible router, instrumentation happens at that level and covers all applications at once.
In all three cases, the point to verify is that the prompt actually sent is recorded, not just the user's question: it is precisely the gap between the two that explains most poor answers from a document search system. With Ollama, the setup below points LANGFUSE_BASE_URL to your self-hosted instance (http://localhost:3000) and declares your public and secret keys as environment variables.
- Centralize your calls with an OpenAI-compatible gateway
- Deploy vLLM to serve multiple users
- Call Ollama from Python
- Ragas: evaluating a local RAG with numbers
- Source: details of SDK, framework, and gateway integrations
#Manage and version prompts
Getting prompts out of the code is the first benefit teams notice. According to the README, a solid server-side and client-side cache lets you iterate on prompts without adding latency to the application. The template lives in Langfuse, the application retrieves it at runtime, and each change creates a version. Fixing a wording issue no longer requires a redeploy, and most importantly, the trace shows which version produced which response: when quality drops after a change, you know which one to blame without having to cross-reference separate deployment logs.
#Test sets and evaluation
Real traces feed test sets: you flag interesting cases—especially failures—and they become a collection against which to replay a new prompt version or another model. That is what lets you seriously answer “is the 14-billion-parameter model enough?” instead of debating it.
Scores come from several sources: users, human annotation in the interface, code-based evaluators, or a model that judges another model's response according to a rubric—the method known as LLM-as-a-judge, which Langfuse supports natively. For the latter case, you need to create an LLM connection. The documentation states that any model following the OpenAI API schema works by replacing the base URL: a local model served by Ollama can therefore act as the judge, provided its address is reachable from the Langfuse container. The method is useful and biased: the authors of the reference paper on the subject (arXiv 2306.05685) describe position, verbosity, and self-preference biases, while measuring over 80% agreement between a GPT-4-type judge and human preferences. It is useful for comparing two versions, not for assigning an absolute score.
A tool dedicated to RAG evaluation, such as Ragas, complements Langfuse rather than replacing it: Langfuse captures the trace and stores the score, while Ragas calculates specific metrics — context faithfulness and answer relevance — that Langfuse can then display like any other score attached to a trace.
#Concrete use cases
- Support assistant with RAG
- An off-topic response can be traced back through the trace: was the retrieved passage relevant, or did the model ignore a correct passage? It settles the question between the two hypotheses.
- Agent that calls tools
- In a multi-step chain—research, computation, writing—the trace shows which step failed instead of leaving you to guess at a global failure with no details.
- Comparing two models before deciding
- Running the same test suite again with an 8-billion-parameter model and then a 27-billion-parameter model produces a quality score and a duration, rather than an impression.
- Tracking a regression after an update
- A prompt or model version change that degrades quality shows up in aggregate scores before a user complains about it.
What these cases have in common: none can be solved by rereading code. The cause of a bad answer is found in the data that passed through the system—the exact prompt, the retrieved passages, and the generated response—and that data exists nowhere else if it was not recorded when the call was made. That is the whole case for installing observability before going live rather than after the first incident.
#What it costs and what it doesn't do
- That doesn't make a model better
- Observability measures; it doesn't fix anything. It tells you where to look.
- It stores your conversations
- Even when self-hosted, prompt content is written to disk, and without a retention policy (an enterprise-edition feature), the data is kept indefinitely. The SDKs do, however, offer masking functions to remove sensitive data before traces are sent.
- It is a stack to maintain
- Backups, updates, migrations: the documentation itself says that Docker Compose does not provide backups. For personal experimentation, that is disproportionate.
- It arrives late, often
- The right time to install it is before deployment, not after the first silent failure.
#FAQ
Is Langfuse free?+
Does it work with a local model?+
Are my prompts sent over the internet?+
How is this different from standard logs?+
Do you need a GPU for Langfuse?+
Can you evaluate a RAG with Langfuse?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.