What Langfuse does
Langfuse is an open-source LLM engineering platform for traces, sessions, prompts, datasets, experiments, scores, annotation queues, metrics, and cloud or self-hosted deployment.
Langfuse offers a broad feedback loop without tying the application to one model framework. A team can trace a RAG request, compare prompt versions, connect model feedback and user reactions, create a dataset from failures, and evaluate a change before rollout. Open-source deployment provides architectural control and avoids a proprietary trace format becoming the only copy of quality evidence. It does not remove operational work: self-hosters own availability, migrations, encryption, identity, backups, data lifecycle, and incident response, while cloud users still must configure capture, access, retention, and regions to match their risk.
Cloud Hobby is free with 50,000 units monthly, 30-day data access, and two users. Core is $29 monthly with 100,000 included units, 90-day access, and graduated usage starting at $8 per additional 100,000. Pro is $199 monthly with longer history and controls; a Teams add-on is $300 monthly. Enterprise is listed at $2,499 monthly with additional administration and support. Usage rates decline at volume, and discounts may apply. Self-hosted software has no Langfuse license charge for ordinary use but incurs compute, database, object storage, network, monitoring, staff, upgrades, and external evaluator or model costs. Verify current unit definitions and regional prices.
Over-collection is the central observability risk. Prompt templates may be safe while interpolated values contain identities, health data, account records, secrets, or full proprietary documents. Apply allowlists and redaction before telemetry leaves the application; separate environments and tenants; restrict exports; rotate keys; and test deletion, backup, and retention behavior. Quality signals are also biased data: thumbs-up responses have selection effects, annotation rubrics vary by reviewer, and model judges can reward verbosity or share a provider's blind spots. Version datasets, prompts, application code, scorers, and models; audit coverage and missingness; measure subgroup performance; track drift; and route uncertain or consequential cases to qualified people.
How Langfuse works
Langfuse receives traces, observations, generations, scores, and session or user identifiers through its SDKs, API, OpenTelemetry, model integrations, or proxies. The interface reconstructs application execution, latency, usage, cost, and feedback; prompts can be versioned and fetched by applications. Teams convert examples into datasets, run experiments, attach model, code, or human scores, build dashboards, and use annotation queues. Langfuse Cloud meters billable units and data access by plan, while self-hosting moves infrastructure, upgrades, backups, security, and scaling to the operator. Recorded telemetry and scores are evidence about instrumented traffic, not a complete or neutral view of quality.
Choose Cloud or self-hosted responsibility
Compare region, identity, retention, unit pricing, support, infrastructure, database, storage, backups, upgrades, recovery, security operations, and external model costs before deployment.
Send only allowlisted trace data
SDK, API, OpenTelemetry, model, or proxy integrations record spans, generations, sessions, usage, and scores. Redact before export, pseudonymize identifiers, isolate tenants, and test trace completeness.
Version prompts, cases, and evaluation
Promote authorized failures into representative datasets, compare application variants, and attach human, model, or code scores. Calibrate rubrics against blinded experts and audit missing and subgroup coverage.
Monitor behavior and evaluator drift
Track latency, cost, errors, quality, usage, and telemetry health; investigate underlying traces, refresh datasets after change, audit access and deletion, and require human release and incident review.
How to set up Langfuse
Choose Cloud or self-hosted responsibility
Compare regions, retention, identity, support, unit volume, infrastructure, backups, upgrades, security staffing, model charges, and recovery objectives before selecting deployment.
Specify a telemetry allowlist
Define necessary attributes, pseudonymous identifiers, redaction rules, prohibited secrets and personal data, sampling, tenant isolation, access, export, retention, and deletion.
Instrument and verify trace integrity
Use a test project and synthetic conversations; confirm span hierarchy, sessions, errors, streaming, retries, cost attribution, sampling loss, and that sensitive values never arrive.
Create calibrated scores and datasets
Curate representative cases with provenance, version prompts and rubrics, compare automated scores with blinded experts, and measure disagreement, bias, and slice coverage.
Operate the feedback loop
Run experiments, require human release review, monitor latency, cost, quality, evaluator and traffic drift, refresh cases, audit permissions, and exercise restore and deletion procedures.
Langfuse FAQs
How much does Langfuse Cloud cost?
Hobby is free; Core lists $29 monthly, Pro $199, Teams add-on $300, and Enterprise $2,499, with usage and optional model costs. Verify current calculator terms.
Can Langfuse be self-hosted?
Yes. The open-source deployment shifts database, storage, scaling, upgrades, authentication, backups, monitoring, security, and incident ownership to the operator.
What is a Langfuse billable unit?
Cloud usage is metered in units under the current pricing definition. Included and graduated allowances vary, so model representative traffic in the official calculator.
Does tracing capture the whole user experience?
No. It captures instrumented server and model activity. Missing spans, sampling, browser behavior, downstream systems, and unobserved user outcomes can leave important gaps.
How should evaluator drift be handled?
Pin and version judge models and rubrics, keep expert-labeled anchors, monitor agreement and subgroup errors, investigate distribution shifts, and recalibrate before trusting trends.
Listing reviewed 2026-07-15. Product details and pricing can change; verify important terms on the provider's website.
Related Developer Tools AI tools
Related AI guides
Reviews
Tell the community what you made, what worked, and what you wish you knew before starting.