What LangSmith does
LangSmith is a framework-agnostic observability and evaluation platform for traces, datasets, experiments, online evals, annotation, prompts, and agent debugging.
LangSmith connects debugging evidence to an improvement cycle. Engineers can inspect an agent trajectory, identify a poor retrieval or tool choice, save the run as a test case, compare a proposed change on a stable dataset, and route ambiguous outputs to domain experts. Framework-agnostic instrumentation supports applications beyond LangChain. This structure is more reliable than ad hoc screenshots, but it captures only what instrumentation emits and what sampling retains. Missing spans, swallowed errors, client-side behavior, external tool effects, and unlogged user outcomes can still make a trace look complete when it is not.
Developer is free for one seat and includes 5,000 base traces monthly before pay-as-you-go. Plus lists $39 per seat monthly with 10,000 included base traces and usage charges; Enterprise pricing is custom with additional hosting, identity, and support choices. Base traces currently have 14-day retention and cost $2.50 per 1,000 after allowances, while extended traces retain 400 days and cost $5 per 1,000; upgrading incurs the difference. Evaluator model calls, application models, deployment, Engine, and other services can add separate costs. Confirm trace definition, ingestion limits, regions, retention, taxes, and current invoice terms.
Prompts, retrieved passages, tool inputs, outputs, feedback, and metadata may contain personal data, credentials, confidential documents, or regulated records. Redact at the source, disable unnecessary content capture, use pseudonymous IDs, scope projects and exports, set short retention, and test deletion and incident procedures. LLM judges can favor their own style, leak benchmark knowledge, mishandle dialects, and drift after model changes. Build representative datasets with rare and harmful cases, blind and randomize comparisons, measure agreement by subgroup, calibrate against qualified reviewers, version rubrics and judge models, inspect disagreements, and never let one aggregate score automatically authorize a consequential release.
How LangSmith works
An application sends asynchronous traces to a LangSmith project through an SDK, API, integration, or OpenTelemetry. A trace groups spans for model calls, retrieval, tools, and custom logic with inputs, outputs, timing, tokens, metadata, and feedback. Teams filter runs, build datasets, compare experiments, version prompts, route examples to annotation queues, and attach code, heuristic, human, or LLM-as-judge evaluators. Offline evals test controlled datasets; online evals sample production runs for continuing quality signals. Retention and billing depend on base or extended traces and plan. Tracing explains recorded execution, not unobserved causality, and evaluator scores require calibration.
Define measurable decisions and safe telemetry
Specify release and monitoring questions, required spans, prohibited fields, source redaction, sampling, consent, tenant and regional boundaries, retention, deletion, incident response, and human owners.
Capture and verify application execution
SDK, API, integrations, or OpenTelemetry send asynchronous traces containing instrumented model, retrieval, tool, and custom spans. Test missing exporters, retries, errors, timing, token attribution, and secrets.
Compare versions with calibrated evidence
Build versioned datasets, run offline experiments, and combine deterministic, custom, human, and judge evaluators. Blind and randomize comparisons, inspect disagreements, and measure subgroup bias and uncertainty.
Turn production failures into reviewed tests
Sample lawful traffic, run online evals, route ambiguous traces to experts, alert on telemetry and quality drift, control retention and spend, and keep accountable people in every release decision.
How to set up LangSmith
Define decisions and telemetry boundaries
List release and monitoring questions, necessary spans, prohibited fields, redaction, consent, regions, retention, access, deletion, sampling, incident response, and human owners.
Instrument a synthetic environment
Create a separate project, enable supported SDK or OpenTelemetry tracing, use test data and scoped keys, and verify parent-child spans, errors, timing, token, and cost fields.
Build a representative evaluation set
Combine authored cases, consented failures, edge conditions, adversarial inputs, languages, accessibility needs, and subgroup slices with provenance and versioned expected behavior.
Calibrate every evaluator
Specify rubrics, compare heuristic and judge scores with blinded expert labels, measure disagreement and subgroup errors, inspect rationales, and set uncertain cases for review.
Gate and monitor conservatively
Compare versions with confidence intervals, require human approval, sample production traffic lawfully, alert on drift and missing telemetry, audit spend, and refresh datasets and judges.
LangSmith FAQs
How much does LangSmith cost?
Developer is free with 5,000 included base traces monthly. Plus is $39 per seat with 10,000, then trace usage; Enterprise and other services are additional.
Do I need LangChain to use LangSmith?
No. LangSmith is framework-agnostic and supports SDK, API, and OpenTelemetry approaches for applications built with other frameworks or custom code.
What is the difference between offline and online evals?
Offline evals compare versions on curated datasets before release; online evals score sampled production runs to detect emerging failures and quality drift.
Are LLM-as-judge scores reliable?
They are imperfect measurements. Calibrate them against qualified human labels, version the rubric and judge, inspect disagreements, and monitor subgroup error and drift.
What telemetry should be excluded?
Exclude secrets and unnecessary personal, confidential, regulated, or copyrighted content; redact before export and apply access, retention, region, deletion, and incident controls.
Listing reviewed 2026-07-15. Product details and pricing can change; verify important terms on the provider's website.
Related Developer Tools AI tools
Related AI guides
Reviews
Tell the community what you made, what worked, and what you wish you knew before starting.