What Braintrust does
Braintrust is an AI observability and evaluation platform for traces, Topics, playgrounds, datasets, experiments, scorers, annotations, online evaluation, and CI gates.
Braintrust centers observability around the idea that production evidence should improve the next test. A support-agent failure can be traced, classified, reviewed, saved as a dataset row, scored against a change, and enforced in CI without translating between unrelated logging and experiment formats. Topics and natural-language analysis help discover patterns across high trace volumes. Automated clustering is still an interpretation: rare but serious failures may disappear into sampling, a topic label can merge distinct causes, and a trace does not reveal customer impact unless the team captures a lawful outcome signal and examines individual cases.
Starter is $0 monthly with included analysis credits, 1 GB processed data, 10,000 scored outputs, 14-day retention, and unlimited users, projects, datasets, playgrounds, and experiments; overages list separate data and score rates. Pro is $249 monthly with larger allowances, 30-day retention, additional controls, and lower usage rates. Enterprise is custom for retention, export, deployment, privacy, identity, and support. Model calls and application infrastructure remain separate. The pricing page may show promotional credits or startup offers, so verify processed-data definitions, score counting, analysis token rates, retention, overages, taxes, deployment, and current contract terms.
Trace payloads may reproduce user messages, model answers, documents, tool calls, metadata, identifiers, or secrets. Apply schema-level allowlists and redaction before logging, use pseudonyms, separate tenants and environments, restrict annotation access and exports, and test deletion and retention. Datasets drawn from reported failures can overrepresent vocal users, while production sampling can miss low-volume harms. LLM scorers can be injected by evaluated text, favor particular tones, or drift with provider updates. Maintain independently authored and consented cases; stratify by language, subgroup, task, and risk; blind human review; measure scorer agreement and uncertainty; monitor topic and traffic drift; and keep humans responsible for release decisions.
How Braintrust works
Applications log traces and spans to Braintrust through SDKs, integrations, or APIs. Logs share a structure with experiments, allowing teams to search production behavior, attach scores and feedback, extract prompts, classify patterns with Topics, and promote selected traces into datasets. An evaluation combines data, a task or application version, and scoring functions; experiments preserve comparable snapshots and can run in code, the UI, or CI. Online scoring asynchronously samples production traces, while human review adds annotations. Braintrust meters processed data, scores, retention, and optional analysis tokens by plan. Search, clusters, and scores reveal measured patterns, not complete causal truth.
Set the quality question and data rules
Document success and harm criteria, release gates, telemetry allowlists, redaction, consent, sampling, tenant isolation, region, retention, deletion, incident handling, and accountable experts.
Trace production and surface candidate patterns
SDKs and integrations record searchable spans and feedback. Topics, queries, and dashboards classify patterns, but teams verify clusters against individual traces, exporter health, sampling bias, and real outcomes.
Promote evidence into comparable experiments
Curate versioned datasets and run tasks with code, human, autoeval, or judge scores. Blind expert review, test injection and order effects, and quantify uncertainty and subgroup disagreement.
Monitor and improve without automating authority
Use CI and online scores as reviewed signals, investigate regressions, refresh cases after drift, control data and score spend, and require humans to approve thresholds, releases, and consequential interventions.
How to set up Braintrust
Define the improvement decision
Write quality, safety, cost, and latency questions; release ownership; telemetry allowlists; redaction; consent; sampling; tenant, region, retention, deletion, and incident rules.
Trace synthetic workflows first
Instrument success, retry, streaming, retrieval, tool, timeout, and failure cases with scoped keys; confirm span structure, missing data, usage attribution, and privacy filters.
Curate versioned datasets
Combine authored specifications, consented traces, rare failures, adversarial cases, languages, accessibility needs, and subgroup slices with provenance, exclusions, and expected behavior.
Validate scorers and Topics
Define rubrics, compare automated results and clusters with blinded experts, test prompt injection and order effects, and measure disagreement, coverage, and subgroup error.
Operate guarded experiments
Run identical datasets across versions, require human approval for CI thresholds, monitor sampled production scores and telemetry health, control usage, and refresh cases after drift.
Braintrust FAQs
How much does Braintrust cost?
Starter is free with included processed data, scores, and short retention. Pro lists $249 monthly plus usage; Enterprise has custom deployment, retention, and support.
What are Braintrust scores?
Scores are outputs from LLM judges, built-in autoevals, custom code, or humans. They measure a defined criterion and require rubric, reliability, and bias validation.
What are Topics?
Topics use automated facets and clustering to surface intents, sentiment, and issues in logs. They aid discovery but can mislabel, merge causes, or miss rare behavior.
Can online evaluation block a bad response before delivery?
Online scoring is described as asynchronous monitoring, so it should not be assumed to prevent delivery. Use separately validated runtime controls for immediate intervention.
How should production traces become test cases?
Sample lawfully, redact and minimize content, preserve provenance, obtain needed consent, avoid selection bias, add rare and adversarial cases, and require expert labeling.
Listing reviewed 2026-07-15. Product details and pricing can change; verify important terms on the provider's website.
Related Developer Tools AI tools
Related AI guides
Reviews
Tell the community what you made, what worked, and what you wish you knew before starting.