Arize Phoenix Review

Trace, evaluate, experiment on, and improve AI applications with OpenTelemetry-based workflows.

Independently researched by AI Toolbox Team · Reviewed 2026-07-15
THE SHORT VERSION

What Arize Phoenix does

Arize Phoenix is an open-source AI observability platform for OpenTelemetry traces, evaluations, prompt versions, datasets, experiments, annotations, and self-hosted or cloud use.

Phoenix uses open telemetry concepts to connect what happened in an AI run with whether it was good. An engineer can inspect a retrieval span, annotate a failure, place it into a dataset, replay a revised prompt, and compare experiments under the same criteria. Tracing evaluators is especially valuable because it makes the measurement process inspectable instead of presenting a score as unexplained authority. Yet trace detail depends on semantic mappings and instrumentor coverage; asynchronous tools, client behavior, hidden provider operations, failed exporters, or incorrect span boundaries may create misleading gaps and cost or latency attribution.

The Phoenix project can be run locally or on infrastructure the team operates, with software and hosting terms verified from the current repository and documentation. Arize AX Cloud Free includes 25,000 spans and 1 GB ingestion monthly with 15-day retention. AX Pro lists $50 monthly with 50,000 spans, 10 GB, and 30-day retention. Enterprise is custom with SaaS or self-hosted options, custom scale, and support. Judge-model calls, storage, network, compute, databases, upgrades, backups, and staff are separate considerations. Confirm project license, cloud regions, span definitions, overages, retention, support, taxes, and the boundary between Phoenix and AX capabilities.

Full prompts and retrieval context may be the most sensitive data in an application. Establish attribute allowlists, content capture controls, redaction at instrumentation, tenant isolation, region and retention policy, encryption, least privilege, audit, export, deletion, and breach response before production collection. An evaluator explanation can sound reasoned while reflecting rubric ambiguity, judge-model bias, prompt injection, position effects, or data contamination. Use deterministic checks where possible; randomize pairwise order; validate on expert-labeled anchors; inspect confusion by language, topic, and subgroup; separate judge inputs from untrusted instructions; monitor model and traffic drift; and require people to resolve uncertain or high-impact failures.

UNDER THE HOOD

How Arize Phoenix works

Phoenix collects OpenTelemetry spans using OpenInference semantic conventions and integrations for model providers and application frameworks. Traces expose model calls, retrieval, tools, custom logic, tokens, latency, errors, sessions, and annotations. Teams replay or version prompts, turn traces into datasets, run controlled experiments, attach deterministic, custom, human, or LLM-based evaluations, and compare results. Evaluator executions are themselves traced so teams can inspect judge prompts, inputs, scores, explanations, timing, and systematic behavior. Phoenix can run locally or self-hosted; Arize AX adds managed retention, online evaluation, monitoring, alerting, and enterprise controls. Instrumentation and evaluators still require validation.

01 · MODEL

Define OpenTelemetry and deployment boundaries

Choose local, self-hosted, or AX scope; specify trace and session semantics, allowed OpenInference attributes, redaction, sampling, tenant isolation, region, retention, deletion, and owners.

02 · OBSERVE

Collect and inspect execution spans

Instrumentors and exporters send model, retrieval, tool, and custom spans over OTLP. Exercise streaming, retry, timeout, async, and failure paths to expose hierarchy gaps and incorrect cost or latency attribution.

03 · TEST

Run traceable, calibrated evaluations

Create datasets and experiments, apply deterministic, human, custom, or LLM judges, and inspect evaluator traces. Test injection, order effects, disagreement, explanations, and errors by language and subgroup.

04 · IMPROVE

Release changes with continuing evidence

Replay and compare prompt or application variants on identical cases, obtain human approval, monitor production signals and drift, inspect alerts, refresh baselines, and audit privacy and infrastructure controls.

YOUR INPUTARIZE PHOENIXREVIEWED OUTPUT
QUICK START

How to set up Arize Phoenix

1

Map Phoenix and AX responsibilities

Choose local, self-hosted, or managed scope; verify license, regions, span volume, retention, online evals, identity, support, infrastructure, upgrades, backups, and costs.

2

Design safe OpenTelemetry semantics

Define trace and session boundaries, approved OpenInference fields, redaction, pseudonymous IDs, sampling, tenant isolation, prohibited content, access, retention, and deletion.

3

Test instrumentation completeness

Run synthetic success, streaming, retry, timeout, tool, retrieval, and failure scenarios; confirm hierarchy, exporter loss, token and latency attribution, and absence of secrets.

4

Build and calibrate evaluators

Create versioned datasets and rubrics, prefer deterministic checks where valid, trace judge executions, blind human labels, and measure disagreement and subgroup performance.

5

Experiment and monitor drift

Compare changes on identical cases, require human release approval, sample production lawfully, alert on quality and telemetry gaps, and refresh baselines after application or judge changes.

COMMON QUESTIONS

Arize Phoenix FAQs

Is Arize Phoenix open source?

Yes, Phoenix is an open-source, local-first platform. Verify the current repository license and account for infrastructure, model, storage, security, and operations costs.

How much does Arize AX cost?

AX Free lists 25,000 spans monthly; Pro is $50 monthly with 50,000 and longer retention; Enterprise is custom. Confirm overages and current terms.

What is OpenInference?

It supplies semantic conventions and instrumentation for expressing AI model, retrieval, tool, and agent activity through OpenTelemetry-compatible traces.

Can Phoenix evaluate LLM judges themselves?

Evaluator tracing records judge inputs, prompts, scores, explanations, and timing so teams can inspect systematic bias and calibrate against human labels.

Does self-hosting guarantee telemetry privacy?

No. Operators must correctly configure network access, authentication, encryption, secrets, backups, logs, retention, tenant isolation, updates, deletion, and incident response.

Listing reviewed 2026-07-15. Product details and pricing can change; verify important terms on the provider's website.

KEEP RESEARCHING

Related Developer Tools AI tools

Related AI guides

COMMUNITY NOTES

Reviews

Be the first to share a detailed review.

Tell the community what you made, what worked, and what you wish you knew before starting.