OpenAI Presence: Enterprise Agent Evaluation Checklist
A buyer-and-deployment checklist for deciding whether OpenAI Presence fits a bounded enterprise workflow and what proof to require before production.
OpenAI Presence is a service-led product for eligible enterprise customers deploying voice and chat agents into defined workflows. It combines policies, approved actions, guardrails, simulations, evaluations, escalation rules, and an improvement process, but it is not a self-serve agent builder. Evaluate it by testing one bounded job against your baseline, mapping every data source and action, defining deterministic approval and escalation boundaries, and requiring production evidence before expanding access.
Key takeaways
- Presence is currently a limited-general-availability, service-led product rather than a self-serve tool
- Its announced scope centers on real-time voice and chat workflows, including customer and internal service
- A suitable first workflow has a clear job, authoritative knowledge, bounded actions, measurable outcomes, and workable escalation
- Policies and model guardrails do not replace authorization, transaction limits, identity checks, or human approval
- Simulations should test normal requests, edge cases, attacks, outages, and handoffs against outcome-level criteria
- Expansion should depend on sustained production evidence, not a successful demonstration or improving average score
Decide whether the workflow fits Presence
Presence is not a general promise to automate an entire department. OpenAI says each deployment begins with a specific job, gives the agent only the knowledge and system access required for that job, and lets the company set policies for actions, approvals, and human takeover. The launch scope is real-time voice and chat, with examples including billing, claims, customer support, outbound sales, and employee IT service. Availability is limited to eligible enterprise customers through a program led by OpenAI engineers and selected integrators.
That makes workflow selection the first buying decision. A strong candidate has meaningful request volume, an authoritative source of truth, decisions that can be expressed as policies, actions that can be bounded, observable outcomes, and a staffed exception path. The agent adds value when requests vary enough to require interpretation but remain narrow enough to evaluate. A fixed rules engine may be cheaper for deterministic work; an employee-facing assistant may be safer when judgment should remain with a person.
Create a one-page workflow contract before discussing integrations. Name the user, trigger, desired outcome, permitted and prohibited work, source systems, actions, approval points, escalation destination, owner, and retirement condition. Record the existing resolution quality, time, cost, transfer rate, repeat contacts, complaints, and serious incidents. Without that baseline, a team can celebrate activity while missing whether customers receive better outcomes.
Turn policies and escalation into testable behavior
Presence combines policies, standard operating procedures, guardrails, approved actions, simulations, evaluation tools, and escalation rules. Treat each as a separate control. An instruction describes desired behavior; a guardrail detects or blocks a class of input or output; authorization decides whether an action may occur; and escalation transfers responsibility. None proves that the others work.
Translate prose policies into decision tables with inputs, allowed outcomes, evidence requirements, limits, and exceptions. Define when the agent must ask a clarifying question, say it cannot help, obtain approval, or transfer. For voice, also test interruptions, accents, background noise, silence, transcription errors, repeated authentication attempts, and the user's ability to reach a person. The transfer should include the relevant context without exposing unnecessary sensitive data or forcing the customer to restart.
Escalation needs an operating commitment. Specify the receiving queue, hours, response target, maximum wait, fallback route, and ownership when no person is available. Measure both over-escalation and under-escalation. An agent that hands off every difficult request may look safe while creating unacceptable queues; one optimized only for containment may conceal rare but severe harm.
Evaluate outcomes before and after launch
OpenAI says Presence uses simulations and graders to check outcomes, policy adherence, tool use, and escalation before launch, then uses production sessions and quality signals to identify improvements. Build your own acceptance plan around those mechanisms. Derive test cases from real request distributions, known errors, complaints, policy exceptions, and expert judgment. Keep a held-out set so the same examples are not used both to tune and approve the system.
Measure the complete workflow: correct resolution, evidence used, policy compliance, authentication, tool arguments, approval, escalation quality, latency, customer effort, repeat contact, cost, and reviewer work. Report high-severity failures separately from averages. Test prompt injection, social engineering, conflicting sources, unavailable tools, stale knowledge, timeouts, duplicate requests, downstream rejection, and partial transactions. Include benign cases so a system cannot pass by refusing everything.
A pilot should restrict users, channels, actions, and transaction size while keeping human takeover staffed. Predefine stop conditions and a rollback that disables the agent or its write permissions without breaking the underlying service. Compare the pilot with the baseline over enough cycles to expose variation. This article applies public documentation to a purchasing and deployment review; it does not claim hands-on access to Presence.
Control change and expansion in production
Production sessions can reveal gaps, and OpenAI says Codex can propose updates that teams test and approve. That improvement loop is useful only when changes remain governed. Version policies, prompts, knowledge, graders, tools, permissions, and escalation rules. Record who proposed and approved each change, which regression set it passed, what population receives it, and how to restore the previous version.
Monitor input mix, resolution and transfer rates, policy violations, overrides, complaints, tool errors, unusual actions, latency, spend, and data-access anomalies. Sample transcripts using a documented privacy and quality process; aggregate metrics alone will not show whether a minority of users experiences a serious failure. Reassess when a model, policy, integration, channel, population, jurisdiction, or business process changes.
Expand one dimension at a time: more users, another request type, a new channel, or a more consequential action. Re-run the threat model and evaluation for each expansion instead of assuming evidence transfers. Keep the option to narrow or retire the deployment. The right outcome may be a hybrid in which Presence resolves bounded requests, deterministic automation executes stable rules, and people retain ambiguous or high-impact judgment.
Practical checklist
- Write the proposed agent job, excluded work, owner, users, channels, and success measures on one page
- Measure the current workflow's volume, resolution quality, handle time, escalation rate, cost, and failure impact
- Map every knowledge source, customer field, credential, tool, action, recipient, processor, and retention requirement
- Separate read, recommend, draft, approve, and execute permissions for each connected system
- Encode monetary, identity, legal, safety, privacy, and account-change rules outside free-form model judgment
- Define when the agent must clarify, refuse, request approval, transfer to a person, or stop
- Build simulations from real request patterns, rare edge cases, policy conflicts, prompt injection, and system failures
- Score correct outcomes, policy adherence, tool use, escalation, customer effort, latency, and reviewer workload
- Pilot with a limited population, narrow actions, explicit rollback, and staffed human escalation
- Review transcript access, data use, retention, residency, deletion, audit, incident, and vendor responsibilities contractually
- Require tested change approval, version history, regression checks, rollback, and post-launch monitoring
- Expand only after the agent meets predefined thresholds across multiple production cycles
Warning signs
- The proposed job is described as handling support generally rather than a bounded class of requests
- A polished demonstration substitutes for a representative baseline, held-out evaluation set, and acceptance threshold
- The agent receives a broad employee account or can execute consequential actions without external authorization
- A human handoff exists on a diagram but has no response target, context-transfer test, queue capacity, or fallback channel
- Policies can change without versioned review, regression testing, approval, and rollback
- The team assumes general enterprise privacy statements answer the exact deployment's data-flow and retention questions
- Success is reported only as containment, automation, or average accuracy while severe failures and customer effort are hidden
Frequently asked questions
Is OpenAI Presence available to every business?
No. OpenAI describes Presence as available to eligible enterprise customers through limited general availability. Deployments are led by OpenAI Forward Deployed Engineers and selected systems integrators, and it is not currently self-serve.
What kinds of workflows does OpenAI Presence support?
The launch announcement describes real-time voice and chat experiences such as customer support, outbound sales, insurance claims, billing, and employee IT service. Confirm the exact supported channels, integrations, regions, and service boundaries for your proposed deployment.
Does OpenAI Presence remove the need for human review?
No. The product is designed around company-defined policies, approvals, and escalation rules. Your organization must decide which actions require deterministic controls or a person and ensure the human path is available and effective.
How should an enterprise evaluate a Presence agent?
Start with a baseline and representative task set. Test outcome correctness, policy compliance, tool selection, authorization, escalation, safety, latency, cost, and reviewer effort across normal, adversarial, and degraded conditions before a limited pilot.
How is Presence different from building an agent with an SDK?
Presence is presented as a deployed product plus implementation and improvement support. An SDK is a code framework your engineering team owns. Compare control, customization, integration, operating responsibility, portability, procurement, and total cost rather than features alone.
Primary sources and further reading
- Introducing OpenAI PresenceOpenAI · July 22, 2026
- A Practical Guide to Building AgentsOpenAI · June 2026
- A Business Leader’s Guide to Working With AgentsOpenAI · June 2026
- Business Data Privacy, Security, and ComplianceOpenAI · Accessed August 4, 2026
- Artificial Intelligence Risk Management Framework: Generative AI ProfileNational Institute of Standards and Technology · July 2024