NIST TEVV-Athlon Framework: How to Build an AI Evaluation
A practical guide to NIST AI 200-2’s four-stage method for designing evidence-based, use-specific AI assessments.
Implement the NIST TEVV-Athlon framework by stating the decision your evaluation must support, identifying stakeholders and operating context, translating important system attributes into measurable Blocks, designing Events and Tools that produce valid evidence, running the assessment under controlled and realistic conditions, and interrogating the results before making a deployment or risk decision. The August 2026 publication is an initial public draft, not a certification standard or universal scorecard.
Key takeaways
- TEVV-Athlon is a tailored assessment method, not a fixed benchmark or compliance certificate
- Its four stages move from goals and scope through construction, measurement, and interpretation
- Blocks identify what to measure; Events create assessment situations; Tools collect and analyze evidence
- Evaluation should cover the deployed system, users, operating context, and important failure conditions
- Results should expose uncertainty, subgroup effects, limitations, and tradeoffs instead of one score
- A decision needs owners, actions, monitoring triggers, and retest conditions
Start with the decision, not the benchmark
TEVV stands for test, evaluation, verification, and validation. It connects an AI system’s technical behavior to its intended use. A test may reveal whether a component meets a condition. Evaluation interprets performance against goals. Verification asks whether defined requirements are met. Validation asks whether the system is suitable for its intended use. The boundaries overlap, so state what decision the evidence must support rather than arguing over labels.
NIST AI 200-2 calls its method an athlon because an AI system is assessed across multiple events, much as an athlete is tested across disciplines. That analogy prevents a common error: treating one benchmark as the complete verdict. A support assistant may need factuality tests, retrieval-permission checks, refusal and escalation scenarios, usability research, accessibility review, latency and cost measurement, and live monitoring. None alone describes whether the service is acceptable.
Begin with a plain-language decision: can this assistant answer billing questions for logged-in customers without exposing another customer’s data, making unauthorized changes, or blocking timely human escalation? Name who will use the result, what action follows a pass or failure, available time and budget, and consequences that matter. NIST’s draft asks about goals, stakeholders, cost, resources, existing techniques, success factors, and challenges. This is scope control, not paperwork.
Translate goals into Blocks, Events, and Tools
In Define & Construct, evaluators turn broad attributes into Metrology Blocks: specific concepts or metrics the assessment will measure. Reliability is too broad to test directly. Useful Blocks might include completion under normal load, recovery after a tool timeout, or the proportion of account changes with valid authorization. Fairness might require separate Blocks for error rates across relevant groups and accessibility barriers.
Define every Block operationally. State its unit, population, sampling frame, procedure, acceptable uncertainty, threshold, and owner. Include qualitative evidence when a number would erase context. A user’s inability to reach a human during a financial dispute may matter more than a small gain in average satisfaction. Do not combine unlike consequences merely to create a dashboard score.
Events are situations in which the system is assessed, and Tools are methods that produce or analyze evidence. An Event might be a fixed test set, simulated multi-turn interaction, adversarial exercise, controlled pilot, or observation of real work. Tools can include benchmark harnesses, red-team methods, rating protocols, user studies, telemetry, statistical analysis, and incident review. Choose the smallest portfolio that covers important failure modes.
Preserve the deployed boundary. If production includes a prompt, retrieval index, memory, policy layer, browser, tools, identity, approval screen, and human handoff, a base-model score is only component evidence. Test components for diagnosis and the full system for operational validity. Record exact versions so results are reproducible and later changes trigger the right retest.
Apply and measure with scientific discipline
Freeze the important protocol before the main run: examples, sampling, exclusions, prompts, evaluator instructions, thresholds, statistical methods, and decision rules. Separate development cases from held-out evaluation. When practical, keep the people tuning the system from seeing every final test case, and use independent review for consequential claims.
Representative data means more than matching topic frequency. Include rare but costly conditions, ambiguous inputs, multilingual use, accessibility needs, out-of-distribution requests, malicious content, tool failures, stale knowledge, permission boundaries, and recovery. Balance ordinary work with stress cases. A system that refuses everything may look safe while failing its purpose; one optimized for average success may hide catastrophic edges.
Validate the measurement process. Check rater agreement, dataset provenance, label quality, missing data, leakage, repeated-user effects, and whether a metric changes when the behavior changes. Report uncertainty when appropriate. Repeat a subset and investigate instability instead of averaging it away.
Human and field testing can raise privacy, labor, accessibility, safety, consent, and research-ethics duties. Obtain appropriate review before recruitment or collection, minimize sensitive data, and define retention. Red teaming also needs authorization, isolated targets, safe success criteria, evidence handling, and a stop procedure. The draft lists methods; it does not remove these obligations.
Synthesize results without manufacturing certainty
Connect results to the original decision. Show performance by Block and Event, not only a blended average. Explain failures, affected people or scenarios, missing evidence, and operational consequences. Compare benefit, risk, latency, cost, and human workload without pretending they share a natural unit.
Interrogate surprising results. A high score may come from duplicated examples, leakage, an easy sample, lenient raters, or a proxy that the system can optimize without improving the real outcome. A low score may expose a genuine limitation or a broken harness. NIST’s draft warns about Goodhart’s Law: once a measure becomes a target, it can stop being useful.
Write conclusions at the strength the evidence supports. An offline English-language test cannot establish suitability for all users or continuous production. A vendor benchmark does not validate your integration. A pilot may support a narrow rollout with monitoring rather than general approval. Record limitations, residual risks, and who accepts them.
Finish with actions: fix and retest, restrict use, add human review, change permissions, improve data, monitor a production signal, or stop deployment. Each action needs an owner, deadline, evidence requirement, and re-evaluation trigger. Model, prompt, retrieval, policy, permission, user population, and environment changes can invalidate prior conclusions.
Use the draft as a flexible design method
TEVV-Athlon complements the AI Risk Management Framework rather than replacing it. NIST maps Govern and Map information into the assessment, treats the athlon as part of Measure, and sends findings into Manage. Governance and context define what matters; evaluation generates evidence; management decides and acts. The loop repeats as the system changes.
Do not present adoption as NIST certification. NIST AI 200-2 is an initial public draft, the AI RMF is voluntary, and the Playbook says its suggestions are not a checklist to follow in full. Legal, contractual, sector, and organizational requirements still apply. Tailoring is valuable only when documented and defensible; removing difficult events is not flexibility.
NIST’s public-comment period runs through October 6, 2026. Organizations can identify unclear definitions, contexts the framework handles poorly, resource assumptions that exclude smaller teams, or reporting elements decision-makers need. NIST requests no proprietary information because comments may be public under FOIA. This guide is based on official documentation, not hands-on validation of a commercial evaluation product.
Practical checklist
- Write the deployment or risk decision the assessment must inform
- Identify decision owners, affected people, operators, reviewers, and domain experts
- Document intended use, prohibited use, operating environment, lifecycle stage, and plausible harms
- Choose system attributes tied to organizational goals
- Define each Block with an operational meaning, method, threshold, and uncertainty treatment
- Construct Events for normal work, edge cases, misuse, adversarial pressure, and recovery
- Select Tools for model tests, red teaming, user studies, field tests, logs, and analysis
- Pre-register datasets, sampling, versions, exclusions, thresholds, and decision rules
- Apply consent, privacy, legal, security, accessibility, and ethics controls before testing
- Analyze failures, subgroup results, uncertainty, costs, and operational consequences
- Record model, prompt, retrieval, tool, policy, data, and environment versions
- Assign every action an owner, deadline, monitoring trigger, and retest condition
Warning signs
- The evaluation starts with a convenient benchmark instead of a real decision
- One score hides severe failures, subgroup differences, or uncertainty
- The test covers a base model while production adds retrieval, tools, memory, and human handoffs
- Acceptance thresholds are chosen after results are visible
- The same examples are used to tune and judge the system without an independent check
- Voluntary draft guidance is presented as certification or legal compliance
Frequently asked questions
What is the NIST TEVV-Athlon framework?
It is a four-stage method in NIST AI 200-2 for designing customized test, evaluation, verification, and validation assessments of AI systems.
Is TEVV-Athlon a required standard or certification?
No. NIST AI 200-2 is an initial public draft, and the AI Risk Management Framework is voluntary. Using it does not create NIST approval or prove legal compliance.
How is TEVV-Athlon different from an AI benchmark?
A benchmark is one possible Tool or Event. TEVV-Athlon starts with an organizational decision and can combine benchmarks with red teaming, system tests, user research, field testing, and logs.
Can TEVV-Athlon evaluate generative AI and agents?
Yes. NIST describes it as adaptable to statistical ML, large language models, multimodal systems, agents, and other AI technologies.
What should a TEVV-Athlon report contain?
Report the decision, scope, stakeholders, system version, Blocks, Events, Tools, data, methods, thresholds, results, uncertainty, limitations, actions, owners, and retest triggers.
Can organizations comment on the NIST draft?
Yes. NIST’s public-comment period closes October 6, 2026. NIST asks commenters not to include proprietary information because comments may be released under FOIA.
Primary sources and further reading
- The TEVV-Athlon Framework for Evaluating AI Systems (NIST AI 200-2 ipd)National Institute of Standards and Technology · August 2026
- The TEVV-Athlon Framework for Evaluating AI SystemsNational Institute of Standards and Technology · Updated August 14, 2026
- Artificial Intelligence Risk Management Framework (AI RMF 1.0)National Institute of Standards and Technology · January 2023
- NIST AI RMF PlaybookNational Institute of Standards and Technology · Accessed August 26, 2026
- AI RMF Playbook: MeasureNational Institute of Standards and Technology · Accessed August 26, 2026