The National Institute of Standards and Technology (NIST) has unveiled a draft framework for evaluating Artificial Intelligence (AI) systems, dubbed TEVV-Athlon, which promises to fundamentally alter how leaders approach AI governance. Larry Marks, a seasoned Governance, Risk, and Compliance (GRC) adviser and consultant, argues that this new framework offers a vital reorientation for organizations that have become accustomed to checklist-style compliance. The core of Marks’s argument centers on the critical distinction between merely demonstrating the completion of a process and truly understanding the outcomes of an evaluation. He cautions that a temptation to transform TEVV-Athlon into yet another compliance checklist risks activating Goodhart’s Law, where a measure, once it becomes a target, ceases to be an effective measure.

The innovative aspect of NIST’s TEVV-Athlon framework, as highlighted by Marks, lies not in what it mandates, but in what it deliberately omits: a universal list of tests that every AI system must pass. This omission underscores NIST’s recognition of the inherent diversity of AI systems, their myriad applications, and the correspondingly varied risks they present. A one-size-fits-all evaluation methodology is deemed impractical and potentially ineffective. Instead, TEVV-Athlon is designed for adaptability and customization, encouraging organizations to tailor their evaluations to their specific objectives and the unique characteristics of the AI systems under scrutiny.

Marks draws upon his extensive career in cybersecurity and technology risk assessments, where a recurring lesson has been the ease with which an organization can master the demonstration of process completion without necessarily ensuring the assessment of relevant risks. The standardization of frameworks, a natural progression in organizational maturity, often leads to the transformation of nuanced questions into simple checkboxes, evidence collection into a bureaucratic exercise, and management approvals into perfunctory sign-offs. In such scenarios, the act of completing the process can eclipse the fundamental purpose of understanding what the process was designed to reveal. NIST, through TEVV-Athlon, appears intent on preempting this pitfall.

The TEVV-Athlon framework guides organizations through four distinct stages, beginning with the articulation of objectives and organizing the evaluation around them. Only after this foundational step do organizations proceed to define what needs to be measured, how those measurements will be conducted, and the interpretation of the resulting evidence. The four stages are: Articulate and Organize, Define and Construct, Apply and Measure, and finally, Synthesize and Interrogate. This structured, yet flexible, approach necessitates a shift in mindset for compliance and risk professionals. The ultimate goal should not be to simply declare that an AI system has "passed TEVV," but rather to ascertain whether the evaluation was thoughtfully designed to yield the crucial information an organization needs about a particular AI system within its intended operational context. This distinction, though seemingly subtle, carries significant practical implications.

The Peril of a "Passing Score": Creating False Security in AI Evaluation

A particularly insightful element of the NIST draft is its explicit discussion of Goodhart’s Law. NIST applies this economic principle to AI evaluation, warning that an overemphasis on optimizing a system’s performance against specific benchmarks may not accurately reflect its real-world efficacy. This cautionary note is of paramount importance for compliance and risk professionals, who are accustomed to leveraging various metrics—risk ratings, control effectiveness scores, key risk indicators—to inform management about risk landscapes. While these measurements are invaluable for distilling complex information into actionable insights, the danger arises when the pursuit of a favorable score becomes the primary objective.

This same predicament can manifest in AI evaluations. An organization might establish performance thresholds, conduct tests, and subsequently conclude that an AI system has successfully met its evaluation criteria. Management, presented with this "passing" result, might then erroneously assume that the associated risks have been adequately addressed. However, the critical question remains: what exactly "passed," and under what specific conditions?

NIST advocates for the use of multiple, complementary evaluation approaches, emphasizing the significance of real-world testing over an exclusive reliance on benchmarks. This is a crucial distinction for compliance functions. The true purpose of TEVV-Athlon should not be to generate a score that offers superficial comfort to management, but rather to produce robust evidence that enhances management’s understanding of the AI system, including its limitations, uncertainties, and residual risks. Consequently, a truly effective evaluation might yield uncomfortable truths, revealing that an AI system performs admirably under certain conditions but lacks sufficient evidence of reliability in others. Such insights, while potentially challenging, are precisely what management needs to make informed decisions.

Shifting the Compliance Paradigm: Beyond "Did It Pass?"

The practical challenge confronting compliance and risk professionals lies in resisting the innate inclination to reduce the complex TEVV-Athlon framework to a simplistic "Did the AI pass?" inquiry. Instead, a more probing question should be: "What did this evaluation actually prove?" This requires moving beyond mere confirmation of testing activities to a deeper understanding of the evaluation’s intent, the rationale behind selected measurements, the underlying assumptions that shaped the assessment, and the specific boundaries of its scope.

This approach aligns perfectly with the structure proposed by NIST. TEVV-Athlon commences with organizational objectives and culminates in the "Synthesize and Interrogate" stage, where results are meticulously interpreted to furnish information that supports strategic organizational decisions. For compliance professionals, this final stage is arguably the most critical. It is not necessary to become a data scientist to effectively challenge an AI evaluation; rather, it is essential to comprehend the reasonable conclusions that management can draw from its findings.

To this end, before relying on the outcomes of an AI evaluation, four fundamental questions should be addressed:

  1. What were the specific, measurable objectives of this evaluation? Understanding the "why" behind the assessment is paramount. Were the objectives clearly defined and aligned with the intended use of the AI system and its associated risks?
  2. How were the chosen metrics and methodologies selected, and what are their limitations? A rigorous evaluation requires a clear justification for the measurement tools employed. It also necessitates an honest acknowledgment of any constraints or potential biases inherent in these methods.
  3. What assumptions underpinned this evaluation, and how might they impact the results? Every evaluation operates under a set of assumptions. Identifying these assumptions is crucial for discerning the validity and applicability of the findings.
  4. What is the scope of this evaluation, and what remains outside of its purview? Clearly delineating what the evaluation covers and, just as importantly, what it does not, is essential for preventing overreliance on incomplete data.

These questions diverge significantly from simply confirming whether required testing protocols were adhered to. They demand an examination of the intricate relationship between the gathered evidence and the business decisions being contemplated. This is particularly pertinent as AI increasingly becomes integrated into third-party products. Organizations may receive evaluation reports from vendors rather than conducting all tests internally. While a vendor might present impressive benchmark results, the core compliance question persists: Do these results provide reliable evidence regarding how our organization intends to deploy and utilize the AI system?

The Enduring Value of TEVV-Athlon: A Tool for Informed Decision-Making

When implemented thoughtfully, the TEVV-Athlon framework can transcend its role as another regulatory control requirement. It possesses the potential to significantly elevate the quality of information upon which management bases its critical decisions concerning AI. Marks emphasizes that his extensive experience with technology risk assessments has consistently revealed a disconnect between the mere completion of an assessment and a genuine understanding of the underlying risks. This realization fuels his appreciation for the NIST AI 200-2 guidance, which underpins TEVV-Athlon, as both interesting and valuable.

NIST is not providing a universal AI testing solution. Instead, TEVV-Athlon offers a structured methodology to help organizations identify what they truly need to know about an AI system and subsequently design an evaluation capable of generating evidence directly relevant to those specific objectives. The framework’s inherent flexibility is not a deficiency that compliance departments need to "fix" but rather a core component of its inherent value.

It is inevitable that pressures will mount to standardize this process. Organizations naturally seek consistency in their compliance efforts. Auditors require verifiable evidence of due diligence. Executives demand clear and comprehensible results. Regulators expect demonstrable proof of appropriate controls. While these expectations are entirely reasonable, the challenge arises when the act of proving process completion overshadows the imperative of understanding what the evaluation is actually communicating. NIST’s explicit reference to Goodhart’s Law serves as a critical admonition. The moment organizations begin prioritizing AI evaluations primarily to achieve acceptable scores, they risk losing sight of the original, fundamental purpose behind selecting those very measurements. This strategic drift can lead to a false sense of security, leaving critical AI-related risks unaddressed. The true power of TEVV-Athlon lies in its ability to foster a culture of critical inquiry and deep understanding, rather than superficial compliance.

By