The evolving landscape of artificial intelligence (AI) within enterprise operations presents a new and critical challenge for contract law. Beyond the traditional question of financial liability when something goes awry, the next generation of AI contracts must definitively address: "Who can prove what happened?" This imperative is driven by the increasing reliance on AI in decision-making processes, particularly in sensitive areas like employment, and the complex evidentiary trails that these systems create.

When an AI-influenced decision, such as a job rejection, faces a legal challenge, organizations may find that their existing contracts, while robust in areas like audit rights, security commitments, and indemnification, fall short when it comes to providing the irrefutable evidence needed to explain the AI’s reasoning. This gap in accountability is becoming increasingly apparent as high-profile litigation unfolds, spotlighting the intricate nature of AI systems and the fragmented ownership of the data and processes that underpin their decisions.

A prominent illustration of this burgeoning problem is the ongoing litigation against Workday, a widely adopted platform for human resources, finance, and IT functions. In the case of Mobley v. Workday, plaintiffs allege that Workday’s algorithmic applicant-screening tools engaged in discriminatory practices based on protected characteristics. While Workday vehemently disputes these allegations and no finding of discrimination has been made by the court, the legal proceedings have illuminated critical issues surrounding the explainability and auditability of AI in hiring.

The legal battle took a significant turn in June when a judge permitted several claims under the California Fair Employment and Housing Act (FEHA) and the Americans with Disabilities Act (ADA) to proceed. This ruling, however, focused on the adequacy of the legal pleadings and did not address the ultimate merits of the case. Nevertheless, it underscored the potential for AI-driven employment decisions to face intense legal scrutiny.

For compliance leaders, perhaps the most consequential development emerged from a May discovery order. This order revealed the intricate ways in which evidence related to AI bias testing can become dispersed across an AI vendor, its customers, and their respective legal counsel. The court determined that certain bias-testing data generated by Workday was protected by attorney-client privilege, as Workday’s lawyers had curated the underlying data and utilized the results to furnish legal advice. Furthermore, the court declined to compel Workday to produce customer applicant data, citing the plaintiffs’ failure to establish that Workday legally controlled such data under federal discovery rules.

Simultaneously, the court mandated Workday’s production of specific EEO-1 and Office of Federal Contract Compliance Programs (OFCCP) documents. These were deemed relevant to Workday’s potential awareness of demographic disparities within the applicant pool. While this order does not establish that Workday failed to maintain legally mandated records, it starkly exposes a broader enterprise risk: the critical records necessary to evaluate an AI-influenced decision may reside with various parties who do not necessarily share uniform data retention obligations, access rights, or litigation strategies. This fragmentation can create insurmountable hurdles in reconstructing the exact sequence of events and the rationale behind an AI’s output.

The Limitations of Bias Audits in Decision Reconstruction

Organizations often refer to "auditability" as a singular control mechanism when discussing AI systems. However, this terminology can be misleading. A bias audit, in its typical form, primarily evaluates outcomes across an entire population of users or data points. Such audits may involve comparing selection rates among different demographic groups, analyzing error rates, or testing whether a system generates statistically significant disparities. While this type of assessment is invaluable for identifying systemic risks and potential biases embedded within an AI model, it often falls short of explaining the specific circumstances and rationale behind an individual decision concerning a single applicant.

Decision reconstruction, in contrast, seeks to answer a more granular set of questions:

  • What specific data points influenced the AI’s output for this individual?
  • What version of the AI model was used?
  • What configurations and parameters were active at the time of the decision?
  • What were the pre-processing steps applied to the input data?
  • Were there any human interventions or overrides in the decision-making process?
  • What were the specific validation and testing results for the model version in use?

Without answers to these questions, an AI vendor might successfully demonstrate that its system performs adequately at a population level, yet be incapable of recreating a specific decision made by that system. Conversely, an employer might diligently retain an applicant’s submission and the final outcome but lack access to crucial vendor-provided information, such as the specific model version employed, its historical configuration, or the evidence supporting its validation. This disconnect highlights a critical flaw in many current AI implementations and contractual agreements.

Legislation in various jurisdictions is beginning to address this gap. A law enacted in New York City, for example, mandates that employers utilizing covered automated employment decision tools must secure an independent bias audit, publish a summary of its findings, and notify affected candidates. It also requires disclosure of information pertaining to the data collected and the employer’s data-retention policies. While these provisions enhance transparency and accountability, they do not automatically guarantee the availability of every technical record necessary for the reconstruction of individual decisions.

California’s regulatory framework takes a more expansive approach to data retention. The state’s automated-decision regulations stipulate that employers and other covered entities must retain employment records, including data generated by automated decision systems, for a minimum of four years. Providers of these automated decision systems are also obligated to retain relevant records for at least four years after the system was last utilized by the employer or covered entity. However, the critical distinction remains: data retention does not equate to assured production. A record can exist without the employer possessing a contractual right to promptly obtain it, interpret its complexities, or furnish it during an investigation or legal proceeding. This highlights that the mere existence of data is insufficient without the contractual mechanisms to access and utilize it effectively.

Integrating an AI Evidence Schedule into Contracts

To bridge this evidentiary chasm, enterprises procuring AI hiring products must incorporate a comprehensive "evidence schedule" into their vendor agreements. This schedule should meticulously delineate which party is responsible for creating, controlling, retaining, and producing each specific category of record associated with the AI system’s operation.

At a minimum, such a schedule should address the following seven critical areas:

  1. Input Data: Clearly define the source, format, and ownership of all data fed into the AI system, including applicant information, relevant metadata, and any pre-processing applied.
  2. Model Versioning: Specify how different versions of the AI model are tracked, stored, and identified, ensuring that the exact model used for a particular decision can be pinpointed.
  3. Configuration Parameters: Detail the specific settings, parameters, and rules that govern the AI model’s operation at any given time, and establish procedures for logging and preserving these configurations.
  4. Decision Logs: Mandate the creation and retention of detailed logs that capture the inputs, internal processing steps, and final outputs of the AI for each individual decision made.
  5. Validation and Testing Data: Outline the requirements for documenting the testing methodologies, datasets, and results used to validate the AI model’s performance, fairness, and accuracy.
  6. Human Review Records: If human oversight or intervention is part of the decision-making process, specify how these interactions, including any modifications or overrides, are recorded and preserved.
  7. Audit Trail of System Access and Changes: Establish a clear audit trail detailing who accessed the AI system, when they accessed it, and any modifications made to the system or its data.

Beyond defining data responsibilities, AI contracts should also establish clear escalation deadlines for the production of evidence and specify the circumstances under which an employer may suspend the use of automated screening tools without breaching minimum-volume or exclusivity commitments. This provides a mechanism for recourse and risk mitigation when issues arise.

Standards as Frameworks, Not Litigation Safe Harbors

While international and federal standards can offer valuable frameworks for structuring AI governance and risk management controls, they should not be misconstrued as substitutes for robust legal protections or litigation safeguards.

Standards like ISO/IEC 42001 provide a comprehensive set of requirements for establishing an organizational AI management system. This includes directives on governance, risk management, transparency, traceability, performance evaluation, and continuous improvement. However, adherence to ISO/IEC 42001 does not certify the legality of a particular AI-driven hiring decision. Similarly, ISO/IEC 42005 offers a complementary framework for documenting AI system impact assessments, assisting organizations in evaluating both intended and unintended consequences throughout the AI system’s lifecycle.

The National Institute of Standards and Technology (NIST) AI Risk Management Framework (RMF) also recommends detailed documentation of intended uses, testing assumptions, evaluation techniques, data practices, and system performance. Given that NIST is actively revising its framework, organizations are advised to monitor these updates rather than rigidly freezing their contractual agreements around a single, potentially outdated, framework version.

The ultimate compliance objective is not merely to insert a list of recognized framework names into a contract. Instead, it is to translate the legal requirements and the principles embedded within these standards into enforceable contractual obligations regarding evidence. Indemnification clauses, while still possessing value, represent a financial remedy rather than a comprehensive accountability architecture. An indemnity clause can effectively cover legal expenses and damages, but it cannot magically recreate an AI-influenced hiring decision when both the employer and the vendor have failed to adequately preserve the crucial elements of the model, its configuration, its output, and any human review involved. The focus must shift from financial recourse to the fundamental ability to explain and justify the actions of AI systems when called into question.

By