The ubiquitous answer to how a company governs artificial intelligence often circles back to a seemingly fundamental safeguard: a human makes the final decision. This assertion, frequently voiced across diverse industries, positions human oversight as the ultimate arbiter, the last line of defense against algorithmic error. However, as highlighted by Adnan Masood, a professor and AI executive, this notion of oversight, while intuitively reassuring, may be as hollow as its absence if not rigorously examined and implemented. The efficacy of human-in-the-loop systems is increasingly being questioned, revealing that the mere presence of a person in the process does not guarantee genuine control or prevent the insidious creep of automation bias.

A stark illustration of these limitations emerged from a 2023 study conducted by researchers at the University of Cologne. In this experiment, a cohort of radiologists was tasked with reviewing 50 mammograms each, aided by an AI assistant. Initially, the AI system demonstrated commendable accuracy, fostering a sense of trust among the participating professionals. This initial success, however, laid the groundwork for a subsequent and alarming decline in diagnostic precision when the AI began to produce erroneous results. For the least experienced professionals, accuracy plummeted from approximately 80% to under 20%. Even seasoned radiologists, boasting over a decade of experience, saw their accuracy rates fall dramatically from 82% to roughly 46%. This significant drop underscores a critical vulnerability: over-reliance on AI, even when initially accurate, can erode human judgment and diagnostic skills.

This experiment provides a compelling real-world portrait of the inherent limitations in human-in-the-loop (HITL) approaches to AI governance. After numerous governance reviews, Masood identifies "automation bias"—the tendency to uncritically accept automated suggestions without verification—as the most frequently cited control mechanism in the field, yet paradoxically, the least rigorously tested. The core issue, as the University of Cologne study suggests, is that a person’s presence within a decision-making workflow does not automatically equate to that person being in actual control of the outcome. The human acts as a validator, but if the system is trusted implicitly, the validation process can become a mere formality.

The Erosion of Verification: Why Approval Becomes Automatic

The cycle of trust and subsequent erosion of critical assessment is a well-documented phenomenon. When AI models consistently deliver accurate results, as seen in the initial stages of the mammogram review, users develop confidence in the system. This burgeoning trust naturally diminishes the instinct to meticulously verify each output. Research has corroborated this, indicating that prolonged reliance on accurate AI can lead to a degradation of underlying human skills. This is particularly concerning in fields where human expertise is critical and nuanced, such as medical diagnostics or complex financial analysis.

Adding to this challenge is the performance pressure inherent in many professional environments. Reviewers are often evaluated based on their throughput—the volume of work they complete within a given timeframe. In such a metrics-driven culture, the act of verifying AI output can be perceived as an impediment to efficiency, registering as underperformance. This creates a perverse incentive: faster processing, even at the cost of thoroughness, is rewarded.

Furthermore, the practical mechanics of AI interaction often favor automated acceptance. Typically, agreeing with an AI’s decision requires a single click, a seamless integration into the workflow. Conversely, disagreeing with the AI often necessitates a more involved process, potentially requiring the completion of a detailed justification form, and in some cases, even a supervisor’s signature. This asymmetry in effort subtly nudges users towards acceptance.

The allocation of blame also plays a significant role in perpetuating automation bias. If an error occurs when a human agrees with the AI, the responsibility is often shared between the human and the system. However, if a human overrides the AI and an error subsequently emerges, the sole burden of responsibility falls upon the individual. This lopsided accountability structure creates a powerful disincentive to challenge the AI, further solidifying the tendency to accept its recommendations.

The real-world consequences of neglecting human judgment in automated systems are starkly illustrated by Australia’s Robodebt program. This initiative automated the process of welfare debt collection, effectively removing the human checks that had previously served as a crucial safeguard against errors. The program was met with widespread condemnation, culminating in a commission’s scathing report, court rulings deeming the debts unlawful, and substantial financial settlements. By June 2026, a court approved a $549 million settlement, in addition to an earlier $112 million award, as a consequence of the flawed automated system. This case serves as a potent reminder that while AI can enhance efficiency, its implementation must be accompanied by robust human oversight that accounts for potential algorithmic inaccuracies and the fallibility of automated processes.

Courts are increasingly recognizing the limitations of AI systems with weak human oversight. In 2023, a European court ruled that when lenders rely predictably on credit scores to make lending decisions, the credit score itself constitutes an automated decision under Article 22 of the General Data Protection Regulation (GDPR). This landmark decision signifies a growing legal and regulatory awareness of the need for meaningful human intervention and transparency in automated decision-making processes, particularly those with significant implications for individuals.

The metric of agreement rate, often cited as a measure of AI system performance, can be misleading. A 99% agreement rate might indicate a highly accurate model that is being correctly validated, or it could simply reflect a "rubber stamp" effect where human reviewers are habitually agreeing with the AI without critical evaluation. The rate alone does not differentiate between these scenarios. The inherent value of human review is measurable, yet this critical assessment is frequently overlooked in favor of simplistic metrics.

Auditing the Human Layer: The Imperative of Testing Oversight

Since 2011, banks in the United States have been operating under SR 11-7, the Federal Reserve’s guidance on model risk management. This regulation mandates critical review of models by qualified individuals who possess both the authority and the incentive to challenge and potentially alter an AI system’s output. The principle embedded in SR 11-7—that human review must be robust, empowered, and incentivized—should be extended to the review layer of AI governance itself. The oversight mechanism, not just the AI, requires rigorous scrutiny.

In the realm of internal audit, most AI oversight programs are designed to document the design effectiveness of controls. This means that the controls exist on paper, outlining the intended safeguards. However, there is often a lack of testing for operating effectiveness—whether these controls are actually functioning as intended in practice. This gap is unacceptable in traditional financial reconciliation processes, and it should be equally unacceptable for the controls that stand between a potentially flawed algorithm and a customer or critical decision.

To effectively audit the human layer, a multi-faceted testing program is essential. A foundational approach involves a three-way comparison: evaluating the performance of reviewers working independently, the AI model operating alone, and the combined approach of reviewers working with the AI. The HITL approach should only be maintained if it demonstrably outperforms the individual components.

One practical method for testing the effectiveness of human oversight is to introduce known errors into live operational queues, analogous to how security teams conduct phishing simulations. The rate at which these seeded errors are caught would become a primary oversight metric. Key performance indicators should include:

  • Override Precision: This measures how often a reviewer’s disagreement with the AI proves to be correct. High precision indicates that when a human intervenes, it is generally for valid reasons.
  • Override Recall: This assesses the proportion of actual AI errors that reviewers successfully identify and flag. High recall suggests that human reviewers are adept at catching the system’s mistakes.

Further enhancements to this testing methodology include:

  • Minimum Dwell Times: Establishing minimum time thresholds for reviewers to spend on a case, ensuring that sufficient time is allocated for genuine verification and preventing superficial reviews.
  • Evidence Logging: Recording whether reviewers actually access and examine the underlying evidence before making a decision, rather than relying solely on the AI’s summary or recommendation.
  • Equal Effort for Agreement and Disagreement: Designing systems where the effort required to agree with an AI is comparable to the effort required to disagree, thus removing the procedural bias towards acceptance.
  • Pre-Decision Judgment Capture: Where feasible, capturing the reviewer’s independent judgment before revealing the AI’s answer. This prevents the AI’s output from unduly influencing the human’s assessment.

Responsibility for AI oversight should be clearly delineated across the three lines of defense commonly employed in risk management:

  1. First Line of Defense (Business Operations): The business unit utilizing the AI is responsible for the day-to-day operation and review of AI outputs.
  2. Second Line of Defense (Risk Management): Risk and compliance functions should be responsible for designing and implementing the risk instruments and controls that govern AI usage, including the testing methodologies.
  3. Third Line of Defense (Internal Audit): Internal audit should independently test the effectiveness of AI oversight controls, employing seeded cases and other rigorous methodologies to provide assurance to senior management and the board.

In scenarios where per-case review is inherently insufficient—due to machine speed, high volume, or the irreversible nature of an action—the practice of superficial review should be abandoned. Instead, the resources allocated to this ceremonial oversight should be redirected upstream. This involves strengthening pre-deployment evaluation of AI models, implementing continuous monitoring of their performance in live environments, and establishing robust appeals processes that genuinely function. Regulators in Europe, with the burgeoning AI Act, are moving towards this more sophisticated, test-driven approach to AI governance.

The implications for risk officers are clear. Those who proactively implement rigorous testing of the human layer in AI governance now will be well-positioned to defend their methodologies and demonstrate the efficacy of their controls in the future. Conversely, those who defer this critical work will likely find themselves in a defensive posture, attempting to justify assumptions and practices that have not been adequately validated when challenges inevitably arise. The future of AI governance hinges not on the mere presence of a human, but on the demonstrable quality and accountability of their engagement with artificial intelligence.

By