Skip to content

Human oversight of artificial intelligence: from declared principle to measurement

Article 14 of the European Union's Artificial Intelligence Act requires high-risk systems to allow effective human oversight. Most organisations can describe their oversight procedure. Very few can show evidence that the oversight actually took place. This page explains the difference and how it is measured.

Reference page maintained by Humanning Lab · Beatriz López López

What Article 14 requires

Regulation (EU) 2024/1689, known as the EU AI Act, entered into force on 1 August 2024 and its obligations apply in stages. Those affecting systems classified as high-risk come into effect between 2026 and 2027, depending on the category of the system.

Article 14 states that high-risk AI systems must be designed and developed so that they can be effectively overseen by natural persons while in use. It is not enough to have someone nominally responsible: the rule requires that person to be able to understand the system's capacities and limitations, detect anomalies, interpret its output correctly, decide not to use it and, where appropriate, stop its operation.

The text also names automation bias explicitly: the tendency to rely automatically on a system's output, especially when that system produces recommendations that look precise. The rule requires the organisation to be aware of that bias and to act on it.

That is where the practical problem appears. Automation bias is not a procedure: it is a behaviour. And a behaviour is not evidenced with an internal policy document, it is evidenced by observing it.

Declared oversight and measured oversight

When an organisation describes its AI governance, it usually presents three things: an org chart with assigned roles, a written review procedure and recorded training. All three are necessary. None of the three proves that oversight took place.

Declared oversight answers the question: what should happen when the system issues a recommendation? Measured oversight answers a different one: what actually happened when the system issued the recommendation?

The distance between the two is where risk accumulates. Someone approving 99% of the model's recommendations in forty seconds is complying with the procedure and, at the same time, not overseeing anything. The file will be complete. The effective oversight will not.

This distinction runs through all of Humanning Lab's work in other domains: we measure what is lived, not what is declared.

How human oversight is measured

Human oversight can be observed because it leaves a behavioural trace. It is measured on the decisions that already occur in the operation, without changing the model or interfering with the workflow.

What is observed are regularities: in what proportion of cases the person departs from the system's recommendation and under what conditions; how much that behaviour varies across teams, shifts and workload levels; which case types concentrate automatic acceptance; what distinguishes the decisions where the person did intervene; and whether that intervention holds over time or decays after deployment.

The result is not a maturity score or a self-assessment. It is a set of behavioural indicators, dated and traceable, describing how oversight is being exercised in that specific organisation, with that specific system.

That set is what an organisation can put in front of a supervisor or an auditor when asked to demonstrate, not assert, that human oversight is effective.

Where the instrument comes from

The instrument does not come from compliance consulting, but from the biology of cognition developed by Humberto Maturana and Francisco Varela.

That framework provides an operational premise: a person's behaviour is not explained by the instruction they receive, but by the structure from which they receive it. A rule does not determine a behaviour; it triggers it, and what happens next depends on the structure of the person receiving it and on the context in which they operate.

Applied to AI governance, this explains why well-written oversight policies produce such different results across organisations, and why the only serious way to know what is happening is to observe behaviour in its operating context.

Where it applies: banking, insurance, healthcare and recruitment

The Regulation classifies as high-risk, among others, systems involved in decisions with direct impact on people's rights and opportunities. These are the four contexts where the instrument applies most clearly.

Banking

The analyst receives a recommendation from the model and decides. The question a supervisor can ask is how many of those decisions were actually reviewed and what distinguishes the ones that departed from the model. Without behavioural measurement, the answer is an estimate.

Insurance

The same pattern, with an aggravating factor: actuarial models have been part of professional practice for decades, so trust in their output is more naturalised and automation bias is harder to detect from the inside.

Healthcare

Human oversight is particularly sensitive here because the professional works under time pressure. Measuring under which conditions review holds and under which it degrades is safety information, not only compliance information.

Recruitment

This is one of the uses explicitly listed as high-risk. The volume of applications pushes towards automatic acceptance of the filter, precisely where the impact on people is direct and where traceability of the human decision is required.

This is a general reference framework, not legal advice. The classification of a specific system depends on its purpose and on how it fits the annexes of the Regulation, and must be determined by the organisation together with its legal counsel.

Frequently asked questions

Who is responsible for human oversight: the developer of the system or the organisation using it?
The Regulation distributes obligations. The provider must design the system so that human oversight is possible, and the deployer must assign that oversight to people with the necessary competence, training and authority, and ensure they can exercise it. Behavioural measurement matters above all to the deployer, because that is where oversight either happens or does not.
What exactly does an auditor or a supervisor ask for?
They ask to be able to verify. The difference between an oversight policy and a measurement of oversight is that the second one can be checked independently: dated indicators, on real decisions, with a described and reproducible methodology.
Does this require modifying the model or the AI system?
No. I measure decision behaviour, not the model. There is no need to intervene in the system or to alter the teams' workflow.
How is this different from an algorithmic audit or a model bias test?
An algorithmic audit examines the system: its data, its performance, its statistical biases. This measurement examines the human who decides on the system's output. They are complementary and answer to different articles of the Regulation.
Does this assess the people who review?
No. What is measured is the regime of the deployment: whether the person deciding has the time, the power and the consequences to do it for real. The records arrive anonymised by the organisation itself and each reviewer appears as a code, never as a name.
How does an organisation start?
With a bounded pilot on a specific deployment. The scope, the period and the volume of decisions are agreed before starting. The indicators are not: they are fixed by the method, and they are the same for everyone. That is the condition for the result to mean anything.

References and published work

Propose a pilot

If your organisation operates AI systems in any of these contexts and needs to demonstrate effective human oversight, the first step is a conversation about the specific use case. I do not sell a compliance opinion: I measure decision behaviour and deliver the indicators that describe it.

Write to propose a pilot