• Stamholmen 175, 1., 2650 Hvidovre, DK
  • +45 26 80 46 42
  • hello@eywa.dk

Category: Ideas & Perspectives

Abstract / Executive Summary

As artificial intelligence (AI) moves from isolated analytical tools into operational workflows, the central governance question is no longer simply whether an AI system is accurate. It is whether people can understand, supervise, challenge and, when necessary, override AI outputs at the points where those outputs can create meaningful consequences. Human oversight is therefore increasingly understood as a socio-technical capability rather than a final human approval step.

This article examines human oversight in AI workflows through current research on human–AI interaction, automation bias, appropriate reliance and AI risk management, alongside emerging regulatory requirements. It distinguishes between human-in-the-loop, human-on-the-loop and more distributed forms of oversight, and argues that effective oversight depends on the design of the entire workflow: task allocation, escalation mechanisms, interfaces, logging, training, accountability and post-deployment monitoring. Evidence indicates that simply placing a human at the end of an automated process does not guarantee meaningful control. Humans may over-rely on AI recommendations, particularly where systems appear confident or where workload and time pressure reduce independent scrutiny.

The article concludes that effective oversight should be risk-proportionate and designed around the specific decisions, consequences and failure modes of an AI-enabled workflow. This approach is consistent with the NIST AI Risk Management Framework, the OECD AI Principles and the European Union AI Act.

Keywords

Human oversight; artificial intelligence; AI governance; human–AI collaboration; automation bias; appropriate reliance; AI risk management; responsible AI

Main article

1. Introduction

Artificial intelligence is increasingly embedded within workflows rather than used as a standalone technology. An AI system may classify incoming information, identify anomalies, generate recommendations, prioritise cases, draft documents, detect potential risks or initiate downstream actions. In these configurations, the quality of the overall process depends not only on model performance but also on how people interact with the system.

This creates a fundamental governance problem: what does it mean for a human to remain meaningfully in control when an AI system performs part of a workflow?

The answer is more demanding than simply requiring a person to approve the final output. A human reviewer may lack sufficient information to challenge an AI recommendation, may not have enough time to investigate every case, or may become overly reliant on an apparently authoritative system. Research on automation bias has demonstrated that people can accept incorrect recommendations from automated decision aids, while more recent work on generative AI similarly identifies both over-reliance and under-reliance as challenges to effective human–AI collaboration (Parasuraman & Manzey, 2010; Buçinca et al., 2021; Passi et al., 2024).

Human oversight is consequently becoming an important component of AI governance. The OECD’s revised AI Principles call for mechanisms that preserve human agency and oversight, while the NIST AI Risk Management Framework emphasises the need to define and differentiate human roles and responsibilities throughout AI system use.

The European Union’s AI Act provides a particularly concrete regulatory expression of this principle. Article 14 requires high-risk AI systems to be designed so that they can be effectively overseen by natural persons, with oversight measures proportionate to the risks, level of autonomy and context of use.

This article examines what effective human oversight means in practice, why nominal human involvement can fail, and how organisations can design AI workflows in which human judgement remains meaningful.

 

2. Conceptual / Theoretical Background

From human-in-the-loop to human oversight

Human oversight is often described using terms such as human-in-the-loop (HITL), human-on-the-loop and human-in-command. These concepts describe different relationships between people and automated systems.

In a HITL configuration, a human participates directly in the decision cycle. The AI may generate a recommendation, but a person is expected to review it before an action is taken. In a human-on-the-loop configuration, the AI may operate with greater autonomy while a human monitors its behaviour and intervenes when necessary. Human-in-command represents a broader governance concept in which people retain authority over the system’s objectives, deployment and ability to be withdrawn.

These categories should not, however, be treated as sufficient descriptions of effective oversight. NIST notes that human–AI configurations can range from fully manual to fully autonomous, and that the appropriate configuration depends on the system and its context. It also stresses that human roles and responsibilities should be clearly defined and differentiated.

The key concept is therefore meaningful oversight: a human must have the authority, information, competence and practical opportunity to influence the AI-enabled process where intervention is required.

Appropriate reliance rather than maximum trust

An important theoretical shift is from asking whether people trust AI towards asking whether they rely on it appropriately.

Appropriate reliance means accepting an AI recommendation when it is sufficiently reliable while rejecting or investigating it when it is not. Both extremes can be problematic. Over-reliance can allow incorrect AI outputs to pass into decisions, whereas under-reliance can prevent organisations from receiving the benefits of a useful system. Research on generative AI identifies this balance as a central challenge in human–AI collaboration (Passi et al., 2024).

This distinction is particularly important for workflow design. An organisation should not aim to make employees automatically trust AI. It should aim to create conditions in which employees can calibrate their reliance according to the system’s capabilities and the circumstances of each case.

 

3. Literature and Evidence Review

Research on automation provides an important foundation for understanding contemporary AI oversight. Parasuraman and Manzey’s review of automation bias found that people can make both omission and commission errors when automated decision aids are imperfect. Their analysis also found that automation-related complacency and bias are influenced by personal, situational and system characteristics. Importantly, these effects were not restricted to inexperienced users.

More recent human–AI research reaches a similar conclusion. Buçinca et al. (2021), in an experiment involving 199 participants, found that cognitive-forcing interventions could reduce over-reliance on AI recommendations compared with simpler explainability approaches. However, the interventions that reduced over-reliance most effectively also received less favourable subjective evaluations. This illustrates an important design trade-off: making an AI workflow safer may sometimes introduce additional cognitive or operational effort.

Research on explainability also challenges the assumption that providing more explanation automatically produces better oversight. Fok et al. (2024) found that explanations do not consistently produce complementary human–AI performance and noted that explanations can, under some conditions, increase over-reliance.

The implication is significant. Transparency is necessary in many AI workflows, but transparency alone is not equivalent to effective oversight. A technically detailed explanation may be less useful than a well-designed indication of uncertainty, relevant evidence, system limitations or circumstances in which human review is particularly important.

This is reflected in risk-management frameworks. NIST identifies clearly defined human roles, documented oversight processes, operator proficiency and impact assessment as elements of trustworthy AI risk management.

The evidence therefore points towards a socio-technical interpretation of oversight. AI systems do not operate independently of organisational structures. Human assumptions, workload, incentives, interface design and institutional accountability can all influence how AI recommendations are interpreted and acted upon. NIST explicitly identifies cognitive and systemic biases across the AI lifecycle as factors that can affect human–AI interaction.

 

4. Analysis and Discussion

Oversight must be designed into the workflow

A common implementation pattern is:

AI generates output → human reviews → human approves or rejects.

Although apparently straightforward, this structure can conceal several weaknesses. If the human receives thousands of AI-generated recommendations, meaningful individual review may become impractical. If the interface presents the AI recommendation as the default answer, the reviewer may simply confirm it. If there is no mechanism for recording why a recommendation was rejected, the organisation may also lose valuable information about system performance.

Effective oversight should instead be distributed across the workflow.

Workflow stage Human oversight question Example control
Input Is the information entering the system appropriate and sufficiently reliable? Data validation and exception handling
AI processing Is the system operating within its intended scope? Performance monitoring and defined use boundaries
Output Can the human understand the relevance and limitations of the output? Evidence, uncertainty indicators and interpretable outputs
Decision Does the human have genuine authority to challenge the AI? Override and escalation mechanisms
Action Can potentially harmful actions be stopped or reversed? Approval gates, rollback and intervention controls
Monitoring Are failures and human overrides being recorded? Audit logs, incident reporting and periodic review
Improvement Is oversight information feeding back into system governance? Model review, workflow redesign and retraining

This approach is particularly relevant to organisations developing integrated digital platforms. Eywa Systems, for example, describes its technology work in terms of responsible, human-centred technology and emphasises governance, transparency and collaboration in environmental and institutional contexts. Its stated solution areas include integrated digital platforms, environmental data analytics, business intelligence, geospatial systems, climate finance platforms and cybersecurity.

In such data-rich environments, an AI recommendation may become one component within a much larger chain of information and institutional decisions. Oversight therefore needs to extend beyond the model itself to the workflow surrounding it.

Risk should determine the intensity of oversight

Not every AI workflow requires the same level of human intervention. An AI tool that recommends document classifications presents a different risk profile from one that contributes to decisions affecting access to public services, employment or other fundamental interests.

The principle of proportionality is explicitly reflected in Article 14 of the EU AI Act, which requires human oversight measures for high-risk AI systems to be commensurate with risk, autonomy and context.

This suggests a practical hierarchy:

  • Low-impact applications: monitoring and periodic quality checks may be sufficient.
  • Moderate-impact applications: structured review, exception handling and documented escalation may be appropriate.
  • High-impact applications: active human decision authority, auditability, intervention capability and stronger accountability mechanisms may be necessary.

This should not be interpreted as a universal classification system. The appropriate controls depend on the actual use case, applicable law and consequences of failure.

Human oversight requires intervention capability

A human cannot meaningfully oversee an AI system if intervention is technically or organisationally impossible.

NIST identifies the ability to shut down, modify or intervene in systems that deviate from intended functionality as part of practical AI safety approaches.

Intervention capability can take several forms: rejecting an AI recommendation, requesting additional evidence, escalating an unusual case, switching to manual processing, suspending an automated function or reversing an action.

The important point is that oversight must have consequences. If a human can disagree with an AI recommendation but the system automatically proceeds regardless, the human is functioning more as an observer than an overseer.

Interfaces can strengthen or weaken oversight

Human oversight is also an interaction-design problem. A system can technically provide an override button while making it difficult to find, or provide explanations without showing the information needed to evaluate them.

Research on cognitive forcing demonstrates why workflow design matters. Interventions that require people to engage more deliberately with AI recommendations can reduce over-reliance, although they may also increase perceived effort.

Consequently, organisations should consider questions such as:

  • What evidence does the reviewer need to challenge an output?
  • How clearly is uncertainty communicated?
  • Which cases require mandatory review?
  • When should the system automatically escalate?
  • How easily can a person override or stop the process?
  • Are reviewers given sufficient time and authority?
  • Are overrides recorded and analysed?

These questions move oversight from a policy statement into an operational capability.

 

5. Challenges, Limitations, and Counterarguments

The first challenge is that human involvement does not automatically eliminate AI-related risk. Humans themselves bring cognitive biases, institutional assumptions and varying levels of expertise into AI workflows. NIST explicitly recognises these human and systemic factors across the AI lifecycle.

Second, oversight can become burdensome. Requiring human approval for every AI action may create bottlenecks, increase costs and encourage superficial approval behaviour. The evidence on cognitive forcing illustrates this tension between more deliberate review and user experience.

Third, explanations have limitations. There is no strong basis for assuming that a more detailed explanation will always lead to better decisions. Evidence remains mixed regarding the ability of explainability techniques to generate complementary human–AI performance.

Fourth, accountability can become ambiguous when responsibility is distributed between developers, deployers, operators and managers. Recent OECD work on AI in the workplace identifies unclear accountability as a concern among organisations using algorithmic management tools.

Finally, empirical research remains uneven across domains. Controlled experiments can reveal important behavioural mechanisms, but their findings may not directly predict behaviour in complex operational environments involving high workloads, institutional incentives and real consequences. Evidence on particular oversight mechanisms should therefore be applied with attention to context.

 

6. Implications

For organisations deploying AI, human oversight should be treated as part of workflow architecture, rather than as an administrative control added after implementation.

Management teams should first identify where AI outputs can materially affect people, resources, safety, compliance or organisational decisions. Those points should then be mapped against appropriate human responsibilities and escalation mechanisms.

Technology teams should design interfaces around the information required for meaningful review rather than simply displaying AI outputs. Logging should capture not only system decisions but, where proportionate and lawful, human interventions and overrides. Such information can support monitoring and continuous improvement.

Governance teams should define who has authority to challenge, suspend or modify an AI-enabled process. This is particularly important where responsibility crosses organisational boundaries.

Training should also extend beyond technical operation. Users need to understand the system’s intended purpose, known limitations, failure modes and appropriate conditions for reliance. NIST identifies operator proficiency and documented oversight processes as elements of AI risk management.

For public institutions and organisations working with environmental, geospatial or other high-value data, the implications are broader. AI may become embedded in systems that consolidate information and support evidence-based decision-making. Eywa Systems’ stated methodology, for example, places assessment, planning, development and integration, monitoring, optimisation and ongoing support within a continuous process. This type of lifecycle perspective is compatible with the broader principle that oversight should continue after deployment rather than end at system launch.

Regulation is also reinforcing this direction. The EU AI Act’s Article 14 establishes human oversight as a specific requirement for high-risk AI systems, while the OECD Principles frame human agency and oversight as part of trustworthy AI.

7. Conclusion

Human oversight is not simply the presence of a person somewhere within an AI workflow. Effective oversight requires that people have the authority, information, competence and practical opportunity to influence the system when necessary.

The available evidence indicates that humans can over-rely on automated recommendations, including when those recommendations are incorrect. It also indicates that explanations alone do not reliably solve this problem. Effective human–AI collaboration therefore depends on broader workflow design: appropriate allocation of responsibilities, risk-proportionate review, escalation pathways, intervention mechanisms, monitoring and organisational accountability.

The strongest practical conclusion is consequently not that every AI system should have a human approving every output. Rather, the level and form of human oversight should correspond to the system’s autonomy, potential consequences and operating context.

As AI becomes more deeply integrated into digital platforms and institutional processes, this distinction will become increasingly important. Responsible AI governance will depend not merely on building capable models, but on designing workflows in which humans remain capable of understanding, questioning and directing the systems they use.

 

References

  1. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), 1–21.
  2. European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). EUR-Lex.
  3. Fok, R., et al. (2024). In search of verifiability: Explanations rarely enable complementary performance in AI-advised decision making. AI Magazine. https://doi.org/10.1002/aaai.12182
  4. National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0).S. Department of Commerce.
  5. Organisation for Economic Co-operation and Development. (2021). Tools for trustworthy AI: A framework to compare implementation tools for trustworthy AI systems. OECD Digital Economy Papers, No. 312
  6. Organisation for Economic Co-operation and Development. (2024). Using AI in the workplace: Opportunities, risks and policy responses. OECD Artificial Intelligence Papers, No. 11.
  7. Organisation for Economic Co-operation and Development. (2024). Revised Recommendation of the Council on Artificial Intelligence.
  8. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410.
  9. Passi, S., Dhanorkar, S., & Vorvoreanu, M. (2024). Appropriate reliance on Generative AI: Research synthesis. Microsoft Research Technical Report MSR-TR-2024-7.