By Madhu Murty, Co-Founder & Head of India Operations, QualiZeal
As enterprises move from experimenting with AI to deploying autonomous AI agents, accountability is emerging as one of the biggest governance challenges. Unlike traditional AI systems that generate predictions or recommendations, agentic AI can independently plan tasks, invoke tools, make decisions, and execute complex workflows with minimal human intervention. That shift fundamentally changes how organisations should think about governance.
The question surfacing in every serious conversation around enterprise AI today is this: how do you validate a system that constructs its own execution path in real time? Not a model that uses a contained path to return an output. Not a deterministic workflow that can be reviewed step by step. But an agent that reasons, selects tools, creates subtasks and executes them—often faster than a human reviewer can meaningfully follow. And suppose the systems involve multiple agents; it becomes hard to determine which agent committed the error, or if the error surfaced from cross-agent interaction.
Most AI governance frameworks acknowledge this challenge. The EU AI Act requires oversight measures proportionate to an AI system’s level of autonomy. NIST’s AI Risk Management Framework assigns organisations clear responsibility for managing AI risks, while Singapore’s governance framework for agentic AI places meaningful human accountability at its core. They all point towards the same objective: ensuring humans remain accountable for AI-driven outcomes. What they are less prescriptive about is how organisations should achieve that when AI systems themselves are increasingly autonomous.
The challenge is not fundamentally regulatory. It is architectural. Agentic AI replaces single-step actions by executing a series of autonomous actions. There is no singular path from input to output—it is a dynamic, branching workflow whose exact route changes from one execution to the next. And every agent in the enterprise environment uses credentials, holds permissions, and is authorized to make decisions in the real-time, Policy defines expectations, governance defines decision rights, but architecture ultimately determines whether accountability can be enforced at runtime.
That distinction matters because existing governance models were largely designed for AI systems that assist human decision-making, not ones that independently orchestrate actions. In fact, when NIST’s Center for AI Standards and Innovation launched its AI Agent Standards Initiative in 2026, it acknowledged that existing frameworks do not adequately distinguish between systems that merely recommend actions and those capable of autonomously executing complex workflows. While standards continue to evolve, enterprises are already deploying these systems into production without having real-world insight into how these systems perform in real-world scenarios and under complex conditions.
The harder question is: what happens when human oversight itself begins to fail?
Automation bias is not a theoretical concern. Decades of research across aviation, healthcare, criminal justice and financial services have consistently shown that people place disproportionate trust in automated systems, often reducing their own scrutiny as confidence in the technology increases. One 2025 study involving 450 clinicians working alongside intentionally biased diagnostic AI found that diagnostic accuracy dropped from 73% to below 62%—not because clinicians lacked expertise, but because the AI subtly influenced their judgement. The more reliable a system appears, the easier it becomes for oversight to turn into rubber-stamping rather than genuine review.
The Dutch childcare benefits scandal offers a similar lesson at an institutional level. Formal oversight mechanisms existed throughout the process, yet systemic failures continued unchecked. Amnesty International later described that oversight as “formal but ineffective”—perhaps the clearest illustration that simply placing humans in an approval chain does not automatically create meaningful accountability.
This is why organisations need to move beyond policy-led governance towards architecture-led governance. Policy remains essential, but sustainable accountability depends on governance being translated into architectural controls.
Rather than relying on reviewers to identify every problematic decision after it occurs, enterprises should design systems that minimise unsafe behaviour from the outset. Agent permissions should be tightly scoped. Tool access should follow the principle of least privilege. High-impact actions should include built-in checkpoints, while comprehensive audit trails should capture not only what happened, but why decisions were made, which tools were invoked, and where human intervention occurred. These controls should be complemented by continuous monitoring and evaluation to ensure governance remains effective as agent behaviour evolves over time.
Importantly, governance itself should become measurable. Singapore’s updated guidance on agentic AI recommends monitoring human override rates and response times as governance metrics, recognising that oversight itself can develop failure modes. Measuring how often humans intervene—and how effectively they do so—provides a far more meaningful indicator of accountability than simply documenting who approved a decision.
From our experience working with enterprise AI assurance programmes, the strongest governance models are not those that depend on reviewers catching every mistake at the end. They are the ones that reduce the opportunity for mistakes through deliberate system design. That also means deciding the assurance dimensions that target distinct risk areas and provide continuous validation until they reach production. Accountability becomes an inherent property of the architecture rather than a responsibility assigned after deployment.
As organisations accelerate their adoption of agentic AI, governance will increasingly depend as much on engineering decisions as policy decisions. Regulations will continue to evolve, but sustainable accountability will come from designing systems that are safe, constrained and transparent long before they reach production.
For agentic AI, accountability is not simply a policy to document. It is an architectural capability that must be engineered from the very beginning and continuously validated throughout the AI system’s operational lifecycle.



