Most conversations about AI agents start with whether to deploy them. That question is usually behind the organisation asking it.

Agents have already arrived, through the same doors AI arrived through the first time. A vendor added an agentic feature to software you already licence. A team built something useful on a framework nobody else in the business uses. Someone in operations wired a tool into a workflow because it saved three hours a week and nobody told them not to.

The organisations with an agent problem are rarely the ones planning an agent strategy. They are the ones that already have forty and believe they have four.

That gap between what is running and what is known is where the governance problem actually lives, and it is a different problem from the one most organisations are preparing for.

What changes between automation and autonomy

Enterprises have automated work for decades and have good controls for it. Understanding why those controls do not transfer requires being precise about what changed.

Traditional automation is deterministic. A robotic process automation bot follows a defined path. Read this field, compare it to that value, if it matches take this action, if not raise an exception. You can read the path. You can test every branch. It does exactly what it was built to do, every time, until something in its environment changes.

An agent is given a goal and works out how to reach it. The path is not defined in advance, which is the entire point. That flexibility is why agents handle work that defeated deterministic automation: messy inputs, exceptions, tasks requiring judgement about which system to use next.

The consequence is the sentence to take from this article. Deterministic automation fails loudly and stops. Agentic automation fails quietly and improvises.

When an RPA bot hits something it was not built for, it throws an error and halts. Someone gets an alert. The process stops, which is inconvenient and entirely safe.

When an agent hits something it was not built for, it tries something else. That is what it was designed to do. The most widely reported incident of this kind in 2026 involved a coding agent that met a minor authentication error, decided to resolve it rather than stop, located a credential carrying far broader authority than its task required, and destroyed a production environment in seconds. Nothing malfunctioned. The agent pursued its goal using access it had been granted.

That is the transition from automation to autonomy in a single event. Not a system that broke, but a system that succeeded at the wrong thing.

Why sprawl is a governance problem

Organisations tend to treat the agent inventory question as housekeeping. Four consequences follow directly from not knowing what you are running.

You cannot risk-rate what you cannot see. Every governance framework in use, including the CBUAE expectations for licensed institutions, starts from an inventory with risk ratings attached. An agent estate nobody has catalogued has no ratings, which means no proportionate controls, which means the agent with access to your payments system is governed exactly as lightly as the one that formats meeting notes.

Per-framework governance does not scale. Agents get built on different platforms by different teams. If governance is implemented separately inside each framework, you have as many governance models as you have frameworks, none of them comparable, and a risk committee that cannot get a single answer about the estate.

Permissions accumulate invisibly. This is the failure mode behind most agentic incidents. Access granted for a build phase, a migration, a one-off integration, and never revoked because revocation was nobody's job. Without a view across the estate, nobody is positioned to notice.

Cost becomes invisible until it is large. Agents consume tokens and make model calls, and unlike a human process there is no natural ceiling on how much. Runaway consumption appears in the incident data alongside the destructive actions: infinite loops, agents retrying indefinitely, workloads nobody attributed to a business owner until the invoice arrived.

That last one is underdiscussed and matters for a reason beyond finance. The instrumentation that tells you what an agent is doing is the same instrumentation that tells you what it is costing. Control and cost control turn out to be the same problem, which means the business case for agent governance has a line in it that most governance business cases do not.

The regional timeline

This cannot wait for frameworks to settle, because agents are being deployed in this market faster than almost anywhere.

In April 2026 the UAE Cabinet approved a programme to move fifty per cent of federal government sectors, services and operations to agentic AI within two years. By June, fifty federal entities were on a ninety-day sprint, each required to take an agentic service from selection through design to implementation planning. Eighty thousand federal employees are being trained. The Presidential Court set itself seventy-five per cent.

The governing principle the programme adopted is worth borrowing regardless of sector: human leads, AI enables.

For the private sector, the relevant instrument remains the CBUAE Guidance Note, which gives licensed institutions a usable vocabulary for the autonomy question. Its human oversight section defines three models: human-in-the-loop, human-on-the-loop, and human-out-of-the-loop, with the last permitted only for low-risk, non-material processes with appropriate controls in place.

That is the decision every organisation deploying agents has to make explicitly, for every agent, and record. The same guidance requires institutions to retain at all times the clear and immediate ability, through human intervention, to cease using a deployed AI system.

Applied to an agent estate, that is a kill switch requirement. Testing whether you have one is a useful afternoon's work.

What governing an agent estate requires

Before naming any product, it is worth setting out what the job is. Five functions, and they hold regardless of what you buy.

Discovery across frameworks. You need to find agents wherever they run, including ones built on platforms your governance team has never touched and ones embedded in software you bought. An inventory that only covers agents built in house is not an inventory.

Identity and scoped permissions. Each agent needs a distinct identity rather than a shared service account, a named human owner, and permissions scoped to what it actually does rather than what it needed once.

Policy enforced at runtime. Design-time approval does not constrain an agent that decides its own path. Boundaries have to hold while it runs.

Evaluation before publication. Agents should be tested against defined measures before reaching production, and the measures that matter are behavioural: does it complete the journey, does it call the right tools, are its answers relevant.

Cost and consumption attribution. Visibility into which agents are driving consumption, attributable to a business owner.

IBM, Microsoft, Google and Databricks are all converging on this same set of primitives. When four platform vendors with different architectures arrive at the same list independently, it usually means the list reflects the shape of the problem rather than any one company's product strategy. Agent governance capability is becoming a standard layer of the enterprise stack rather than a differentiator.

Two layers, integrated

IBM's approach splits this across two systems with different jobs, and the split is instructive even for organisations that end up buying something else.

watsonx Orchestrate is the operating layer. Its agentic control plane provides a centralised place to observe, govern and optimise agents across the enterprise regardless of where they were built or where they run. That includes cross-platform discovery, which scans connected environments and brings agents built elsewhere under the same control plane rather than requiring teams to rebuild governance per framework. It supports agents built natively, in Langflow and LangGraph, on the open A2A protocol and in environments such as Amazon Bedrock.

Operationally it handles runtime policy enforcement, so agent behaviour stays within defined boundaries while the agent is running. Evaluation before publication against completion, tool-call accuracy and answer relevance. Credential health monitoring, which catches broken or missing connections before they cause failures. A governed catalogue of prebuilt and partner agents with lifecycle controls. And visibility into token usage and LLM calls, which is where the cost question gets answered.

watsonx.governance is the enterprise risk layer, and the integration extends the picture rather than filling a gap. Orchestrate governs the agent estate operationally. Connecting it to watsonx.governance brings that estate into the same enterprise AI risk and compliance environment as everything else you govern: the model inventory, the risk register, control mappings to recognised frameworks, and integration with OpenPages where AI risk needs to sit inside existing GRC rather than beside it.

The practical value of that connection is that your risk committee stops receiving two separate stories. Agents are not a special category reported separately from the rest of the AI estate. They are part of one picture, rated on one scale, evidenced in one place.

What no platform gives you

This applies to any vendor in the category rather than to one.

Control planes give you primitives. Policy enforcement, observability, identity, evaluation, audit trails, cost attribution. Those are necessary and genuinely hard to build yourself.

They do not decide who owns a given agent. They do not define your escalation path when one misbehaves. They do not tell you which actions require a human and which do not, which is a judgement about your customers and your risk appetite. They do not design the handoff between an agent and a person, or specify what gets recorded at that handoff. And they do not answer who is accountable when an agent acting on a customer's behalf gets it wrong.

That is organisational design, and no vendor ships it.

The CBUAE oversight models give you the vocabulary for the autonomy decision. Classifying each agent against it, and having someone sign that classification, is work only your organisation can do. Buy the platform before you have done it and you end up with excellent instrumentation pointed at an estate nobody has decided how to govern.

Where to start

Four steps, in order, and only the last is a purchasing decision.

Discover what you have, across frameworks, including vendor-embedded agents. Expect the number to be higher than the executive team assumes. That conversation is frequently the most valuable output of the exercise.

Run the permission audit. For each agent, what can it reach today against what it needed on the day it was provisioned. Uncomfortable, high value, and almost never done.

Set autonomy levels explicitly, using the CBUAE vocabulary. Classify by reversibility and customer impact. Irreversible and customer-affecting is where a human belongs, regardless of how well the agent performs in testing. Record the decision and who made it.

Then choose tooling, once you know the shape of the estate and what you have decided about it. The conversation is shorter and considerably cheaper in that order.

The shift underneath

Arvind Krishna made a point at IBM's Think conference this year that applies well beyond IBM's customers. The enterprises pulling ahead are not the ones deploying more AI. They are the ones redesigning how the business operates around it.

For governance, the redesign is this. With deterministic automation you were governing a process. You could read it, test it, and know that what it did yesterday is what it will do tomorrow.

With agents you are governing a decision-maker. One that pursues goals, chooses its own route, and improvises when the route is blocked. The DIFC Commissioner's framing, that such a system sits in a position substantially similar to an employee within your organisation, is useful precisely because it tells you what kind of control to build.

You supervise employees continuously. You scope their access to their job. You require a second signature on the consequential things. And you can remove their access on the day you need to.

The question is not whether your agents are performing well. It is whether you could say, tomorrow, what every one of them is permitted to do and who decided.

DISCLAIMER - This article describes the regulatory landscape as understood at the date of publication and is provided for general information. It does not constitute legal advice. Organisations should obtain advice from qualified UAE counsel on their specific circumstances.

Frequently Asked Questions

What is the difference between RPA and AI agents?

Robotic process automation follows a defined path and fails by stopping when it meets something unexpected. An AI agent is given a goal and determines its own route, so when it meets an obstacle it tries an alternative rather than halting. That flexibility handles messier work, and it means failures are quieter and harder to detect.

What is agent sprawl?

Agent sprawl is the accumulation of AI agents across an organisation faster than anyone can catalogue them. Agents arrive through vendor product features, business teams building on platforms they already have, and developer tooling. The result is an estate with no inventory, no risk ratings and no consistent controls.

What is an agentic control plane?

A centralised layer for observing, governing and optimising AI agents across an enterprise, regardless of which framework they were built on or where they run. Typical functions include discovery of agents across platforms, runtime policy enforcement, evaluation before publication, credential health monitoring and visibility into consumption and cost.

How do you control AI agent costs?

Through visibility into which agents are driving token usage and model calls, attributable to a business owner, combined with runtime policy limits. Runaway consumption from retry loops and unbounded workloads is a recognised failure mode, and the instrumentation that reveals what an agent is doing is the same instrumentation that reveals what it costs.

How much autonomy should an AI agent have?

Decide per agent, by reversibility and customer impact. CBUAE guidance offers a usable vocabulary of human-in-the-loop, human-on-the-loop and human-out-of-the-loop, with the last appropriate only for low-risk, non-material processes with appropriate controls. Irreversible actions affecting customers warrant human involvement regardless of testing performance.

Aligne AI works with banks and insurers across the UAE and GCC to build AI governance that satisfies multiple regional regimes from a single control set, aligned to CBUAE supervisory expectations, ISO/IEC 42001 and the NIST AI Risk Management Framework. Get in touch to discuss your regional estate.

Blog

Our latest news

Stay Informed: Engage with our Blog for Expert Analysis, Industry Updates, and Insider Perspectives

All Posts
Services Image
Govern, Orchestrate, Build: How the Enterprise AI Question Fits Together
Most organisations solve three AI problems separately. They are one problem seen from three positions, which is why programmes stall.
Read Details
Services Image
Who Governs the Code the AI Wrote? Security and Auditability in Agentic Development
Your institution owns the code in production regardless of who wrote it. What that means now that agents write a meaningful share of it.
Read Details
Services Image
The COBOL Cliff: Modernising Mission-Critical Systems When the Expertise Is Retiring
The hard part of legacy modernisation was never the rewriting. It was understanding what the code does. That constraint has changed.
Read Details

Ready to Take the First Step?

let’s design the governance framework your AI strategy deserves

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
bg elementbg elementLet's Talk