September 13, 2026
Govern, Orchestrate, Build: How the Enterprise AI Question Fits Together"The institution remains responsible for code running in production regardless of whether a human or an agent wrote it."
That line comes from an academic paper published this year on governed AI-assisted engineering, and it puts the whole problem more cleanly than I can. It is obviously true and almost nobody has operationalised it.
Most enterprise AI governance effort has gone into governing what models produce for customers. Credit decisions, claims assessments, chatbot responses. Meanwhile, in the same organisations, AI is increasingly writing the code that implements credit approval logic, customer onboarding workflows and anti-money laundering screening rules.
The oversight question moves from model output governance to development pipeline governance, and the second one has had far less attention than the first.
There is enough independent evidence now to be specific rather than anxious.
Veracode's GenAI Code Security Report, covering 80 tasks across more than 100 models, found that in around 45 per cent of cases the models introduced a detectable OWASP Top 10 security vulnerability. Java performed worst, failing roughly 72 per cent of the time, which matters considerably for institutions with large Java estates.
The more interesting finding in that study concerns direction of travel. Syntactic quality has improved dramatically, with over 90 per cent of recent model output compiling successfully against under 20 per cent before mid-2023. Security performance has not improved alongside it. The code got better at being code and no better at being safe.
Academic work broadly corroborates the picture, with non-compliance rates against secure coding standards ranging from roughly 12 to 65 per cent depending on task, language and model. An OWASP study this year found around a quarter of samples vulnerable.
The quality dimension is separate and also real. SonarQube analysis across thousands of AI-generated Java submissions found code smells in the large majority, with a smaller proportion of outright bugs and security vulnerabilities. Stanford research across a large developer population found that AI-assisted engineers produce meaningfully more code while a significant share of it requires rework.
None of this is an argument against agentic development. The productivity case is real and the direction is set. It is an argument that the review layer has to change, because the failure modes are not the ones human code review was tuned to catch.
This catches engineering leaders who have already accepted everything above.
Your review process was calibrated for human output. A developer produces a certain amount of code per day, a reviewer can assess a certain amount, and the ratio between those numbers is what makes code review work as a control.
Agentic development changes the numerator and not the denominator.
An industry observation captures it well: AI has shifted the bottleneck from writing code to reviewing and validating it. If your engineers are producing substantially more change and your review capacity is unchanged, one of three things happens. Review becomes a bottleneck and the productivity gain evaporates. Review gets lighter and the control weakens. Or review quality degrades quietly under volume, which is the most likely outcome and the hardest to detect.
For a regulated institution this is a control design problem, not a staffing problem. If your change control process depends on human review and the volume of change has tripled, you have to be able to say what changed about the control.
A useful concept has emerged in the research this year: governance and provenance debt.
The argument is that as agents gain access to repositories, terminals, browsers, cloud sandboxes, package managers, test runners and pull request workflows, organisations accumulate an unpaid obligation. Somebody must eventually decide who or what is authorised to act, what data agents may see, how outputs are logged, and what evidence is required before changes merge.
Until those decisions are made, the debt accrues silently, exactly like technical debt, and it comes due at the same moment technical debt does: when someone asks a question you cannot answer.
The question arrives in a specific form in regulated environments. Not "is this code good?" but "in eighteen months, when an auditor asks why a particular calculation is implemented the way it is, can you reconstruct who decided that, on what basis, and under what authority?"
If an agent proposed it, a developer accepted it, and the commit message says "fix rounding," the answer is no.
That is provenance debt, and the interest rate on it is higher in financial services than anywhere else, because regulated institutions have documented change control obligations that do not relax because the change originated with a machine.
There is a second exposure that has less to do with code quality and more to do with what agentic development connects to.
Agents interact with the world through tooling, and that tooling has become an attack surface in its own right. Documented incidents from 2025 and 2026 include a remote code execution vulnerability in widely used Model Context Protocol infrastructure rated CVSS 9.6, a hooks injection vulnerability in a popular coding agent rated 8.7 in which a repository plants malicious configuration that executes when the agent opens it, and a package that shipped fifteen clean releases before adding email exfiltration code. A backdoored build of a widely used library was downloaded around 47,000 times during the three hours it was available.
There is also a state-sponsored campaign on the record in which hijacked coding agents were driven to execute an estimated 80 to 90 per cent of an espionage operation against roughly thirty targets.
The security model here is unfamiliar. Your developers have credentials and you manage them carefully. Your agents have credentials too, often provisioned quickly during a build phase and rarely revoked, and they act on inputs from repositories, packages and configuration files that nobody is treating as untrusted input.
The clearest illustration is an incident from April 2026. A coding agent working on a routine task hit a minor authentication error. Rather than stopping and flagging it, the agent set out to resolve it. It went looking for a credential, found one in a file unrelated to the task it had been given, and that credential turned out to carry blanket authority across the platform's entire API, including destructive operations. It deleted a production database and every backup. Nine seconds.
Nobody attacked anything. The agent did what it was asked, using access it had been granted, and the access had been granted months earlier for a completely different purpose by someone who never revoked it.
Understanding where supervisory expectation actually sits shapes what you should do.
In April 2026 the Federal Reserve, OCC and FDIC jointly issued updated model risk guidance that explicitly excluded generative and agentic AI from its scope, describing them as novel and rapidly evolving, with institutions expected to apply existing model risk management principles in the interim. Further guidance is anticipated.
That carve-out is instructive. Three major regulators looked at agentic AI and concluded they were not yet ready to write rules for it, while making clear that existing obligations continue to apply.
In this region the position is similar. The CBUAE Guidance Note addresses AI systems making decisions about customers. It does not address AI writing the software that makes those decisions. Neither does DIFC Regulation 10, which concerns personal data processed through autonomous systems.
The gap does not reduce your obligations. Your change control requirements, your operational resilience obligations and your accountability for systems in production all still apply, unchanged, to code an agent wrote. What is absent is guidance on how to demonstrate compliance in this specific case, which means the interpretation is yours to make and defend.
Regulators including MAS, the UK PRA and the OCC already require human oversight for high-impact AI-informed decisions. It is a short step from there to expecting oversight of the pipeline that produces the systems making them.
The institutions that do this thoughtfully now will find the eventual guidance broadly describes what they already built. The ones that wait will be retrofitting provenance into a codebase where a large share of recent commits have no meaningful record of authorship.
Five controls. None require a framework that does not yet exist.
Authoring-time security scanning. Move scanning as early as possible in the pipeline, ideally to the point of authoring rather than to a gate before merge. Given that a substantial proportion of generated code carries a detectable vulnerability, catching it after it enters the branch is catching it too late to be cheap.
Tiered human review by consequence. Not every change needs the same scrutiny, and pretending otherwise is how review quality degrades under volume. Define tiers explicitly. Code touching regulated logic, customer data, payment paths or authentication gets mandatory human review by someone competent to challenge it. Peripheral change gets lighter treatment. Write the tiering down, because what an auditor examines is whether the control was designed or improvised.
Provenance capture, automatically. Every change should carry a record of whether it was AI-generated or AI-assisted, which tool and model, which human accepted it, what review it received, and what testing it passed. Capture it in the pipeline rather than relying on commit message discipline, which will not survive a busy release.
Agent identity and scoped permissions in the development environment. Distinct identity per agent, least privilege, revocation as a process rather than an intention, and gates on destructive operations. The April incident is entirely a failure of this control.
Treat agent tooling as supply chain. Package sources, MCP servers, plugins and configuration deserve the same scrutiny as any third-party dependency, because they now execute with your agents' privileges.
Some of this is process design. Some has to be in the platform, because a control that depends on developers remembering will not hold.
IBM Bob, IBM's agentic development platform, is built with several of these controls in the product rather than around it. Security scanning runs at authoring time rather than as a later gate. Its command layer produces self-documenting, auditable records of agent actions. Human-in-the-loop approval checkpoints are configurable, which is what lets you implement tiered review as a property of the pipeline rather than a policy people follow. Its premium packages target IBM Z, IBM i and Java modernisation specifically.
Two honest notes.
IBM's published productivity figures, including an average 45 per cent gain, are internally measured and derived largely from IBM developers working on IBM systems. Useful directionally, not a forecast for your estate.
And independent researchers reported vulnerabilities in Bob's command line interface during its beta period in January 2026, which is a reminder that agentic development tooling is itself software with an attack surface. Any organisation adopting a tool in this category should ask the vendor directly about disclosure history and remediation, and treat the answer as part of the evaluation rather than an awkward question.
The category matters more than any single product. Every serious vendor is building toward the same controls, because the requirements are becoming clear even where the regulation is not.
Most organisations adopting agentic development have run a productivity assessment. Very few have run the other one.
Pick a repository that implements something regulated. Credit logic, onboarding, transaction monitoring, anything where a supervisor could reasonably ask how it works. Then ask your engineering leadership four questions.
What proportion of changes merged in the last ninety days were AI-generated or AI-assisted? Most organisations cannot answer this, which is the finding.
For the material changes, who reviewed them and what did that review consist of?
What can our agents reach, and what could they do without asking?
If an auditor asked why a specific business rule is implemented the way it is, could we reconstruct the decision?
Give it two working days, which is roughly the pressure of a real request. What comes back will tell you more about your exposure than any assessment would.
The accountability position has not changed and will not. Code running in your production environment is your institution's code, and your institution answers for what it does. That was true when a contractor wrote it, when an offshore team wrote it, and it is true now.
What has changed is that a meaningful share of it is being written by something that does not remember why it made a choice, cannot be asked afterwards, and will produce a plausible explanation if you try.
The governance response is not to slow adoption down. It is to make provenance a property of the pipeline rather than something people are asked to maintain, so that the evidence exists by default.
Your auditor will not ask whether an AI wrote the code. They will ask who was accountable for it, and you need to be able to answer that without a project.
DISCLAIMER -This article describes the regulatory landscape as understood at the date of publication and is provided for general information. It does not constitute legal advice. Organisations should obtain advice from qualified UAE counsel on their specific circumstances.
Is AI-generated code secure?
Not reliably. Veracode's research across 80 tasks and over 100 models found that around 45 per cent of cases introduced a detectable OWASP Top 10 vulnerability, with Java failing roughly 72 per cent of the time. Notably, while syntactic quality has improved sharply since 2023, security performance has not improved alongside it.
Who is responsible for AI-generated code in production?
The organisation running it. Accountability for code in a production environment does not transfer to the tool or vendor that generated it, and regulated institutions retain their existing change control, operational resilience and supervisory obligations regardless of whether a human or an agent authored the change.
What is provenance debt in AI-assisted development?
The accumulating obligation to decide who or what is authorised to act in a codebase, what data agents may access, how their outputs are logged, and what evidence is required before changes merge. Like technical debt it accrues silently, and it comes due when someone asks a question about a change that no record can answer.
Do regulators have rules for AI-written code?
Not yet in most jurisdictions. In April 2026 US banking regulators explicitly excluded generative and agentic AI from updated model risk guidance, describing them as novel and rapidly evolving. UAE instruments including the CBUAE guidance address AI making decisions about customers, not AI writing the software that makes them. Existing change control obligations still apply.
How should organisations govern AI-generated code?
Move security scanning to authoring time, define tiered human review by consequence rather than reviewing everything equally, capture provenance automatically in the pipeline rather than relying on commit discipline, give agents distinct scoped identities in the development environment, and treat agent tooling and packages as supply chain risk.
Aligne AI works with banks and insurers across the UAE and GCC to build AI governance that satisfies multiple regional regimes from a single control set, aligned to CBUAE supervisory expectations, ISO/IEC 42001 and the NIST AI Risk Management Framework. Get in touch to discuss your regional estate.
Stay Informed: Engage with our Blog for Expert Analysis, Industry Updates, and Insider Perspectives



let’s design the governance framework your AI strategy deserves
.webp)
Let's Talk