Thought Leaders

AI Governance: A Problem for Most, an Advantage for the Few

mm
Add Unite.AI to your preferred sources on Google

Enterprise AI is accumulating risk faster than anyone is accounting for it. AI agents are beginning to touch customer data, production code, and consequential business decisions, and in most organizations, they are governed by controls designed for a very different era of technology. At the same time, the external supports enterprises might have leaned on are thinning out. Leading AI labs have scaled back voluntary safety commitments, and regulation remains fragmented across jurisdictions. The trust burden is landing on the companies deploying AI, because ultimately they are the ones who will suffer the negative outcomes when and if they surface.

I do strategic AI work with companies across nearly every industry, and I talk to CEOs and boards daily. The consistent pattern I see is as follows: deployment accelerates while the control infrastructure simply does not keep pace. The pattern is consistent enough that I consider it structural rather than any one firm’s failure, and I will make a prediction: we are not far off from a CEO having to explain that a big mistake cost the business millions, and they will blame AI. I would politely suggest that is not an excuse. 

The good news is that I am fortunate to see firsthand a handful of companies that treat governance differently, and I believe those firms are building a long-term advantage that compounds over time, while others waiting on the sidelines are at risk of becoming obsolete. Understanding what the few see starts with understanding where the current trajectory leads.

We Are Headed Toward a Problem

Today’s deployments are small enough that human oversight can absorb the risk. Deloitte’s latest State of AI in the Enterprise research found that only one quarter of surveyed organizations had moved at least 40 percent of their AI experiments into production. Most enterprises are still early on the curve, running a handful of agents in narrow use cases with people tracking every piece of output.

The trajectory is what matters. As deployments scale from two agents to 20, to 200, to 2,000 and beyond, complexity compounds. Each new agent brings its own behavior, plus its interactions with every system it touches and every other agent it works alongside. At two agents, trust is easy: you know the team, and you review the outputs yourself. At 20, trust has to become a process, with defined owners, testing, and escalation paths. It can still work, but it is becoming harder to sustain. At 200, informal review becomes theater, because no leadership team can eyeball that much activity. And at 2,000, the human in the loop is virtually useless, as the interactions among agents, data, and business processes form a web beyond what manual oversight can cover. At that scale, trust needs to also exist at scale and be embedded in the firm’s technical architecture.

The Cost of Waiting

The missing trust infrastructure is also one of the reasons for the disappointing returns. In PwC’s latest Global CEO Survey, 56% of chief executives said AI had produced neither increased revenue nor reduced costs in the past year. In my experience, even the returns being reported look less convincing once you examine the underlying economics. Most firms count hours saved and call it return on investment. Hours saved become financial return only when the organization converts that capacity into something more valuable: greater output, better quality, avoided cost, faster growth, etc.

That conversion requires trust. An organization can move human judgment to higher-value work only when leadership trusts the systems taking over the lower-value work. Absent that trust, companies accumulate technical debt that does not scale, and over time, it meaningfully impacts productivity. It is the worst of both worlds: increased spend with little to show for it. Trust is the foundation on which these returns accrue.

The Elevator Problem

The linkage between trust and adoption is not a new one, and we can go back to the mid-20th century for a reasonable reference point. Driverless elevators existed decades before anyone would ride them. The technology worked, but passengers would step into a car, see no operator, and step back out. What changed was not the elevator, but the safety controls around it. After a 1945 operators’ strike shut down office buildings across New York, the industry engineered trust into the machine itself: the emergency stop button, the telephone line, doors that reopened when they met resistance, and a recorded voice explaining what to do. Within a generation, riding an elevator without a human at the controls became so ordinary that walking into an elevator with a person manning the buttons felt antiquated. The barrier to removing the human from the loop was never capability. It was trust, and the industry solved it by building visible, verifiable controls into the system and putting them where every passenger could see them.

Enterprise AI is at the step-in, step-back-out stage. Leadership teams see agents working without a human at the controls and hesitate, and the hesitation is rational: most organizations have not yet built the equivalent of the stop button, the phone line, and the door that will not close on you. The companies that engineer those controls into their systems will be the ones whose people, customers, and regulators step in and stay in.

Vendors Can Only Carry Part of This

Many leaders hope their AI vendors will carry the trust burden. Vendors can provide model-level safeguards and security controls, and the good ones do. Deciding whether a particular deployment is appropriate for your data, your regulatory obligations, your customers, and your risk tolerance remains your job, because your vendor lacks the context to do it. They have never seen your obligations, and they have no way to price what a bad outcome costs you.

A vendor’s safety commitment covers their model, not your use of it. It also is not permanent: as the past year has shown, those commitments can be revised or withdrawn. Regulated industries settled this question long ago for third-party models, and the standard they landed on is the right one: the institution that deploys the model owns the risk.

What the Few Are Doing

The companies pulling ahead are making three moves:

  1. They are building control infrastructure now, sized for the scale they intend to reach rather than the scale they are at, because controls have to exist before the agents that need them. 
  2. They are embedding risk classification and mitigation at the start of the deployment process rather than as a review gate at the end, so every use case is classified and controls are proportionate to risk from day one. 
  3. And they are standing up independent governance with real authority to verify these systems behave as intended, kept flexible enough to evolve with the technology. Independence makes the verification credible. Flexibility keeps it current.

It is worth pointing out that few regulators have explicitly established rules around the use and oversight of generative AI, but that does not mean we have nothing to look to for guidance. Regulators have well-established rules for managing traditional models, and we can use those as a guide. In April, the Federal Reserve, OCC, and FDIC issued revised model risk guidance reflecting decades of practice, a principle-based approach that includes model inventories, risk-based controls proportionate to exposure, independent validation, ongoing monitoring, and enterprise accountability for vendor products, and they directed institutions to apply these same practices to AI on their own. The message is clear: the principles that work are known, and the formal AI rules are not coming soon. Enterprises that wait for regulatory clarity will be waiting while the few act. This is why the heavily regulated industries, so often described as slow, enter this era with a head start. They spent decades building exactly the trust machinery everyone else now needs, and they are not waiting for permission to use it.

The Question That Matters

The divide is already forming. Most companies treat governance as a cost: something absorbed, endured, and minimized. The few treat it as the infrastructure that lets them move people to higher-value work and scale with confidence. Every leadership team is on one side or the other, whether they have chosen deliberately or not, and the gap is compounding quietly ahead of the scale that will expose it.  Nobody thinks twice about stepping into an elevator today, because trust was built into it. The companies that do the same for AI will scale past everyone still deciding whether to step in.

Jeff McMillan, founder & CEO of McMillanAI and former Head of Firmwide AI at Morgan Stanley, is one of the few IT leaders who has scaled global AI systems at major Fortune 100s, and can speak to over 100 highly successful use cases. McMillanAI works with enterprise leaders across a variety of industries, including banking, fintech, healthcare, education, entertainment and more to scale AI ethically and responsibly. McMillanAI services help ensure organizations understand the current state of their AI usage and where it can be improved at every level, from leadership to frontline workers. He is also an adjunct professor at Columbia, where he teaches AI strategy to MBA candidates and executives to ensure the next generation entering the workforce has the AI skills and know-how to succeed in an ever-changing workplace.