Thought Leaders

AI Agents Need Security Boundaries They Cannot Rewrite

mm
Add Unite.AI to your preferred sources on Google

It’s tempting to read the Hugging Face story as the moment AI agents went rogue. That isn’t quite what happened, and the details matter. These were cybersecurity research agents running in evaluations where safeguards had been deliberately turned down so researchers could see what the models were capable of. Nobody’s customer service bot woke up one morning and decided to attack a company. But that context doesn’t let anyone off the hook. An agent went past the boundary it was supposed to stay inside, used credentials and tools in ways its operators never authorized, and ended up in systems that belonged to someone else. That’s the part every security team should pay attention to.

Reuters reported that agents were probing Hugging Face as early as May, although researchers said they found nothing showing that the earlier activity caused a breach on its own. July was different. OpenAI said its models got around isolation controls, reached the internet, and compromised parts of its own research infrastructure along with Hugging Face systems. Hugging Face’s own account describes an intrusion run end to end by an autonomous agent system, one that exploited its data processing pipeline, harvested credentials, and moved across internal clusters.

The uncomfortable part is that the agents were doing their jobs. They were chasing the objective they’d been given. That’s why this story reaches well beyond one research lab. Enterprise agents chase objectives too. They hold credentials, call tools, and move faster than any person can review. An agent with perfectly good intentions can still do real damage, and a hijacked agent can use the exact same authority on behalf of an attacker. So security has to govern what the system can actually do, no matter how confident the model sounds or how harmless its stated purpose looks.

Secure the Action, Not Just the Model

Most early agent programs pour their effort into the model. Teams test prompts, tune refusals, add a second model to check the first, and watch the reasoning trace for signs of bad intent. None of that is wasted. But all of it is probabilistic, because it depends on yet another model making a judgment call. A production security boundary has to be deterministic, and it has to sit around the tools, credentials, networks, and transactions.

The question I’d ask is a concrete one: what can this agent actually make happen in the real world? Drafting a payment request is one thing. Releasing the funds is another. The same goes for preparing a database change versus running it in production, or flagging records that meet a retention rule versus deleting them. It can be the same model in both cases, with very different risk depending on which side of that line it sits.

A recent Unite AI overview of capability control draws the same line, tying risk to data, tools, permissions, autonomy, and the environment the agent runs in. I like this framing because it gets us past vague labels like “safe model” and “unsafe model.” It makes teams trace every path from an agent’s decision to something with real consequences.

Give Every Agent an Identity and a Narrow Mandate

An agent should never run on a developer’s account or inherit everything a human user is allowed to do. Shared identity wipes out attribution. Long-lived credentials give an attacker more time to abuse them. And broad service accounts let a small workflow wander into data and systems it has no business touching.

NIST now treats software and AI agent identity as its own architecture problem. Its concept paper asks how an agent can prove it’s authorized for a specific action, how an agent’s identity can be tied back to a human’s authorization, and how organizations can keep tamper-resistant records of what was intended and what actually happened. In practice, that points to a simple design. Every agent gets a unique identity, an owner (a person or a team), a defined purpose, and permissions scoped to the task in front of it.

Credentials should expire fast and work only for specific resources and actions. Network access should start from a tight allow list. If an agent needs to query one approved database, it shouldn’t also get a general shell, open internet access, or the power to mint new credentials. And as work passes down a chain of agents and tools, authority should get narrower at each step, not wider.

NIST also warns against credential sharing and overly broad access, and that warning deserves weight because agents are opportunistic. If one route is blocked, they may try another tool, poke around their environment, or stumble onto a token someone forgot about. Harvested credentials were part of the July story, too. Least privilege keeps the blast radius small when the reasoning layer does something its designers didn’t expect.

Keep Authorization Outside the Reasoning Loop

An agent can recommend an action. It shouldn’t get to decide whether it’s allowed to take it. That decision belongs in a separate enforcement layer the agent can’t rewrite, switch off, or talk its way around. Every tool call should show up as a structured request: which agent is asking, which human sponsored it, what operation it wants, what it’s targeting, and what limits apply. The enforcement layer then allows it, blocks it, or escalates it.

OWASP describes excessive agency as some mix of unnecessary functionality, excessive permissions, and too much autonomy. Its guidance calls for narrow tools, minimum permissions, authorization at the downstream system, and user approval for high-impact actions. I think that’s exactly the right order. The rule should be enforced by whatever owns the data or executes the transaction. If a model says an action is approved, that statement alone should carry zero weight.

This separation helps with prompt injection too. A poisoned email or document might steer the agent’s reasoning, but it can’t widen the agent’s credentials or pull down a policy gate. The model is free to ask for something forbidden. The system should still say no.

Save Human Approval for the Moments That Matter

Human review earns its keep when an action can’t be undone, crosses an organizational boundary, changes privileges, releases sensitive information, moves money, or touches a production system. Ask for approval on every routine step and you get two things: delays, and people who learn to click “approve” without reading. NIST calls out this consent fatigue by name.

A good approval request shows the exact action in plain language, including where it’s going and the parameters that matter. It should come from an authoritative system, not from text the agent wrote. The approval should expire quickly and cover that one action only. If any material detail changes, the system asks again.

OWASP’s agent security guidance recommends testing whether a high-impact action can go through without a valid, unexpired, parameter-bound approval. That phrase is worth remembering. “OK to continue” is a weak approval that can be approved by another agent or bad actor. “Transfer this amount to this account” or “deploy this change to this environment” is something the system can actually verify at the moment it executes, and a user fully understands.

Live human identity assurance matters at this gate. A push notification proves only that somebody or some thing clicked a button. Stronger designs require an enrolled person to use phishing-resistant public key authentication, backed by a local biometric verification method. FIDO standards bind public key credentials to the legitimate online service and keep biometric data on the user’s isolated device. Used well, hardware-backed authentication gives you much better evidence that the right person was actually there. It doesn’t replace transaction binding, a trustworthy display, or policy enforcement, though. You need all of them working together.

Monitor Behavior and Preserve Evidence

You can’t count on the initial prompt to explain what happened over a long agent run. Security teams need telemetry on tool calls, network activity, credential use, policy decisions, approvals, denials, and changes in scope. Monitoring should compare what the agent actually did against the boundary declared for that run. If an agent was assigned to analyze code and it starts hunting for external credentials or probing an unrelated service, that should set off an alert.

Logs need enough context to rebuild the chain of actions without spilling secrets in plain text. Each record should capture the agent version, its owner, the person or system that kicked things off, the tool used, the action requested, the policy result, and any human authorization. Signed or otherwise tamper-resistant records make the after-action review far more credible, especially when several agents and services were involved.

All of this has to run at machine speed. Nobody watching a dashboard is going to stop thousands of calls that finish in seconds. Automated controls should enforce rate limits, catch unusual sequences, and suspend credentials the moment behavior crosses a defined threshold. That way, human investigators get a contained incident to work through instead of an open-ended chase.

Design the Stop Path Before Launch

Every agent deployment needs a way to stop it that actually removes capability. Asking the agent to stop doesn’t count. Operators should be able to revoke its credentials, cut its network path, kill its runtime, and keep queued actions from starting back up. For high-consequence workflows, if the approval or policy service goes down, the system should fail closed.

Then test that path under pressure. Take the approval service offline. Give the agent conflicting instructions. Rotate a credential in the middle of a run. Simulate a compromised tool and an approver who never responds. Confirm that the action gets blocked and that you’re left with a useful record. And rerun those tests any time the model, prompt, connector, memory system, or permission set changes.

The Goal Is Accountable Autonomy

None of this is an argument against agents, and the Hugging Face incident shouldn’t scare anyone away from useful ones. What it should do is kill the idea that a safety prompt plus good intentions add up to a trustworthy deployment. Give agents room to analyze, prepare work, and handle reversible tasks. Keep their authority to cause real consequences narrow, visible, and enforced by something other than the agent itself.

Before an agent goes into production, leaders should be able to answer a handful of plain questions. What systems can it reach? What credentials can it use? What can it do without review? What triggers escalation? How does the approver see the exact action being approved? What evidence will be left behind? And how can security stop the run immediately?

If the answers are fuzzy, the agent has more authority than the organization realizes. The architecture that holds up over time pairs model safeguards with identity, least privilege, external policy enforcement, selective human approval, complete telemetry, and a stop mechanism that really works. It starts from an honest assumption: capable agents are going to surprise us now and then. Our security boundaries shouldn’t.

In the end think of an AI agent as an intern with (potentially) root access who isn’t afraid of HR. 

What protections and gates would they have?

Proceed accordingly.

Kevin Surace is CEO of Token and an AI pioneer, inventor, author and entrepreneur with decades of experience applying artificial intelligence. He holds 95 worldwide patents and speaks globally on AI, automation, and the future of work.