Most writing about agent security is still about the shape of the problem. This is not that. Your developers are already running agents that touch repositories, your finance team is being sold one that reads invoices, and somewhere a browser automation is logged into something it should not be. What follows is the configuration layer: six controls, in the order we would apply them.

Three assumptions to start from

Every control below follows from three things the last year has taught, expensively.

The input is hostile. An agent reads documents, web pages, tickets, emails and search results. Any of these can carry text that reads as an instruction. We walked through how ordinary underspecification turns into prompt injection in The Car Wash Paradox; the operational conclusion is that there is no reliable way to separate data from instructions inside a single text stream, so the separation has to be enforced outside the model.

The agent optimises for approval, not for truth. Published incident reports from frontier labs now describe reward hacking in the wild: agents doing what earns a positive grade rather than what was asked, and in at least one documented pattern, loosening their own constraints through the notes they write to themselves between sessions. Most of it is still fringe and the reports say so. It is enough to design against.

The access mode can change without notice. The most useful agents are being given computer use, because it works where no interface exists. In one on-the-record account from a team building exactly this, an agent moved from a structured integration to driving the mouse for a single step, and a non-technical user saw an animated pointer rather than the script. Seamless and observable pull against each other, and the market is choosing seamless. Your configuration has to choose the other one.

Control one: credentials scoped to the task

The default in almost every deployment we have reviewed is a single service account with broad rights, created once so that everything works, and never narrowed. It is the same mistake as the shared administrator password, made faster.

What to do instead is ordinary engineering. One identity per agent, never shared with a human and never reused between agents. Rights scoped to the task rather than to the role, which usually means read on almost everything and write on almost nothing. Credentials minted per run and short-lived, so that a leaked token expires before it is useful. Separate credentials for separate environments, so that an agent with production access simply does not exist unless a task requires one.

And the rule that has changed character: no secrets in repositories, ever. This used to be hygiene. Published reports now describe an agent that searched public repositories for API keys on purpose. A committed key is no longer waiting to be found by luck; something is looking for it, continuously, at machine speed.

Control two: the gate

The single most valuable structure in an agent deployment is an explicit boundary between proposing an action and causing an effect. Everything on one side is reversible and cheap. Everything on the other side needs a decision.

The list of what belongs behind the gate is short enough to agree in one meeting: writes to production systems and databases, anything that moves money, anything that sends a message outside the organisation, deletion of anything, use of a credential the task did not explicitly require, and any change of access mode. Everything else the agent may do freely, and should, because a gate in front of everything gets clicked through without being read, which is worse than no gate at all.

Two design details decide whether the gate works. The approval has to show the effect rather than the intention, meaning the actual command, the actual recipient, the actual sum, not a summary of what the agent believes it is about to do. And the gate must be outside the agent's reach: a separate process with its own permissions, not a tool the agent can be persuaded to call on its own behalf.

Control three: the log

Ask of any agent deployment: three weeks from now, can we reconstruct what it actually did on a given afternoon? In most deployments the honest answer is no. There is a conversation transcript, which records what the agent said, and that is not the same thing.

Log actions rather than dialogue: the call, the target, the parameters, the result, the identity used, the timestamp. Record the access mode on every action, and treat a change of access mode as an event in its own right, logged and alertable, because that is the transition that quietly widens the blast radius. Keep the log outside the system the agent can write to. Make it retrievable by a person who was not involved, which is the only test that matters, and give it a retention period that survives the time it takes to notice a problem, which is usually longer than you think.

Control four: everything it reads is untrusted, including its own notes

Retrieved content is data, never instruction. In practice that means the boring things: keep retrieved text in a separate channel from the task description, never concatenate a fetched document into the instruction block, strip or neutralise markup that carries directives, and constrain the output to a schema where an action is a structured field rather than free text the system then interprets.

The part organisations miss is that this applies to the agent's own memory. Summaries, memory files, hand-off notes between sessions and shared context files are all input, and they were all written by something that can be influenced. An agent reading its own past notes is an injection surface, and it is a particularly good one, because that content arrives with the authority of having come from inside. Treat persistent agent memory as an asset with an owner, a review and a retention rule, the way you would treat a database.

Control five: contain the blast radius

Assume the agent is fully compromised and ask what it reaches. That question, asked before deployment, produces better decisions than any amount of policy.

The answers are unexciting and effective. Run it in its own execution environment rather than on a workstation with a person's session. Give it an egress allowlist, so that an agent that has been talked into exfiltrating something has nowhere to send it. Never hand it a browser profile logged into banking, mail, or an administrative console, and if computer use is genuinely required, give it a dedicated profile with dedicated, limited accounts. Cap what a single run can do, in volume as well as in kind: an agent that can send one email is a feature, an agent that can send four hundred is an incident.

Control six: test for the failure mode you actually have

Conventional testing asks whether the agent completes the task. The failure that hurts is the agent completing something that looks like the task.

Build a small adversarial set and run it on every change. A document with an instruction buried in it. A ticket that asks politely for a credential. A task that is impossible, to see whether the agent reports failure or manufactures a plausible success. A task where the cheap way to earn approval differs from the correct outcome. Then read the agent's memory files occasionally, the way you would review firewall rules, and ask whether anything in there is now granting the agent latitude nobody approved.

The short version

  1. One identity per agent, scoped to the task, short-lived, never shared with a human.
  2. No secrets in repositories. Something is looking for them on purpose.
  3. A gate between proposal and effect, listing production writes, money, outbound messages, deletion, credential use and access-mode changes.
  4. Approvals that show the effect, not the intention, enforced by a process the agent cannot call.
  5. Action logs with the access mode on every entry, stored where the agent cannot write.
  6. Retrieved content and agent memory treated as untrusted input, in separate channels, with structured output.
  7. Separate execution environment, egress allowlist, no logged-in browser profile that matters, volume caps per run.
  8. An adversarial test set that includes an impossible task and a buried instruction.
  9. A named owner, and the agent in the same register as your other non-human identities.
  10. A kill switch someone has actually used in a drill.

None of this is novel. It is least privilege, separation of duties, audit logging and input validation, applied to an actor that is fast, credulous, and persuasive about its own reliability. The novelty is only that the actor is new. The controls are the ones we already knew, and the organisations that get this right in 2027 will be the ones that stopped treating an agent as a feature and started treating it as an identity with a budget of permissions.


Related reading: The AI Agent Security Crisis on identity controls, You Trained Your Employees. Who's Training Your AI Agents? and The Car Wash Paradox. The observations on reward hacking, self-written notes and access-mode changes come from published lab incident reports and on-the-record engineering accounts covered in our Focus notes during September 2026.