Here's a pattern I keep seeing.
A team wires up an AI agent that can do real things — send emails, run commands, query and modify the database, call external APIs. The demo is magical. It reads a request, figures out the steps, takes them, reports success. Everyone's impressed, and it ships.
Then comes the first incident. It emails the wrong list. It runs a destructive command against the wrong environment. It reads a web page that quietly tells it to do something nobody asked for, and it obliges. Suddenly the magical demo is a very un-magical cleanup, and someone's asking how this was allowed to happen.
Here's the thing: the demo was never the hard part. Getting an agent to do something is genuinely easy now. The hard part — the part that almost always gets skipped in the rush — is everything that keeps the agent from hurting you when it inevitably does the wrong thing. And it will do the wrong thing, because it's a probabilistic system acting in an unpredictable world.
So this is the checklist I think belongs before you let an agent touch anything that matters. Not the fun part. The part that separates an agent you can actually deploy from a liability with a good demo.
1. Least privilege — the agent can only reach what it strictly needs
This is the highest-leverage guardrail, and it's the one most worth getting right first, because it makes entire categories of disaster simply impossible.
An agent that cannot reach production cannot wipe production. An agent that cannot send money cannot be talked into sending money. An agent with no access to a system can't be the cause of an incident in that system, no matter how confused or compromised it gets. So the question to ask before anything else isn't "what should this agent be able to do?" — it's "what is the least it needs to do its job?" Then give it exactly that and nothing more.
Almost every catastrophic agent story, when you trace it back, turns out to be a permissions decision someone made without quite noticing — handing over credentials that could reach production at all, or a tool scope broader than the task required. The destructive action got the headline, but the over-broad grant was the actual mistake, made quietly, long before. Least privilege is the guardrail that catches the error at the point where it's cheapest to prevent.
2. Human approval for consequential actions — gated by blast radius
Some actions you can let an agent take freely. Some you absolutely should not let it take unattended. The skill is in drawing that line well.
The irreversible or high-impact ones — send, spend, delete, export, deploy — should pause for a human. But two things make or break this guardrail. First, don't gate everything. An agent that asks permission forty times a session trains the human to click "approve" without reading, and now your approval step is theater — the click records attendance, not consent. Reserve the gate for actions that actually warrant it.
Second, the approval has to be meaningful. "Approve this action? y/n" on something the human can't evaluate is a rubber stamp. Show the blast radius: what this affects, what it will cost, what it's about to change, what the relevant history is. Give the approver something concrete enough to actually reject. An approval made blind isn't a control; it's a liability with a signature on it.
3. Treat everything the agent reads as untrusted
Anything your agent ingests — web pages, emails, documents, a teammate's file, the output of a tool, a code comment — can carry instructions. This is prompt injection, and it's not a solved problem: you cannot reliably detect malicious instructions hidden in natural language, because the space of ways to phrase them is endless and attackers adapt to whatever filter you deploy.
So don't put your faith in detecting the bad input. Put the control on the action instead. The set of consequential things an agent can do (send, spend, delete, export) is small and you can enumerate it — so gate those few calls with hard checks that don't depend on the model's judgment about whether the content it just read was trustworthy. The practical rule of thumb: an agent becomes dangerous when it has private data access, exposure to untrusted content, and the ability to act externally, all at once. Remove any one leg of that triad and a successful injection has far less it can reach.
4. A reviewer that can actually say "no" — and has been tested saying it
For consequential actions, a second check — a validation step, or a separate "judge" agent whose job is to review the proposed action before it executes — adds real protection. A system that has to convince an independent reviewer is much harder to push into a bad outcome than one that just acts on its own first impulse.
But here's the part people skip: a reviewer that always approves is not a control. It feels like safety, it generates a reassuring log, and it stops exactly nothing. Before you trust a reviewer, hand it a deliberately bad action — one you know should be refused — and confirm that it actually blocks it. A guardrail you've never watched engage is not a guardrail; it's a hope with good UI. Test the "no," or you don't have one.
Better still, make that test permanent: have a deliberately-failing probe run on a schedule and emit a receipt like any other check, so "the reviewer has been tested saying no" becomes a standing, provable property of the system instead of a memory of one afternoon. (Credit: @slabb.)
5. An independent audit trail — because the witness can't be the suspect
When something goes wrong, you'll go to the logs to reconstruct what happened. And here's the trap: if the logs were written by the agent, they were written by the very thing that misbehaved. A confused or compromised agent doesn't produce a broken log that tips you off — it produces a clean one, a tidy record of a bad decision, which is worse, because a clean log makes you stop looking.
So the record has to be produced somewhere the agent doesn't control — the infrastructure or supervisor layer around it, not the agent's own self-report. Seal it so that tampering leaves a visible gap rather than a silent edit. And capture not just what the agent did but what it believed at the time — which environment it thought it was in, which target it thought it was acting on — because the action is usually defensible given a wrong belief, and the belief is the part that actually explains the incident.
A quick, disclosed note on what this looks like in practice: I work on xenition.com, an AI workspace, and this exact cluster — approval gates, an independent audit log, a second reviewer — is something we had to build in deliberately rather than bolt on afterward. In our setup an agent's consequential actions pass through an approval gate, every step is written to an audit log the agent itself doesn't author, and a separate judge-agent reviews an action before it runs. I'm not holding it up as the answer — plenty of stacks assemble these pieces differently, and the right shape depends on your system. I mention it only as a concrete illustration that none of this is theoretical: the approval gate, the independent record, and the second reviewer are things you can actually ship today, and increasingly things you should.
6. Blast-radius limits — assume it will go wrong, and bound how bad
The guardrails above reduce how often things go wrong. This one accepts that something eventually will, and makes sure no single mistake is catastrophic.
Put hard limits around the agent: spending caps, rate limits, quotas on how many actions it can take before it has to check in. Prefer reversible actions by default — soft deletes over hard ones, staged rollouts over all-at-once, drafts over sends. The mindset shift is the important part: stop trying to guarantee the agent never errs (you can't), and start guaranteeing that when it does, the damage is small, contained, and recoverable. A mistake that costs you a reversible change and a shrug is a completely different thing from one that costs you a database.
7. Observability that measures the effect, not the report
Finally, watch the right thing. A dashboard built to answer "did the agent report success?" is, by construction, a dashboard built to trust the witness — and we just covered why the witness can't be trusted. The agent will happily report a valid-looking success while pointing at the wrong target, and your green dashboard will tell you everything is fine.
So measure the world, not the agent's account of it. Did the invariant hold? Does the total still balance? Is the referenced record actually there? Is the system still in a consistent state? These checks are more annoying to write than "did it return 200," and they are the only ones that catch the failure where the report looks perfect and the reality is wrong. Measure the effect, not the self-report.
The takeaway
Capability is the easy part. It's also the exciting part, which is exactly why it gets all the attention and all the demo time — and why the guardrails, which are none of those things, get left for "later," which often means "after the incident."
But an agent's real value in production was never what it can do. It's what it can do safely, repeatably, and recoverably — what it can do without becoming the thing you spend next quarter cleaning up after. The seven items above aren't the glamorous part of building with AI. They're the part that decides whether the glamorous part survives contact with the real world.
Build the brakes before you build the engine. Or at the very least, before you take it out on the highway.
If you've shipped an agent to production, I'm genuinely curious: which of these did you have in place before your first incident — and which one did you add right after it taught you the hard way? Most of us learned at least one of these the expensive way. Which was yours?