AI Agent Security Best Practices: Eight Checks
Eight checks a security team can run on its own AI agents this quarter, and which of the eight Intercis answers today.
- Measured: 40 of 47 payloads, our own corpus
- Not measured: latency
- Not seen: hosted MCP tool calls
An agent rollout arrives at a security team as a pile of questions nobody has time to research. These eight checks are the short version, and none of them requires buying a product. Where Intercis answers one of them, a sentence says so and carries the limit. Where it does not, a sentence of the same weight says that.
Count the agents, and the tools each one holds
The inventory is the check the other seven depend on. One row per agent: who owns it, which provider API it calls, which model, and every tool it can reach. The tool column is the interesting one, because it is the blast radius written down. An agent holding a shell tool and a cloud credential is a different risk from one that can only read a wiki, and that does not show up in a count of agents.
Two sources usually get you most of the way: the provider billing, which knows which API keys are in use, and your own deployment configuration. Put an owner's name beside each row; an agent with no owner is the first finding of the exercise.
Intercis builds part of this list for you once traffic runs through it: one row per governed tool call, carrying the agent, the tool and a short excerpt of what the call held. The limit is that the list only covers what crosses the Anthropic and OpenAI API wire. An agent pointed somewhere else never appears in it, and a tool a hosted MCP server runs for the model never crosses that wire at all.
Narrow the credential, not the prompt
The principal your database sees is not the agent. It is the credential the agent was handed. A system prompt telling an agent to stay inside one output directory is a string. A role that can only read, a key scoped to one bucket, a write path limited to one directory: those are boundaries the agent cannot reason its way around.
Walk the inventory from check one. For each tool, ask what the widest thing that credential can do is and whether the agent's job needs it. The common finding is a credential scoped to a team rather than to a task, because it was issued before anyone imagined an agent using it.
Intercis does not issue, hold or narrow your credentials. It can refuse a call at the wire, and a refused call leaves the credential exactly as wide as you left it.
Close the network path to the provider
An agent reaches a control point because its base URL says so, and a process that can change its own base URL can call the provider directly. So the control is worth what the network rule beside it is worth. Allow outbound traffic to api.anthropic.com and api.openai.com from the control point and from nowhere else on the network your agents run on.
This is the check no vendor can do for you, and it is the one that decides whether the others survive contact with a compromised process.
For us it is the first of the five costs we publish under what Intercis does not cover: being in-path is a setting, not a law. Without that rule we are a control a compromised process can walk around.
Put the decision outside the agent process
A library the agent imports is code that process controls. It can be skipped, reconfigured or never called, and none of that is visible from outside. A control in a separate process cannot be switched off by the code it is judging. That is the architectural argument, and check three is its condition.
Ask where the decision is computed and what the agent runtime actually receives. A verdict returned beside the action leaves the action in the response, and the runtime is then free to execute it.
Intercis is that separate process. The proxy reads the model's response, policy runs on the tool calls in it, and once enforcement is on a denied call is taken out of the response before the agent runtime can execute it. The limit is check three.
Keep a record the agent cannot edit
One row per decision, written when the decision is made, in a store the agent's credentials cannot write to. That is the difference between a log and an audit trail. Ask your own build, or anything you are evaluating, three questions: who holds UPDATE and DELETE on that table, what a row contains, and how you would notice a missing one.
A log the agent process can reach is not evidence about the agent. A log nobody has tried to verify is not evidence either, so write the verification down and run it before the week you need it.
Ours is one row per governed event, append-only for every client role, chained with SHA-256 per tenant, and there is a script that verifies the chain independently of the proxy. The limit is that it is tamper-evident and not tamper-proof: anyone holding database-owner rights can rewrite rows and recompute the hashes, or drop the newest rows, and nothing outside the database anchors the chain to catch either one. What it does catch is a change that left the hashes after it alone. It does not stop that either.
Decide now how you stop one agent
A kill switch is a question asked at three in the morning, so the answer has to exist beforehand. Three different controls go by that name and they are not interchangeable: stop the process, revoke the credential, or refuse the agent's next call. Write down which one you have, who may operate it, and how long it takes. Then try it once, on a working day, with someone watching.
Ours is manual and per-agent. An operator sets the agent inactive, and from the next validation each tool call that agent makes is rejected with a 403 before it reaches a provider. Revoking the tenant's API key is the tenant-wide version. The limits: it stops the next call rather than the one already in flight, nothing scores an agent against a baseline, and nothing here terminates a running session by itself.
Test the control with your own payloads, then publish the miss rate
A control nobody has tried to get past is a claim. Write the payloads that matter in your environment, run them through the control, count what it missed, and put the number where your own team can read it. Then do the harder half: run ordinary safe work through it and count how often it flags something that was fine. A control that interrupts good work gets switched off, and then it decides nothing at all.
Ours are on the validation report. The deterministic layer catches 40 of 47 payloads in the in-repo agentic-threat corpus, 85.1% pattern-layer recall, with seven documented misses, five of them obfuscations of one destructive command. On 4,764 governed tool calls from our own Dogfood tenant, all of them labelled benign, today's rules flag 737: 15.47%, Wilson 95% 14.47% to 16.52%. The limit is that the corpus is written by the defender, so it measures regression stability rather than an attacker who has read the rules.
Write down what happens to the calls nobody has decided about
Some calls are neither clearly fine nor clearly forbidden, and a control that only allows and denies will guess at those. The practice is to name the path a person is on: who is asked, on which channel, what the timeout is, and what happens when the timeout expires. If no such path exists, write that down rather than assume one, because the assumption is what turns an undecided call into an allowed one.
We do not run one. The verdict state pending exists for actions that would need human review, no reviewer approval loop is wired to it, so there is no escalation step to record. A call is allowed, denied or observed in line.
Where that leaves the eight
Five of the eight Intercis answers with a limit attached: the inventory, the decision outside the process, the record, the kill switch and the measured miss rate. Three it does not answer at all: your credentials, your egress rule, and what happens to a call nobody has decided about. Two of those three are yours whichever way you decide about us.
What Intercis does not see
Intercis reads the tool calls that cross the Anthropic and OpenAI API wire. When a hosted MCP server runs a tool for the model, that call never crosses the wire, so the proxy does not see it and no event row is written for it. Being in path is also a configuration setting: a process that can change its own base URL can route around the proxy unless network egress control forces the traffic through it.