How Intercis works, and what it costs you

The whole case, for the engineer who has to sign off on it: the request path, what a deny does to the response, what we counted and where, the price, the costs, and a check you can run on our published files in a minute.

What it does

Your agent's base URL points at the proxy. The proxy forwards the request to Anthropic or OpenAI on your key, judges each tool call in the answer, and takes a denied one out of the response before your runtime reads it. Every call it judges writes a hash-chained audit row.

What it costs

It is in the path, so if we are unreachable your agents stop. It sees only the LLM API wire, not tools the provider runs on its own servers. It holds a stream until it has judged it. When the classifier cannot answer, the call goes through and the deny rules keep enforcing. It is hosted only.

Check the chain yourself, one minute The five costs and three gaps

The request path, in words: your agent runtime sends its request to the Intercis proxy, which forwards it to Anthropic or OpenAI under your own provider key. The answer comes back carrying a tool_use block. The proxy judges that block, deletes it, and hands your agent runtime a response that carries text only. Two branches leave the proxy, and neither is a step on the way. One writes the audit row, hash-chained, one chain per tenant, for each judged call the database accepts. The other is drawn as a dashed line: when no deny rule matches, the proxy asks a classifier, and that call goes to the same provider on an Intercis account rather than yours, carrying the tool name and the full tool input.

The dashed line is the one a security reviewer will ask about. When no rule matches, the classifier call leaves our boundary on our provider account, carrying the tool name and the whole tool input.

Other tools return a verdict. This one takes the action off the wire.

A verdict is advice, and something downstream still has to act on it. Intercis edits the response: the content array becomes one text block, so the tool call is gone before your runtime reads the message. There is nothing left for it to ignore.

What the model returned

{ "role": "assistant",
  "content": [
    { "type": "text",
      "text": "Clearing old releases." },
    { "type": "tool_use",
      "name": "bash",
      "input": {
        "command": "rm -rf /var/app/releases"
      } }
  ] }

The agent runtime would have run this block. Abridged: model and usage are copied through from the upstream response and are left out of both panes.

What your agent runtime received

{ "role": "assistant",
  "content": [
    { "type": "text",
      "text": "⛔ Intercis blocked this
  action.\nPolicy: shell-rm\nThis tool
  call was denied by your organisation's
  AI governance policy and has been
  logged for audit review." }
  ],
  "stop_reason": "end_turn",
  "stop_sequence": null }

That text is the string apps/proxy/main.py writes, with your policy name in place of shell-rm. The two \n are its only line breaks; the rest is wrapped to fit this column. No tool call is left to run.

It did not run.

40 of the 47 payloads in our own corpus end that way on the deny rules alone. Seven get past the rules, and five of those are the same rm written five ways.

The seven the deny rules miss

A deny rule is a regular expression over the flattened tool input. A command assembled while it runs, and a command sitting inside a file the agent was told to read, are both outside what that can see. They carry names in our test suite; we do not print them here.

Seven documented bypasses. Five are obfuscations of one command, each rewriting it so the rule's word never appears as typed. Two are tool poisoning, where the destructive instruction sits inside a file or document the tool call reads rather than in the call itself. We do not publish the techniques.

85.1% is the rules on their own. The classifier is the layer that reads a call no rule matched, and we have no measured number for it against these seven. The last two are a gap in both layers: we judge what the agent asked to do, not what comes back to it.

A score, an alert or a log line would all have left that block in the message, for something downstream to catch in time.

Watch a denial happen in the demo lab, on a real agent

What you change, and what the proxy does with the call

Two changes: the base URL and 2 headers, of which 1 is required: the key. The other names the agent, and a key bound to one agent does not need it. No code inside the agent changes, and we install nothing on the machine it runs on.

  1. Point the agent at the proxy

    In Claude Code they are environment variables:

    ANTHROPIC_BASE_URL=https://api.intercis.io
    ANTHROPIC_CUSTOM_HEADERS="x-intercis-key: ik_live_...
    x-agent-id: your-agent-id"
    ANTHROPIC_API_KEY=your own Anthropic key, unchanged

    One Name: Value per line in the headers variable. Your provider key stays yours: we forward it and we do not keep it. We store an HMAC-SHA256 of the Intercis key under a 256-bit secret, and you see the plaintext once. An unbound key is the identity: whoever holds it can claim any agent you registered. Bind a key to one agent and the proxy refuses any request naming another.

  2. Rules decide first, before any model does

    The deny rules run first. They are regular expressions, so the same call gets the same verdict every time, and you can read the one that stopped you. Every rule name starts with its category: shell, cred, exfil, privesc, logtamp and thirteen more.

    When no rule matches, an intent classifier reads the call and answers in one word. It is asked only about tools that can act: execution tools, file-writing tools, and any mcp__ tool your runtime executes. A prompt-injection scanner runs alongside both, observe only. A tenant runs in observe, where the verdict is logged and nothing is blocked, or in enforce, where a deny deletes the call. Pilots start in observe.

  3. Every call it judges writes a row, unless our database is down

    A governed call crosses the Anthropic or OpenAI API wire and is one your runtime would run. A hosted MCP call crosses no wire of ours and writes no row. These are the columns this page argues about, on the row the denial above writes:

    created_at
    2026-09-12T14:07:52Z
    agent_id
    your-agent-id
    action
    bash
    target
    rm -rf /var/app/releases
    verdict
    deny
    policy
    shell-rm
    injected
    false
    source
    anthropic-messages
    classifier_status
    not_run
    raw_request
    the full tool input, credential-shaped strings masked
    event_hash
    sha256 over ten of this row's fields and the hash of the row before it

    A worked example, not a capture: the column names are our own, the values are the ones the example above would write. not_run is what classifier_status says when a deny rule matched, because the classifier is asked only when none did. The row carries more columns than these eleven, and the demo page defines the rest, including command_fingerprint.

    The audit write never blocks the call. If our database is unreachable we still block, in line, and the deny row waits on local disk until the database comes back.

What we counted, and where

Each figure was counted at the commit under the table, by the command beside it. The repository is private, so for these six you get the command and the count. The chain check below is the one you run yourself, on files we publish.

Repo counts on 28 September 2026, at commit 78e1765, which landed the same day; the corpus run at commit 87ccd47 on 12 September 2026.
What was counted Figure
Deny rules in the deterministic layer apps/proxy/deny_list.py 111
Prompt-injection patterns, observe only apps/proxy/injection.py, in 4 families 13
Detection on our own agentic-threat corpus scripts/detection_bench.py: 40 of 47 payloads caught, 7 documented bypasses 85.1%
Database migrations infra/supabase/migrations/ 63
Tables with Row Level Security on the same migrations, counted distinct 18
Test cases collected in the proxy suite pytest --collect-only -q apps/proxy/tests 3,802

We wrote that corpus, so 85.1% measures whether we have regressed, not whether an attacker can get past us. It holds no safe payloads. That number is below, on our own traffic.

How often we would have stopped work that was fine

15.47% 737 of 4,764 governed tool calls

Our own coding-agent traffic: tenant Dogfood, agent claude-code-mac, every governed call from 25 August to 4 September 2026 except 31 August, the day a four-hour classifier outage left 25.6% of that day's calls unanswered and voided it under our own runbook. The tenant was in observe, so nothing was removed; the 737 are what today's rules flag, and 516 of them are what enforce mode would have deleted. The other 221 are injection-scan rows, which are recorded next to the request and never block. We labelled every bucket benign, by bucket and not call by call, and the page that works it out says which is which.

When a rule is wrong on your traffic, an owner or an admin writes an exclusion naming one policy, and optionally one agent and one literal string, so the calls you named go through and the rule keeps enforcing on everything else. Nothing saves without an end date, or with a reason under twenty characters. Both of those are constraints in the database, not manners in the form.

How the 15.47% was worked out

What the chain proves, and what it does not

Each row carries a SHA-256 hash over its own fields plus the previous row's hash, one chain per tenant, written by a trigger before the insert. That catches an edit, or a row deleted from the middle, that did not also recompute every hash after it: the next row's link stops matching. It catches nothing else, and it cannot prove a row was never written. No client role holds UPDATE or DELETE on that table; we hold the service role and it bypasses that, so the chain is tamper-evident rather than tamper-proof: somebody holding our database credentials could recompute the chain, drop the newest rows or append rows of their own, and only a hash recorded outside our database would catch it. Nothing we ship puts one there.

Each event recorded from 2026-09-20 is hashed over every one of its columns except raw_request, the tool input a nightly job empties at 90 days. That is the ten fields below plus classifier_status and classifier_reason, command_fingerprint, the cloud and Claude Code attribution columns, the token counts and the model name. classifier_reason is where a fail-open is recorded, and on those rows it is inside the chain.

Events recorded before 2026-09-20 are hashed over ten named fields: id, tenant_id, agent_id, action, target, verdict, policy, injected, source, created_at. Those rows keep that format for good and are never re-hashed, because recomputing the hashes over rows already written is the act the chain exists to detect, so the other columns of a row that old stay outside its hash. The trigger we publish carries both formulas, so this is a sentence you can check rather than take from us.

The row is kept for the life of the tenant. A nightly job erases raw_request 90 days on, which is why that column is outside the chain: hashing it would break every chain the night the first row aged out. target is not erased. It holds the command, the path or the query, cut at 500 characters, and it stays as long as the row does.

What you pay

One price per registered agent, month to month. No quote and no call to find out what it is.

Team

$200 a month, first agent

$190 a month for each additional agent. Five agents is $960 a month.

  • An agent is one registered identity.
  • The 90-day pilot is this price, not free and not a discount, and it opens in observe mode.
  • You get the hosted proxy, on the Anthropic and OpenAI routes we govern today.
  • The dashboard: the event log with filters and CSV export, the agent register, the API keys, and the policy page where enforcement mode and exclusions live.
  • Design partners pay 40% less, which is $576 a month for the same five agents.
  • Month to month. Cancel any time, including during the 90 days.
Start a 90-day pilot

Enterprise

Custom Quoted per fleet

There is no self-hosted build today. The proxy, the dashboard and the audit log run on our machines, which is the question this card exists to answer. A self-hosted deployment of the proxy has been rehearsed end to end in our own lab; the licensing for one is undecided.

  • Everything in Team, quoted per fleet, and hosted by us the same way.
  • The card checkout stops at 20 agents, which is where the quote starts.
  • The routes are the ones Team gets: Anthropic's messages route, OpenAI's responses and chat completions.
  • Rate limits are per tenant, 300 requests a minute by default, with a per-agent override. They are counted in memory on each proxy instance.
  • A pilot agreement and a data processing addendum exist as drafts. No lawyer has read either, and both are written for a customer in the United States. If your data is European, British or Swiss, we are not ready for you yet.
Start a 90-day pilot

What it costs you

5 costs and 3 gaps. The costs are the ones we publish under what Intercis does not cover: being in the path, seeing only the LLM API wire, buffered streaming, a new trust boundary and the provider routes. The gaps follow from them: if we are unreachable your agents stop, the classifier fails open, and the audit log can have holes. The security page has the longer list and the detail under each one.

  • Being in the path is a setting, not a wall

    Your agent reaches us because its base URL says so, and a compromised process can change that line. Pair us with an egress rule that blocks direct calls to api.anthropic.com and api.openai.com. Without that rule we are a control it can walk around.

  • If we are unreachable, your agents stop

    We are in the path, so a call that cannot reach us does not reach the provider either. Nothing queues and nothing routes around us. The proxy is one hosted service behind a health check on /health, and it restarts on failure up to ten times. There is no second region, no replica and no failover in anything we ship.

    The egress rule above is what makes that total. With it in place you cannot point the base URL back at Anthropic and keep working, and that is the trade you are making when you close the bypass. We publish no uptime number and no SLA, because nothing in our code measures one.

  • When the classifier cannot answer, the call goes through

    Nothing on that path raises. The call is allowed, the row records the status and a typed reason, and an alert goes out carrying the provider and the upstream status, so an allow because the control was down does not read like an allow because the call was judged safe. The deny rules underneath keep enforcing. Above 5% unavailable across a rolling 200-call window holding at least 20 calls, the proxy pages us.

  • We only see the LLM API wire

    Tools the provider runs on its own servers never cross that wire. Hosted MCP calls, web search, file search, code interpreter and image generation are neither judged nor logged.

  • Streaming is buffered before it is judged

    To judge a tool call we have to see all of it, so the proxy holds the whole stream and replays it after judging. That adds the full stream duration before your agent sees any output. Nothing in our code measures that delay, so the shape is all we can give you: a deny rule is a compiled regular expression run over at most 65,536 bytes of the flattened tool input, and the one extra model call happens only when no rule matches, and only for execution, file-writing and mcp__ tools. Past that cap the scan stops, the call is denied, and no exclusion releases it.

  • We are a new trust boundary, and hosted only

    Your prompts and tool inputs pass through our proxy in plaintext. The classifier call runs on our own provider credential, so the tool name and input reach Anthropic or OpenAI under an Intercis account, not your agreement with them. There is no self-hosted build you can buy today. We have rehearsed a self-hosted deployment of the proxy end to end in our own lab, and what is undecided is the licensing, not the mechanism.

  • Keeping up with each provider's routes is our work, and it lags

    We govern Anthropic's messages route and OpenAI's responses and chat completions routes. A new provider, or a new route shape at an old one, is work we do first, and that traffic is not governed until it is done.

  • Our audit log can have holes, and the chain will not show them

    A database outage stops the log, not the blocking. Deny rows wait on disk; the allow and observe rows from that window are gone for good, and a row that was never written leaves nothing behind. A pilot runs in observe, so the rows a pilot most wants are the ones in that set.

The long version is on the security page

What Intercis does not cover

What a reviewer can check today

Before you pay us anything, and without an account or a form: verify our published sample chain on your own machine, recomputing every row's hash. Three files. Three fetches and one check.

curl -O https://intercis.io/evidence/verify_event_chain.py
curl -O https://intercis.io/evidence/sample-chain.json
curl -O https://intercis.io/evidence/event_hash_trigger.sql
python3 verify_event_chain.py sample-chain.json

sample-chain.json is five rows we built to be read, not five rows out of anybody's tenant, and we rewrote the third one on purpose: its verdict reads allow under a policy that denies. The verifier recomputes each row's hash from that row's own fields and the hash of the row before it, stops at the first one that does not match, and exits 1. It is standard library Python with no network calls, so you can read it before you run it.

Then put that row back. Open the file, change the third row's "verdict" from "allow" to "deny", and run the same command: the line ends ok (5 events, chain verified from genesis) and it exits 0. Breaking it and unbreaking it is the whole proof, and it takes a minute.

That is our script against our file, which on its own proves nothing, so the database code that writes the hash is the third file: event_hash_trigger.sql, the functions and the trigger copied out of the migration, with its own header naming which columns the formula covers and which it does not. Read one against the other and check the formula matches.

What you cannot do yet is run this on your own rows. There is no self-serve export of the hashes: the dashboard's CSV carries eight columns and neither hash column is one of them, and no endpoint hands out the head hash. Today the check runs on the file we publish.

The demo lab: a real agent, a denial on the wire, and the row it wrote

What we do not have

Intercis is not SOC 2 certified and no audit is underway. A SOC 2 and ISO 27001 evidence-package export is a roadmap item with no date; nothing on this page waits on it. No customer, logo or analyst placement is named here, because there are none to name. The one number we publish against ourselves is the 15.47% above, measured on one agent over 10 days.

Intercis is one person. That is the smallest number of people who can see your tool inputs, and the smallest number who can fix a bug. These are the machines your prompts and tool inputs land on: Railway runs the proxy, Supabase holds the database and the audit log, Vercel serves this site and the dashboard, Sentry takes our errors and Stripe takes the payment.

Read the security page

When the policy says no, what reaches the runtime?

Your agents are already running, and the platform you pay for sells a control for them too. Put whichever one you already pay for next to this page and look for the sentence that says what reaches the runtime after a deny. Ours is near the top of this one. Installing this is a base URL change and 2 headers, only 1 of them required, and you can cancel in the middle of the 90 days.

How it compares

One question decides most of these: when the policy says no, what does the runtime receive? Each page below answers it for one vendor from their own public pages, dated, and says where those pages do not.

Why the control sits outside the agent