How to test an agent boundary

Three differentials, one A/B rule, and the control key I forgot twice.

An agent boundary is the thing that decides what runs: the approval/permission layer, the sandbox, and the configuration the agent reads from the directory it was pointed at. All three fail in the same way — the component that enforces the boundary runs in the environment it polices, and takes its parameters from untrusted input.

Crash fuzzing is the wrong tool here

On a well-tested command parser, 19.5M and 1.6M libFuzzer executions produced zero crashes. The parser was fine. The bug was downstream: what the parser decided vs. what the shell did. A crash oracle cannot see that, because nothing crashes — the command just means something other than what the permission layer thought.

So the oracle has to be differential.

Three differentials

  1. Belief vs. reality (the shell). Build a fake PATH where every leaf command is a logger and wrappers (bash, env, find, xargs, script, timeout…) exec the real binary with PATH still pointed at the stubs. You get the true execution chain, not the first token. Nothing destructive ever runs.
  2. What the vendor pinned vs. what git still executes. Reproduce the vendor's own arguments and environment — from source, or from the strings in a shipped binary — then run the commands they actually issue, inside a repo carrying a marker program.
  3. What the repository gets to say. Walk the agent's source for the config files it reads from the workspace; for each, does this let the repository choose an executable? Then ask whether the product has a workspace-trust concept at all.

The A/B rule

Every finding is an A/B pair: identical payload, one variable different — the suspected evasion. Same repo, same argv, with and without the sink pinned. Anything weaker gets triaged down, because “it executed something” is not a finding; “it executed something because of this one difference” is.

The control key (the mistake I made twice)

A sink that does not fire is only meaningful if a pinned sink also failed to fire in the same run. My first two probes reported false negatives because I checked the exotic sink and forgot the control: the marker never ran at all — I had broken the fixture, not found a result.

Every run now carries reference: everything pinned. If that profile fires, the run is invalid.

Why this generalises

The three differentials are the same shape: two things that should agree, one under the attacker's influence.

If you build or audit an agent, the question for each component is one line: where does this take its parameters from, and who controls that?

Tools