If your agent outputs JSON — and most tool-calling agents do — you probably validate it against a schema. Required fields present, types correct, enums in range. That layer is necessary, and it rejects malformed output cheaply. But a schema validates the shape of an output. It says nothing about the intent of the values inside that shape, and in agent systems the dangerous failures are semantic.

What schema validation actually guarantees

A JSON schema for a shell-execution tool might look like this:

{
  "command": "string",      // required
  "args": ["string"],       // required, minItems 0
  "timeout_ms": "integer"   // >= 1
}

Every one of these payloads passes validation:

{"command": "ls", "args": ["-la"], "timeout_ms": 5000}
{"command": "rm", "args": ["-rf", "/"], "timeout_ms": 5000}
{"command": "curl", "args": ["http://169.254.169.254/latest/meta-data/"], "timeout_ms": 5000}

The schema was designed for well-behaved callers. A model is not a well-behaved caller — it is a text generator whose outputs are influenced by user prompts, retrieved documents, and tool results it fetched a moment earlier. Shape-correct data can carry a destructive command, a cloud-metadata URL for SSRF, an access key pasted into a free-text field, or prompt-injection instructions smuggled inside retrieved content.

A field-by-field example

Consider an agent that manages infrastructure tickets and can call a run_diagnostic tool with a structured payload. The schema requires a host string, a checks array, and a notify email string. Now consider the output the model actually produces after reading an internal runbook that an attacker influenced:

{
  "host": "10.0.0.4:2375",                // Docker daemon, not the ticket host
  "checks": ["network", "disk"],
  "notify": "[email protected]",
  "notes": "AKIAEXAMPLEKEY / wJalrExampleSecret"  // field exists in schema
}

Every field type checks. The schema is satisfied. Semantically, the output targets an internal administrative port (SSRF against the Docker API), exfiltrates a result to an outside address (integrity failure on the notification action), and carries a credential in a free-text field (leakage). Schema validation’s answer to all three is “valid.” This is not a hypothetical path — it is the shape of the disclosed MCP ecosystem incidents our detection rules are built from.

The seven dimensions of an output

The Correctover Conformance Standard (CCS) models agent output across seven verification dimensions: Structure, Schema, Latency, Cost, Identity, Integrity, and Security. Schema validation covers the shape part of Structure and Schema — roughly two dimensions. It has nothing to say about:

Teams that ship schema-only validation are signing off on an output contract while leaving the security, integrity, and identity clauses unchecked.

Why keyword matching makes the problem worse

The obvious next step is to add detectors — grep for exec, flag eval, block curl. Naive substring matching fails in both directions.

False negatives are straightforward: ex""ec, obfuscated arguments, indirect invocation, and context-dependent abuse all slip past a substring. False positives are the subtler problem, because they train teams to ignore the scanner. exec() inside a mocked test harness, a safe-eval library call in a sandboxed REPL tool, and subprocess.run with a fixed argument array are not vulnerabilities — but a keyword scanner flags all of them.

When we ran a keyword-based pass over the public AgenticX repository, it reported 27 CRITICAL findings. Running the same code through our contextual semantic intent engine reduced those 27 CRITICAL findings to zero: every flagged call site was an execution inside a sandboxed or mocked context that a pattern-only rule cannot distinguish from production use. Noise at that volume means the scanner gets turned off.

Semantic intent checking, explained

Semantic intent checking evaluates what the output does in context rather than what strings it contains. In Correctover’s engine, 116 detection rules reason about the data flow from model output to sensitive sinks — shell execution, network egress, filesystem writes, credential stores — and classify each path by its surrounding context: whether a shell call uses a fixed command with argument arrays, whether execution is wrapped in a sandbox, whether a URL points at an allowlisted host or a link-local metadata address.

The same token sequence gets different verdicts in different contexts:

The rules are grounded in publicly disclosed CVEs from the MCP ecosystem — the exploit patterns that have actually shipped in agent tooling, rather than hypothetical strings.

What “intent” means technically

It is worth being precise, because “semantic” is an overloaded word in AI tooling. The engine does not ask an LLM to judge whether an output feels safe. It applies deterministic, data-flow rules: identify where model-controlled values enter the output, trace each value to the sinks the output can drive (process execution, network requests, file paths, log records), and evaluate the guards on that path (allowlists, quoting, sandboxing, argument isolation). The verdict is reproducible for a given code version and input — which is what allows the same rules to run in CI as tests and in an audit as evidence.

Context classification is where the 116 rules earn their keep. exec in a test fixture that mocks the process module never reaches a real kernel; exec in a request handler with a tainted argument does. A network call to an allowlisted SaaS host is integration; the same call to a link-local address is SSRF. The engine reads the surrounding call structure — invoked function, arguments form, environment, and guard presence — rather than matching substrings, which is why the AgenticX pass went from 27 keyword-driven CRITICAL verdicts to zero contextual ones.

Fast enough to run in line

A second validation layer only helps if it runs on every output. For reference, our microbenchmarks of the core verification path — timed around the regex and core check execution only, excluding network and model time — measure a P50 around 7.5 microseconds, with the Node-side binding measuring around 2.7 microseconds P50 for the same core check scope. These are microbenchmark figures for the verification primitive itself, not end-to-end request latency; the point is that intent-level checks cost less than a rounding error against a multi-second model call, so there is no performance argument for skipping them.

A layered output contract

The practical setup is two gates, not either/or:

  1. Schema gate: reject malformed output immediately — cheap, strict, unforgiving.
  2. Semantic intent gate: inspect values for injection, SSRF, credential leakage, and prompt injection in context, using rules that understand sandboxed and safe-eval patterns.

For runtime verification of agent outputs in Python, the ccs-verifier package on PyPI implements the CCS dimension checks. For configuration-level hygiene around MCP servers, npx correctover-scan covers the deployment side (package: correctover-scan; source: GitHub).

Schema tells you the output is well-formed. Intent checking tells you it is safe. Production agents need both.

Free 1-page audit summary — reply with your repo link

Want a quick read on your own agent or MCP setup before a full review? Send us a repository link and we’ll return a free 1-page scan summary covering the findings our semantic engine flags, with concrete fix pointers.

Get your free 1-page summary

Details: correctover.com/agent-audit.html