“We tested it and it works” is not a security result. Teams building agents have the same blind spot as teams building any complex system: the author of a tool handler knows what the handler is supposed to do, which makes it genuinely hard to see what a model can make it do. An independent audit exists to produce evidence about the second question — and, increasingly, to produce it in a form a buyer or reviewer can read.
The blind spot that creates the demand
Agent teams test against the happy path because the happy path is what ships the demo: a user asks for a file conversion, the right tool runs, the right answer comes back. Security failures live on the unhappy paths that internal testing under-samples — a prompt that tries to redirect the agent, a document containing embedded instructions, a malformed tool result the parser coerces into something dangerous. An audit is a structured visit to those paths, performed by reviewers whose only prior assumption is that model output is untrusted.
What an independent agent output audit covers
An agent output audit reviews the complete path from model output to side effects: the agent code, MCP server configurations, tool definitions and descriptions, and a sample of runtime traces. Findings are organized against the seven CCS verification dimensions — Structure, Schema, Latency, Cost, Identity, Integrity, and Security — and the security analysis runs a corpus of 116 semantic intent rules over the code and traces.
The rule set focuses on the exploit classes that matter in agent deployments:
- Command injection — model-controlled text reaching shell or script-execution sinks in tool handlers.
- SSRF — tool-driven requests to internal or cloud-metadata addresses, including redirect-based bypasses.
- Cloud credential leakage — secrets in configs, logs, error messages, or free-text outputs.
- Prompt injection — untrusted fetched content steering tool calls, and injection text passing through outputs.
The rules are grounded in publicly disclosed CVEs from the MCP ecosystem: each detection corresponds to an exploit pattern that has appeared in shipped agent tooling, which keeps the signal tied to real failures rather than theoretical ones.
What the deliverable contains
A Correctover Agent Output Audit runs over roughly five working days and returns a PDF report with:
- Findings ranked by severity, each with the affected file or tool and the exact trace or code path.
- Concrete remediation guidance — the specific code or configuration change that closes the finding.
- A dimension-by-dimension coverage map against CCS, showing which controls are verified and which gaps remain.
- A reproducible methodology section, so an internal reviewer can re-run the checks and confirm results.
Why independent
Two reasons. First, context asymmetry: an external review does not know what the code “should” do and so reads what it actually does — including the tool handler that safely wrapped a binary until a shell: true landed in last month’s refactor. Second, transferability: enterprise security teams and procurement reviewers do not accept self-attestation for controls that touch production data and credentials. A report from an independent review is evidence that travels; a Slack message from the agent team is not.
How to self-check first
An audit goes faster — and finds fewer embarrassing basics — when the configuration and wiring are already clean. Three self-service checks cover the ground you can cover yourself:
1. Scan every MCP configuration
The open-source correctover-scan CLI inspects MCP config files for the deployment-level failure modes: credentials inline in JSON, plaintext transport, missing timeouts, SSRF exposure, absent kill switch, unpinned dependencies.
npx correctover-scan # auto-detect configs in the repo
npx correctover-scan -f json # machine-readable output
Package: correctover-scan on npm; source: github.com/Correctover/correctover-scan.
2. Gate the repository in CI
The correctover-scan GitHub Action runs the same checks on every push and pull request, can fail the build on critical findings, and emits SARIF for the GitHub Security tab. Once green, the baseline cannot silently regress.
3. Verify outputs at runtime
For the output side rather than the config side, ccs-verifier on PyPI implements the CCS dimension checks for agent responses — structure and schema first, then the security and integrity checks that catch payloads a schema accepts. Wire it into your agent’s tool-call path so every output passes a verification contract.
What self-checking does not cover
Be clear-eyed about the boundary. Config scanning finds wiring mistakes. Runtime verification catches bad outputs as they occur. Neither one answers the code-level, cross-tool questions that require reading the agent’s source and traces in context: whether tainted model input reaches a shell sink in a specific handler, whether a tool’s description plus fetched content can steer an agent off-policy, or whether the combination of two individually-passing tools creates a confused-deputy path. Those require semantic analysis — in our case, the 116-rule intent engine applied to your repository — and human review of what it flags.
How findings get rated
Every finding in the report carries a severity derived from two factors: reachability (can model-controlled data actually reach the sink?) and authority (what does the process have access to when it fires?). A shell sink with tainted arguments in a tool that holds no credentials and runs in a disposable container is serious but bounded; the same sink in a process holding cloud production keys is critical. SSRF that can only reach public hosts is a hygiene issue; SSRF one redirect away from the instance metadata endpoint is critical. This keeps the report prioritized by exploitable impact rather than by how alarming the code looks.
What to hand the auditor (and why it helps)
Audit quality scales with access. The useful inputs are: the agent repository (or the relevant tool-handler subset), all MCP configuration files across dev and production, tool and prompt definitions, a sample of runtime traces showing real tool calls, and a short description of the execution environment (where tools run, what credentials they hold, what network they can reach). With those, the review can trace taint end-to-end instead of guessing at deployment topology. Clean the traces of third-party personal data first; credentials found in the repo are themselves a finding and will be flagged as such.
After the report
The report is a snapshot, and agent code changes weekly. Teams get durable value from it by turning each finding into one of three artifacts: a code fix merged with the audit as reference, a CI gate (the scanning action for configuration, runtime verification for outputs) that prevents the same pattern returning, or an accepted-risk entry with an owner and a review date. Findings that do not become one of those three reappear in the next audit — which is also how reviewers measure whether a security program is real.
Traces: the artifact teams forget
Code review shows what a tool handler can do; traces show what the agent actually does. A representative trace set — full tool-call sequences with arguments and outputs, captured from staging or a recorded evaluation suite — reveals the behaviors static reading misses: chained tool calls where the output of one fetch becomes the argument of the next exec, retry loops that amplify a malformed result, and descriptions that nudge the model toward tools it should rarely use. When collecting traces, include adversarial sessions deliberately (injection in fetched content, out-of-policy user requests), because happy-path traces only confirm happy-path behavior and the audit’s value is in the unhappy paths.
Self-check vs. audit vs. runtime: three layers, one pipeline
The three layers are not alternatives, they are stages of the same assurance pipeline. Self-service scanning shifts left and removes the mechanical findings at zero marginal cost. The independent audit is a point-in-time, human-validated assessment that produces externally credible evidence. Runtime verification is the continuous control that catches what neither static review nor a snapshot can: behavior in production, on inputs nobody anticipated. Teams that buy only the audit get a report that rots; teams that run only self-checks get findings nobody validates in context. The pipeline works because each layer feeds the next — audit findings become runtime rules, and runtime anomalies that recur become audit focus areas.
A sensible sequence
- Run the local scan and fix everything critical. Add the CI action.
- Add runtime output verification on tool calls.
- Request the independent audit for the code-and-trace layer, ahead of an enterprise review, a funding diligence step, or a major production rollout.
The self-check steps cost you an afternoon and remove the findings you did not need an auditor for. The independent review then focuses on what automated self-checking structurally cannot see.
Free 1-page audit summary — reply with your repo link
Want a quick read on your own agent or MCP setup before a full review? Send us a repository link and we’ll return a free 1-page scan summary covering the findings our semantic engine flags, with concrete fix pointers.
Details: correctover.com/agent-audit.html