Building a Security Auditor Into a Coding Agent
An AI coding agent will happily write you an Express route, a database query, and a Dockerfile in one turn. What it has not traditionally done is turn around and ask whether any of that is safe to ship. The agent that produces the code and the tool that reviews it have lived in different worlds.
CodeBuddy now closes that gap with a security-auditor sub-agent: a read-only specialist that scans your code for vulnerabilities, hunts for committed secrets, and checks your dependencies against known advisories, then reports findings tagged with the relevant CWE and OWASP categories and ranked by exploitability. It never edits and it never fixes. Anything it recommends goes back through the normal diff-review flow, where you stay in control.
The scanners themselves are the boring part. The interesting part is a single design principle that shaped every decision underneath them.
Key Takeaways
Section titled “Key Takeaways”- For a security tool, a false negative is the worst failure mode. Worse than a false positive, and worse than no tool at all, because a clean report you cannot trust removes the caution that would otherwise protect you.
- A scan that did not actually run must never read as clean. Every place the auditor could quietly under-report (an engine missing, a check timing out, a file skipped) is surfaced as a loud, explicit note instead of a silent green result.
- A secret scanner must never leak the secret. Matches come back redacted. The raw value never enters the model’s context, the logs, or the editor’s Problems panel.
- Read-only means read-only by construction, not by good intentions. The auditor holds scanning and read tools and no write tool, and those tools are granted to it by exact name so they cannot leak into other roles.
- No install, cross-language. Dependency auditing works out of the box against a public vulnerability database, covering more than just the Node ecosystem, and every engine degrades to a simpler one when the richer option is not present.
The shape of it
Section titled “The shape of it”The auditor is one read-only sub-agent backed by three tools, each of which returns findings in a single shared shape (a rule id, a CWE, an OWASP category, a severity, a file and line, a message, and a remediation). Findings render both in chat and in the VS Code Problems panel, so they behave like first-class diagnostics rather than a wall of text.
scan_codeis static analysis (SAST). It looks for the classics: weak or broken hashing, disabled TLS verification, command injection, and dynamic code execution. When a full analysis engine is available it uses it for broad, multi-language coverage; otherwise it falls back to a small, high-precision built-in rule set.scan_secretsfinds committed credentials: provider API keys, tokens, private keys, and high-entropy strings that look like secrets in a credential-shaped context.scan_dependenciesis software-composition analysis (SCA). It checks your installed dependencies against known advisories and flags unpinned versions, across languages, not just JavaScript.
You do not invoke any of this by name. You ask for it: “do a security audit of this workspace,” or “scan for injection bugs and leaked secrets,” or “audit my dependencies.” The agent picks the right tools.
The principle: a clean report is a promise
Section titled “The principle: a clean report is a promise”Most tools fail loudly. A formatter that crashes formats nothing, and you notice. A test runner that cannot start prints an error, and you notice. A security scanner is different, because its success case and its silently-did-nothing case look identical from the outside. Both say “no issues found.”
That symmetry is the whole problem. If a secret scanner reports a workspace clean when it never actually opened the file with the key in it, it has not merely failed to help. It has done something worse: it has told you the code was checked when it was not, so you ship with more confidence than if the tool had never existed. The tool removed the very caution that would have saved you.
So we treated one rule as non-negotiable: a scan that did not actually cover something must never render as one that did. Everything else about the auditor is downstream of that.
What the principle actually buys you
Section titled “What the principle actually buys you”Stated as a slogan it sounds obvious. In the implementation it means resisting a dozen small temptations to swallow a failure and return an empty result.
Failures are loud, not silent. If the SAST engine is not installed, if the dependency database cannot be reached, if a check times out on a large repository, the auditor does not hand back a blank list. It returns an explicit note: “results are incomplete, do not treat this as a clean audit,” and it says which part did not run. An empty findings list from a real scan and an empty findings list from a scan that never happened are two completely different facts, and the report keeps them different.
Coverage is honest. When the auditor skips a file (too large, binary, or explicitly allowlisted) or truncates a very large result set, it says so, in the same place it reports findings. A mostly-skipped scan should read as a mostly-skipped scan, never as a clean one. The moment “we did not look here” is quieter than “we looked and found nothing,” the report starts lying by omission.
Secrets stay secret. The corollary on the secret-scanning side is that the tool must find the credential without becoming the thing that leaks it. So scan_secrets reports a redacted preview only. It tells you enough to locate and rotate the key and nothing more. The full value never reaches the model, the transcript, or the logs. A secret scanner that pastes the secret into a chat window has solved one problem by creating a worse one.
Read-only by construction
Section titled “Read-only by construction”The auditor’s job is to report, so it should be structurally incapable of doing anything else. It is defined as a read-only role that holds the three scan tools plus a handful of read-only helpers for cross-referencing (diagnostics, symbol search, the dependency graph, git history) and no write tool at all. The scanning tools are handed to it by exact name, which means they cannot accidentally end up in the hands of a different, write-capable specialist.
There is one careful distinction inside that. Two of the scanners are purely local and safe to run even in the most restricted, side-effect-free modes. The dependency scanner is different, because checking dependencies against known advisories means reaching a vulnerability database over the network. That single outbound step is treated as exactly that: a network action, excluded from the side-effect-free modes and routed through CodeBuddy’s outbound-safety checks like every other request the agent makes. “Read-only” and “makes a network call” are not the same promise, so we do not let one quietly stand in for the other.
No install, and it still works
Section titled “No install, and it still works”A security feature that only runs after you install three CLIs is a security feature most people never turn on. So the auditor is built as a ladder. Dependency scanning prefers a locally installed scanner if you have one, falls back to a public vulnerability API that needs no install and no signup, and falls back again to your package manager’s own audit. Static analysis prefers a full engine (on your machine or in a container) and falls back to a compact built-in rule set. At every rung, if the richer option is not there, the auditor uses the simpler one, and, per the principle above, tells you it did.
The payoff is coverage across ecosystems out of the box, so a Python or Go service gets a real dependency audit, not a polite “install something first.”
The quiet philosophy
Section titled “The quiet philosophy”There is a temptation, building tools for an agent, to optimize for the happy path and let the edges fail softly. For most tools that is fine. A code formatter that occasionally no-ops is a minor annoyance.
A security tool does not get that grace, because the cost of its quiet failure is not annoyance, it is a false sense of safety at exactly the moment safety matters. The engineering that makes the auditor trustworthy is almost entirely about the unhappy paths: the timeout, the missing engine, the skipped file, the unreachable database. Getting those to fail loudly, every time, is less glamorous than the scanning itself and far more important.
A green checkmark is a promise. The only job that matters is making sure the auditor never makes that promise it did not keep.
The security auditor ships in CodeBuddy today. Ask it to audit your workspace, and read the notes as carefully as the findings.