Skip to content

Blog

Building a Security Auditor Into a Coding Agent

Building a security auditor into a coding agent: a read-only sub-agent runs SAST, secret, and dependency scans and reports CWE/OWASP findings, built on the principle that a scan that did not run must never read as clean.

An AI agent that writes code should be able to check it, too. But a security tool has a failure mode no other tool has: a false ‘clean’ is worse than no tool at all, because it manufactures unearned confidence. Here is how we built a read-only security auditor into CodeBuddy around that single idea.

Introducing CodeBuddy: A Deep-Agent Coding Assistant for Your Editor

Introducing CodeBuddy, a deep-agent coding assistant. Any model, in your editor.

Most AI coding tools autocomplete a line and stop. CodeBuddy is a deep agent: it plans the task, delegates isolated sub-jobs to eight specialist subagents, runs ~30 tools against your real files, verifies the result, and retries (failing over to another provider) when a step breaks. Open source, runs in your editor, uses your own keys. Here’s what that means and why we built it.

Building a Coding Agent with LangGraph Deep Agents

Building a coding agent with LangGraph Deep Agents: one createDeepAgent call, five seams (model, tools, backend, subagents, middleware) returning a compiled LangGraph graph.

A single LLM loop can’t reliably run a multi-step coding task. CodeBuddy is a deep agent instead: one createDeepAgent call composed from five seams (model, tools, file system, subagents, middleware) returning a compiled LangGraph graph, with four layers of self-healing. This is the whole architecture from the inside, grounded in the real code and file paths.

The Prompt Is the Smallest Part: Context Engineering as a Middleware Stack

Context engineering as an ordered middleware stack: seven layers (token budget, PII redaction, output sanitizer, compaction, skills, persona, domain policy) narrowing onto a single model call at the core.

Open the harness and the system prompt is 40 lines you could read over coffee. Print what the model actually receives on the same turn and it’s 4,000 tokens: a persona block, a skills catalog, a summarized history, tool schemas, spotlighted retrieval. None of it is in the prompt file. That gap is context engineering, and if you let it be ad-hoc string concatenation it rots. Here’s how we made it an ordered middleware stack instead, and why the order is load-bearing.

Giving Agents Memory: An Async ETL Pipeline, Not a Vector Database

Conversation → extract → consolidate → store → inject: memory as a pipeline

A customer tells your agent their budget on Monday. On Tuesday they come back through a different channel and it has forgotten. ‘Give it memory,’ everyone says, as if that means plugging in a vector database. It doesn’t. Memory is an extract-consolidate-store-inject pipeline that runs after the turn, and almost all the hard parts are policy, not storage. Here’s how we built one, grounded in the actual code.

Recursive Dispatch: When Agents Write the Plan and Other Agents Do the Work

Agents That Call Agents: recursive dispatch in practice

Ask an agent to review 300 files and somewhere around turn 40 it starts losing the plot. The fix isn’t a smarter model; it’s a different architecture. Here’s how recursive dispatch works, why we wired it into CodeBuddy, and everything that surprised us along the way.