Building a Coding Agent with LangGraph Deep Agents
Ask a single-loop coding assistant to “refactor the authentication module to use OAuth2” and watch it come apart. It tries to hold the entire plan, every file it has touched, and every decision in one context window. It forgets step three by step seven, loses track of what it already changed, and starts inventing file paths that were never there.
The failure isn’t the model. Give the same model a harness that plans on paper, delegates to fresh contexts, and reads files instead of memorizing them, and it finishes the task. CodeBuddy is that harness: an open-source coding agent that lives in your editor, built on LangChain’s Deep Agents running on LangGraph. This post is the architecture from the inside: the one call everything hangs off, the five seams that feed it, and the production traces that shaped each one. Every code excerpt and file path below is from the shipping extension.
Key Takeaways
Section titled “Key Takeaways”- The harness is the product, not the prompt. A “shallow” agent is one LLM loop calling tools until it answers. A “deep” agent externalizes what the loop can’t hold: it plans to a todo list, offloads state to a file system, and delegates subtasks to isolated subagents. Deep Agents gives you those primitives; CodeBuddy’s job is to bind them to a real IDE without the model ever knowing the IDE is there.
- The interesting engineering lives in the seams. Bridging an AI framework to a live editor surfaces problems a chat wrapper never sees: a backend that routes virtual paths to real files, subagent permissions that fail loudly instead of silently leaking a write tool, and middleware that must add per-turn context without invalidating the prompt cache. Get that last one wrong and you pay for it, literally.
- Reliability is a design layer, not a try/catch. Models hallucinate, providers rate-limit, contexts overflow. CodeBuddy classifies the failure and picks a matching recovery across four layers. An auth error is not retried (it would just fail again); a rate limit trips provider failover with a per-reason cooldown.
CodeBuddy’s core is a single createDeepAgent() call composed from five seams: a model factory (eight providers, bring your own key), a tool substrate (~30 tools, each wrapped in four cross-cutting layers), a CompositeBackend that routes /workspace to real editor files, /docs to cross-session storage, and everything else to ephemeral state, a pool of eight isolated subagents with role-scoped, permission-fenced tools, and a middleware stack (the Deep Agents harness plus three custom layers). The call returns a standard compiled LangGraph graph, so streaming, checkpointing, and human-in-the-loop come for free. Persistence uses a custom sql.js checkpointer (native SQLite doesn’t survive extension bundling). Resilience is four layers: fail-open tool recovery, provider failover, context compaction, and checkpoint restore. The patterns are portable; the war stories are ours.
The harness, not the model
Section titled “The harness, not the model”Deep Agents is LangChain’s “batteries-included” agent harness. It sits on top of LangGraph the way a pre-built car sits on top of an engine: LangGraph gives you the graph runtime (state, checkpointing, streaming, human-in-the-loop); Deep Agents wires in the four things every long-horizon agent needs and every from-scratch build reinvents badly.
- Planning. A
write_todostool lets the agent decompose a task into a checklist and tick items off as it goes. The plan lives in state, not in the model’s head, so a twelve-step task doesn’t degrade as the window fills. - Sub-agents. A
tasktool spawns a specialist with its own context window, its own filtered tools, and its own middleware. It runs, returns a singleToolMessage, and is destroyed. No cross-task context pollution. - A file system.
ls,read_file,write_file,edit_file,glob, andgrep, backed by a pluggable storage layer. This is how the agent survives without a million-token window: it writes things down and reads them back. - Detailed prompts. The system prompt is assembled at runtime from project rules, memories, and skills, not baked in as a static string.
Crucially, createDeepAgent returns a compiled LangGraph graph. That single fact is why CodeBuddy gets token-by-token streaming, durable checkpoints, and human-in-the-loop interrupts without building any of them: they are LangGraph features, and CodeBuddy’s agent is a LangGraph graph.
One call, five seams
Section titled “One call, five seams”Here is the actual composition point, src/agents/developer/agent.ts:
return createDeepAgent({ model: this.model, // seam 1 tools: this.tools, // seam 2 systemPrompt: this.getSystemPrompt(), backend: backendFactory, // seam 3 store, checkpointer: checkPointer, name: "DeveloperAgent", subagents, // seam 4 interruptOn: interruptConfig, middleware, // seam 5 permissions: this.buildFilesystemPermissions(),});Everything else in this post is one of those five arguments. Let’s take them in turn.
Seam 1: the model
Section titled “Seam 1: the model”CodeBuddy is bring-your-own-key. A single factory, buildChatModel() in chat-model.factory.ts, switches over eight working providers and returns a ChatAnthropic | ChatOpenAI | ChatGroq | ChatGoogleGenerativeAI:
anthropic · openai · groq · gemini · deepseek · qwen · glm · locallocal points at Ollama or any OpenAI-compatible endpoint, so your code can stay entirely on your machine. The factory is also a small museum of provider quirks, each comment a scar from a real failure:
- Default to
gpt-4.1, notgpt-4o. On some accountsgpt-4o404s asmodel_not_found, and worse, it “fails to terminate the agent loop, spins to the recursion limit.” - OpenAI’s SDK reads
apiKey, not theopenAIApiKeyalias in the type definitions, which the runtime silently drops, producing a baffling “Missing credentials.” - Groq wants
baseUrl(lowercase); the wrong casing is silently ignored and you get a 401 only once you’re in agent mode.
When multi-model mode is on, a resolver routes each subagent to a tier (a power tier of anthropic/openai, a fast tier of the rest) and wraps every override in a FallbackChatModel back to the parent model, so a flaky secondary provider degrades instead of failing the run. Per-provider reasoning knobs are resolved by pure capability modules: Anthropic models take adaptive thinking on 4.6+ but the older budget_tokens shape (plus a maxTokens bump) before that, because sending budget_tokens to a 4.7/4.8 model is a hard 400. OpenAI’s reasoning detection is a deliberately narrow regex, /^o\d/, so it matches o4-mini but never gpt-4o.
Seam 2: the tools, wrapped four deep
Section titled “Seam 2: the tools, wrapped four deep”Around 30 tool factories are registered in ToolProvider: file and AST-aware editing, terminal execution, ripgrep and semantic search, LSP queries, a full debugger integration over DAP, web search, a browser, git, and more.
The part worth studying is what happens to every tool on the way in. Each one is wrapped in four cross-cutting layers, and the order is the point:
caller → audit → dedup-cache → rate-limit → original invoke- Rate-limit is innermost so it counts real invocations for runaway-loop detection.
- Dedup-cache wraps outside the limiter on purpose, so a cached hit “isn’t a real tool invocation as far as runaway-loop detection is concerned.”
- Audit is outermost, so it “sees rate-limit blocks and cache hits alongside real invocations,” all written to
.codebuddy/logs/audit.jsonl.
The wrap is idempotent: each tool gets tagged with a Symbol("codebuddy.rate-limit-wrapped") and re-wrapping is a no-op, so the same tool passing through the pipeline twice never gets double-audited. In plan mode, the tool list is filtered to a read-only allowlist and the filesystem permissions deny all writes globally, so “planning” genuinely cannot mutate your repo.
Seam 3: the file system that bridges two worlds
Section titled “Seam 3: the file system that bridges two worlds”This is the seam I’m proudest of, because it’s where an AI framework meets a real editor and neither has to know.
Deep Agents’ file system is pluggable, and CodeBuddy passes a backend factory that builds a CompositeBackend: a router that dispatches virtual paths to different real backends.
const defaultBackend = new StateBackend(stateAndStore); // ephemeral agent stateconst routes = {};if (stateAndStore.store) { routes["/docs/"] = new StoreBackend({ ...stateAndStore }); // cross-session docs}routes["/workspace/"] = vscodeFsBackend; // real editor filesreturn new CompositeBackend(defaultBackend, routes);| Route | Backend | Purpose |
|---|---|---|
/workspace/* | VscodeFsBackend | Real IDE files, read and write actual code |
/docs/* | StoreBackend | Persistent, cross-session documents |
/ (default) | StateBackend | Ephemeral, in-graph agent state |
The agent calls read_file("/workspace/src/auth.ts") and has no idea it just touched your disk. It calls write_file("/docs/adr/001.md") and has no idea that persisted across sessions. That indifference is the whole trick.
Seam 4: eight subagents that can’t step on each other
Section titled “Seam 4: eight subagents that can’t step on each other”CodeBuddy ships eight specialist subagents, each a fresh agent with its own window, its own filtered tools, and a single-task lifecycle:
| Subagent | Job | Can write? |
|---|---|---|
code-analyzer | analyze structure, find bugs and anti-patterns | read-only |
reviewer | review changes for quality, security, performance | read-only |
architecture-expert | explain the codebase’s actual architecture | read-only |
doc-writer | write technical docs | /docs/** only |
architect | design architecture, write ADRs | /docs/adr/** only |
debugger | investigate errors, surface a root cause | yes |
tester | design and run automated tests | yes |
file-organizer | move, rename, refactor directory layout | yes |
Tools are filtered per role, and permissions are declared per role in a ROLE_PERMISSIONS map. Two design choices here are worth stealing.
The type system enforces that new roles get a policy. The permissive roles are marked with a literal "permissive" sentinel rather than just omitted, so adding a ninth role without deciding its permissions is a compile error, not a silent grant. You cannot forget.
Permissions fail loud, because the framework’s don’t cover everything. Deep Agents’ built-in enforcePermission guards the filesystem tools, but tools that write through other paths (terminal, VCS, browser) aren’t covered by that filesystem policy. So CodeBuddy scopes those per role via the tool list and runs its own write-leak detector at assembly time to catch a mismatch:
const violators = (s.tools ?? []).map((t) => t.name) .filter((n) => WRITE_LEAK_TOOL_NAMES.has(n));if (violators.length > 0) { getLogger().warn(`[subagents] SECURITY: '${s.name}' has restricted ...`);}A read-only subagent holding a write-capable tool gets flagged the moment it’s constructed, not the first time it does damage. There’s a matching fix on the other side: a read-only role like code-analyzer used to be told to persist its findings to /docs/code-reviews/, then get denied by enforcePermission mid-turn and crash the whole run. Now the prompt is generated from the actual policy, so a read-only subagent is instead told to return its findings.
Here’s a task moving through the whole system:
Seam 5: middleware, and the cache tax
Section titled “Seam 5: middleware, and the cache tax”The Deep Agents harness bundles its own middleware (filesystem tools, the subagent task tool, automatic summarization, and, for Anthropic and Bedrock, prompt caching of the static prompt). CodeBuddy adds its own layers via createMiddleware: memory (AGENTS.md), skills (.codebuddy/skills/), and custom middleware for session-awareness and per-turn context. (The wider argument that context assembly is an ordered middleware stack is its own post: The Prompt Is the Smallest Part, with Giving Agents Memory on the memory pipeline.)
The constraint these layers work under is the cache tax: Anthropic bills the first time it sees a prompt prefix and discounts every reuse, so any per-turn context has to be appended after the stable prefix or it re-bills the whole prompt every turn. So CodeBuddy’s session-awareness note (which files were already read) is delivered as a trailing message rather than baked into the system prompt, and volatile grounding rides LangGraph’s read-only runtime context channel instead of the checkpointed state. Both layers are strictly fail-open: on any error they pass the request through unchanged.
Persistence: a checkpointer that survives bundling
Section titled “Persistence: a checkpointer that survives bundling”createDeepAgent takes a checkpointer and a store. The obvious choice, @langchain/langgraph-checkpoint-sqlite, uses better-sqlite3, whose native binding does not survive being bundled into a VS Code extension. So CodeBuddy ships a custom BaseCheckpointSaver backed by sql.js (WASM SQLite):
This replaces the native-module-dependent
@langchain/langgraph-checkpoint-sqlitewhich usesbetter-sqlite3and fails to load in bundled VS Code extensions. sql.js is already a project dependency used by SqliteVectorStore.
It reuses a dependency already present for the vector store rather than adding one, and it means every agent turn is checkpointed to a real, queryable SQLite database that ships inside the extension. The same interruptOn config that enables human-in-the-loop is deliberately minimal: by default it gates delete_file for approval, since gating every edit would make the agent unusable. Other writes are shown as reviewable, undoable diffs — and requireDiffApproval gates all of them behind explicit approval when you want that.
Reliability: four layers, each matched to a failure
Section titled “Reliability: four layers, each matched to a failure”Production agents fail in categorically different ways, and the recovery has to match the category. Retrying an auth error just fails again; retrying a rate limit after a cooldown works. CodeBuddy layers four distinct mechanisms:
| Layer | Handles | Mechanism |
|---|---|---|
| Tool recovery | tool errors, malformed steps | fail-safe retry: only genuinely transient signals retry, with exponential backoff |
| Provider failover | rate limits, auth, billing, overload | classify into one of 8 reasons, cool down, switch provider |
| Context compaction | window overflow | 3-tier summarize-and-compact, keep the last few messages |
| Session restore | unrecoverable errors | resume from the sql.js checkpoint |
What the architecture buys
Section titled “What the architecture buys”None of this matters unless it does something a single loop can’t. Each of these real tasks is 20 to 50 tool calls across several subagents, more than any one window could hold:
- “Refactor the payment module.” Plans the steps, delegates the survey to
code-analyzer, rewrites the code itself, runs the suite withtester, and loops on failures until it’s green. - “Debug this production error.” Reads the stack trace, attaches to the debugger over DAP, inspects live variables, traces the root cause, writes the fix.
- “Add i18n support.” Scans every user-facing string, creates the translation files, updates the components, tests each locale.
The recursive-dispatch pattern pushes this further, letting the agent write a plan as code and fan work out to subagents through a sandboxed interpreter.
The numbers
Section titled “The numbers”| Metric | Value |
|---|---|
| Agent framework | LangGraph Deep Agents (deepagents) |
| Composition | one createDeepAgent() call, five seams |
| Subagent specialists | 8, isolated and permission-fenced |
| Core tools | ~30 (plus 6 filesystem primitives, plus MCP) |
| LLM providers | 8 (7 cloud plus fully-local) |
| Storage backends | 3 (IDE files, cross-session docs, ephemeral state) |
| Checkpointer | custom sql.js (WASM SQLite) |
| Self-healing | 4 layers (tool recovery, failover, compaction, restore) |
Try it
Section titled “Try it”CodeBuddy is open source (MIT) and runs in VS Code, Cursor, Windsurf, and VSCodium.
ext install fiatinnovations.ola-code-buddy- CodeBuddy: github.com/olasunkanmi-SE/codebuddy
- Deep Agents: github.com/langchain-ai/deepagentsjs · docs
- Related reads: context engineering as a middleware stack · giving agents memory · recursive dispatch