Skip to content

Sandboxed scripts

CodeBuddy runs user-authored scripts inside a QuickJS WASM sandbox — a third trust tier that lets you customize behavior without exposing Node capabilities to arbitrary code.

TierRuns whereTrust
Extension host codeNode.js, full APIFull — shipped by us
User/team scriptsQuickJS WASM sandboxPartial — behavior only; no I/O except via curated host bindings
LLM outputNever executed directlyUntrusted — sanitized and wrapped

Every user script runs in the middle tier: it can compute, transform, and read (via bindings), but cannot write files, spawn shells, or reach the network on its own.

ConventionPurposeBindings available
skill.js next to skill.mdSkill scripts — extend a skill with programmatic logicNone (string I/O)
.codebuddy/transforms/<tool>.jsRewrite a tool’s raw output before the LLM sees itNone (string I/O)
.codebuddyignore.jsPredicate fallback / complement to the glob-based ignore fileRead-only pack
{{js: EXPRESSION}} in AGENTS.md / rules.mdDynamic prompt fragments (e.g. {{js: host.gitBranch()}})Read-only pack
.codebuddy/commands/*.jsCustom slash commands (/my-command)Read-only pack

All five ship today.

Bindings are the only way a sandboxed script reaches CodeBuddy or the OS. Each is Zod-validated at the argument boundary and permission-tagged for audit.

Read-only pack (available in every non-string-only context):

BindingReturns
host.readFile(path)File contents (path validated against workspace root)
host.grep(pattern, opts)Ripgrep matches over the workspace
host.glob(pattern)Path list matching the glob
host.gitBranch()Current git branch
host.workspaceInfo()Workspace root, project type, framework detection

Every path argument runs through WorkspaceIdentityService.validatePathWithinWorkspace() before hitting the filesystem — no escape to /etc/ or ~/.ssh/, TOCTOU-resistant.

  • File writes — every write flows through the normal tool path with diff-review approval. There is no host.writeFile.
  • Network access — no host.fetch, no XHR, no sockets. If you need HTTP, use an MCP server.
  • Arbitrary shell — no spawn, no exec, no shell strings. Terminal work goes through the modal-approval tool path.
  • Long-running scripts — per-execution budgets on wall-clock time, memory, host-call count, stack size, and I/O size are enforced at the WASM level (default script timeout ~5 s). Scripts that exceed a budget are terminated cleanly.
  • Cross-eval persistence — fresh VM per invocation (call mode). State does not leak between calls.
  • Fresh VM per invocation — every call creates a new QuickJS runtime + context and disposes it afterward. There is no warm pool; state never leaks between calls.
  • Interrupt handler — a wall-clock deadline (Date.now() + wallTimeMs) plus a signal-abort check. Runaway scripts terminate at the deadline; the calling code falls back to its non-sandboxed path. The handler is installed on every execution and cannot be disabled.
  • Fail-open on script errors — a broken user script never blocks the agent; failures surface as a warning in Output > CodeBuddy and the caller falls back to its default behavior. The sandbox is added value, not a gate — so a transform is best-effort and should not be relied on as a security boundary (e.g. as the sole means of redacting secrets from tool output).

Create .codebuddy/transforms/read_file.js in your workspace to filter every read_file tool result before the LLM sees it:

// input: the raw tool output as a string
// return: the transformed string (return input unchanged to no-op)
export default function (input) {
// Strip lines starting with "DEBUG:" from any file read
return input
.split('\n')
.filter((line) => !line.startsWith('DEBUG:'))
.join('\n');
}

Save the file. Next time the tool runs, the transform picks up. No restart. No configuration.

Same shape for skill scripts (skill.js beside skill.md), slash commands (.codebuddy/commands/deploy.js), and predicates (.codebuddyignore.js) — each has its own contract; see the linked pages.

LLM-driven code interpreter (experimental)

Section titled “LLM-driven code interpreter (experimental)”

The sections above cover user-authored scripts. QuickJS also powers a second, distinct surface: an LLM-driven REPL (@langchain/quickjs CodeInterpreterMiddleware) where the agent writes and runs its own JavaScript to orchestrate tool calls.

  • Experimental, on by default — codebuddy.experimental.codeInterpreter ships true; in-REPL task() fan-out (codebuddy.experimental.codeInterpreterSubagents) also defaults true. Set either to false as an escape hatch.
  • Read-only tool surface — the REPL’s pre-tool-call set is read-only (get_diagnostics, ls, grep, search_symbols, search_vector_db); writes still route through the normal modal-approval tool path.
  • Gated off in Plan mode — attach condition is codeInterpreterEnabled && !planMode, so a read-only Plan turn never spins up the REPL. See Modes.
  • Fenced fan-out — when in-REPL subagent dispatch is on, per-subagent filesystem permissions and tool-list scoping keep a prompt-injected write from reaching the workspace. See Subagents.
  1. Every binding has a Zod-validated signature — no argument confusion via prototype pollution or type coercion.
  2. Every path argument goes through validatePathWithinWorkspace() — symlink-resolved, workspace-scoped.
  3. Adding a new binding requires a security-review checkbox on the PR — the audit surface only grows deliberately.
  4. WASM isolation — even a JIT bug inside QuickJS cannot reach the extension-host process directly. Memory and control flow are contained.
  5. No host binding may call write-side tools (Terminal.executeAnyCommand, EditFileTool.execute, DeleteFileTool.execute). Enforced at review time and by the binding framework’s own contract.
  • Security model — how the sandbox fits the broader threat model
  • Project rules — where {{js: …}} fragments live
  • Skills — how skill.js extends a skill
  • MCP — the other integration path: external process vs sandboxed script