Skip to content

Giving Agents Memory: An Async ETL Pipeline, Not a Vector Database

Conversation → extract → consolidate → store → inject: memory as a pipeline

A customer tells your assistant their budget and their preferred language on Monday. On Tuesday they come back through a different channel, and it greets them like a stranger. The model is excellent; the product still feels amnesiac.

“Give it memory” usually gets read as “plug in a vector database.” In a production multi-tenant agent runtime we built on LangChain’s Deep Agents (createDeepAgent, the JavaScript sibling of LangGraph), memory turned out to be an async ETL pipeline that runs after the conversation ends: extract, consolidate, store, then inject on the next turn. Almost every hard problem lives in the policy (what to remember, when a fact changes, whose memory this is), not in the storage.

Architecturally it’s a Deep Agents graph, and the memory system hangs off two of its seams: a wrapModelCall middleware on the read path, the run-terminal hook on the write path. It’s the same runtime as the previous post on recursive dispatch, viewed from a different seam. This post walks the whole memory system: the store, the consolidation engine, the identity that keys it, the pre-turn injection, and the evals that keep it honest. Every file path and code excerpt below is from the shipping implementation.

Memory as an async ETL pipeline: after a turn, a debounced scheduler triggers an engine that makes one LLM call to extract and consolidate facts, writes them to the store (pgvector, customer profile, watermark), and injects them into the next turn, all keyed on a verified, channel-independent identity.

  • Memory is a pipeline, not a database. The vector store is the easy 20%. The other 80% is the engine: a single structured LLM call after the turn that extracts candidate facts, reconciles them against what’s already known, and writes them down in a strict order. Recall within a session is just the session; memory is about the next conversation.
  • Identity is the memory key, and a security boundary. Sharing memory across web and WhatsApp only works if you key on a verified, channel-independent subject (a hashed phone), never a self-asserted one. Get this wrong and “personalization” becomes “reading a stranger’s history.”
  • If you can’t measure memory, you don’t have memory. You have a liability. We built three separate evals: precision/recall/F1 on what gets extracted, MRR/nDCG on what gets recalled, and a duplicate-rate metric on the merge step. The pure scorers run in CI; the live end-to-end numbers run on demand.

After a conversation goes quiet, a debounced background job reads the new dialogue, makes one structured LLM call that folds extraction and consolidation together, and persists the result in three ordered writes: episodic memories → a structured customer profile → an idempotency watermark. The ordering is the correctness guarantee: there’s no cross-store transaction. Long-term memory lives in Postgres + pgvector; the durable profile lives in a typed SQL table with an optimistic lock; every table is tenant-isolated with Row-Level Security. On the next turn, a fail-open middleware injects a compact persona block into the system prompt. The pipeline is built on LangGraph/deepagents but the pattern is stack-agnostic. It ships gated off by default behind two independent flags.


Why “just add a vector database” falls over

Section titled “Why “just add a vector database” falls over”

Dropping a vector store next to your agent gives you retrieval. It does not give you memory, for three reasons that only show up in production.

  • Recall within a session is already solved. The conversation is in the context window; the checkpointer persists it per thread. You don’t need semantic search to remember what was said four turns ago. Memory only earns its keep at the next contact (tomorrow, or on another channel), which is exactly when the session context is gone.
  • Writing is a decision, not a dump. If you embed every message, you get a pile of duplicated, contradictory, decaying facts. “I want a beachfront resort in Bali” and, three weeks later, “actually somewhere quieter inland” both being equally retrievable is worse than useless. Someone has to decide what to keep, what a new statement overrides, and what to forget.
  • Structured facts want a schema, not a blob. “What’s this customer’s budget?” should be a SQL column your analysts can query, not a similarity search over prose. A profile and an episodic log are different surfaces with different access patterns.

So the interesting design work isn’t “which vector DB.” It’s the machinery that turns a raw conversation into durable, deduplicated, queryable knowledge, and does it safely, per tenant, without blocking the response.

The shape: an ETL job that runs after the turn

Section titled “The shape: an ETL job that runs after the turn”

Here’s the whole system on one page. The response path stays hot; everything memory-related happens off it.

sequenceDiagram participant U as User participant E as Agent Engine participant Sc as Consolidation Scheduler<br/>(debounced, per thread) participant P as Consolidation Pipeline participant M as MemoryStore (pgvector) participant Pr as customer_profile (SQL) participant W as watermark U->>E: turn E-->>U: response (streamed) E->>Sc: schedule({tenant, thread, subject}): fire-and-forget Note over Sc: coalesce turns,<br/>wait for the thread to go quiet Sc->>P: consolidate(job) P->>W: read watermark P->>P: read dialogue AFTER watermark P->>P: one structured LLM call<br/>(extract + consolidate) P->>M: 1. write episodic memory ops P->>Pr: 2. save profile (optimistic lock) P->>W: 3. advance watermark

The five moving parts, in the order data flows through them: a store (where facts live), an engine (the pipeline that writes them), a scheduler (when it runs), an identity (whose facts these are), and an injector (how they come back). Then the evals that measure all of it. Let’s take them in turn.

A running agent has four distinct persistence surfaces, and conflating any two of them causes pain later. They are separated by table and by key structure:

SurfaceTableKeyed byWhat it holds
Checkpointeragent_checkpoints(tenant_id, thread_id, …)Short-term per-thread run state
Long-term memorymemory_store(tenant_id, namespace_path, key)Cross-thread episodic memories + embeddings
Corpus RAGcorpus_chunks(tenant_id, corpus_id, chunk_id)Knowledge-base document chunks
Customer profilecustomer_profile(tenant_id, subject_key)Structured, BI-queryable per-customer facts

The MemoryStore port (src/runtime/agent/ports/memory-store.ts) is a namespaced key-value store with an optional semantic index. Its method surface is deliberately small:

put(ns: readonly string[], key: string, value: unknown, text?: string): Promise<void>;
get(ns: readonly string[], key: string): Promise<unknown>;
list(ns: readonly string[]): AsyncIterable<{ key: string; value: unknown }>;
search(ns: readonly string[], query: string, k?: number): Promise<SearchHit[]>;
delete(ns: readonly string[], key: string): Promise<void>;

Two design choices matter here. First, the namespace is a tuple whose first segment must be the tenant ID (end-user memory lives at [tenant, "user", userId, "memories"]) and adapters re-check that prefix rather than trust the caller. Second, put takes a separate optional text argument: the value you get back is kept distinct from the text that gets embedded. You store structured JSON but embed a clean, metadata-prefixed projection of it.

The pgvector adapter (src/runtime/agent/adapters/memory/memory.repository.ts) writes the embedding with raw SQL inside the same tenant-scoped transaction as the upsert, and searches with cosine distance:

SELECT "key", "value", 1 - ("embedding" <=> $1::vector) AS score
FROM "memory_store"
WHERE "tenant_id" = $2 AND "namespace_path" = $3
AND "embedding" IS NOT NULL
ORDER BY "embedding" <=> $1::vector
LIMIT $4

Embedding on write is best-effort: it returns a vector, null (to clear a now-stale vector, search-invisible beats search-wrong), or undefined (leave it untouched). A write never fails because it couldn’t be embedded. And when no embedding backend is configured for a tenant, search() throws a typed NotSupportedError that every caller is written to catch and degrade from: semantic recall becomes recency-ordered list(). That degradation is a documented contract, not a bug. (In the current build only OpenAI embeddings (text-embedding-3-large at 1024 dimensions) are actually wired; other providers are stubs. Be honest about that in your own port.)

The customer profile is the other half of the store, and it’s a plain SQL table on purpose (1787500000000-CreateCustomerProfile.ts):

"tenant_id" text NOT NULL,
"subject_key" text NOT NULL,
"preferred_language" text,
"locale" text,
"lifecycle_stage" text,
"last_interaction_at" TIMESTAMP WITH TIME ZONE,
"attributes" jsonb NOT NULL DEFAULT '{}'::jsonb,
"version" integer NOT NULL DEFAULT 0,
CONSTRAINT "PK_customer_profile" PRIMARY KEY ("tenant_id", "subject_key")

It’s a hybrid: a typed, BI-queryable core (preferred_language, locale, lifecycle_stage, last_interaction_at) plus a domain-owned jsonb attributes bag for everything else, and a version column for optimistic locking. We chose a table over a graph database deliberately: the whitepaper default for a user profile is a structured contact-card of facts, and the one thing a table loses, temporal reasoning, you recover with valid_at/invalid_at columns, not a new database.

Every one of these tables carries the same isolation discipline: a non-superuser application role, ENABLE and FORCE ROW LEVEL SECURITY, and a tenant_isolation policy predicated on current_setting('app.current_tenant', true). Tenant scoping isn’t application logic you can forget to write; it’s enforced in the database.

The engine: one structured call, three ordered writes

Section titled “The engine: one structured call, three ordered writes”

This is the part people skip, and it’s the part that makes memory real. The pipeline (LlmMemoryConsolidationPipeline, src/runtime/agent/adapters/memory/llm-consolidation.pipeline.ts) has a single entry point, consolidate(job), and it folds extraction and consolidation into one structured LLM call. Given the topics a domain cares about, the existing memories, the current profile, and the new dialogue window, the model emits both a set of episodic operations and a profile patch at once.

The output schema has two fields: episodicOps (an array of CREATE/UPDATE/DELETE operations) and a profile patch. The load-bearing detail is how the structured attributes are typed:

attributes: z
.array(
z.object({
key: z.string().describe("Stable snake_case key for a durable domain fact…"),
value: z.string().describe('The value as a concise string, e.g. "90000", "ready_to_book".'),
}),
)
.describe("Durable structured domain facts as key/value pairs… Re-emit a key to update its value."),

Why an array of {key, value} pairs and not a map? Because strict structured output can’t express an open-ended object with arbitrary keys; every property has to be declared up front. An array of pairs is the workaround for “I don’t know the keys in advance.” For the same reason, the typed profile fields are modeled as z.string().nullable(): required-but-nullable, not optional. Strict mode demands every field be present; “no signal this turn” is encoded as null, not as an absent key. That one rule (“model optional data as nullable”) shapes the entire schema.

The domain, not the kernel, decides what’s worth remembering, through two ports it declares on its contribution: MemoryTopic (“what is meaningful to remember,” with few-shot examples) and ProfileAttribute (the structured keys to prefer). These get rendered into the prompt. The kernel owns the contract; the domain owns the content. An assistant that declares no topics is a no-op: the pipeline advances the watermark and returns without an LLM call.

Concretely, these are declared as sibling arrays on the domain’s HarnessContribution: memoryTopics next to profileAttributes, alongside its prompts, tools, and skills. Here’s how a travel-planning assistant might fill them in:

memoryTopics: [
{
label: "trip_preferences",
description:
"The traveller's destination interests, budget, trip dates, and travel style (luxury, backpacking, family-friendly).",
examples: [
{
dialogue:
"user: We're thinking Portugal in early May, two adults, mid-range budget, and we love food and walking tours.",
memories: [
"The traveller is considering Portugal in early May for two adults on a mid-range budget.",
"The traveller enjoys food experiences and walking tours.",
],
},
],
},
],
profileAttributes: [
{
key: "budget_tier",
description: "The traveller's stated budget level: budget, mid_range, or luxury.",
example: "mid_range",
},
],

A MemoryTopic is small: a stable label, a natural-language description that guides the extractor, and optional few-shot examples. The load-bearing detail is what those examples extract: natural-language memory sentences, not {key, value} pairs. A topic teaches the model what an episodic memory reads like; the {key, value} array from the schema above is a separate surface, the structured profile patch, whose keys the sibling profileAttributes declares as a key, a description, and one example value apiece. And examples is optional: a precise description alone carries topics a plain sentence already pins down, and it’s the nuanced ones (where “budget” might mean per-night or whole-trip) that earn a worked example.

Drift outside the declared attribute set is logged, not rejected (observed / fail-open), so a domain can start with a handful of keys and widen the contact card as it learns what deserves a column.

Before the call, the pipeline has to decide which existing memories to show the model; you can’t dump thousands into a prompt. The selection policy (selectExistingMemories, a pure exported function) is a recency backbone with a relevance reserve:

// 1. Recency backbone: leave `relevanceQuota` slots for relevance.
for (const m of byRecency) {
if (selected.size >= max - relevanceQuota) break;
selected.set(m.key, m.content);
}
// 2. Relevance reserve: window-relevant memories the backbone missed.
for (const key of relevanceKeys) {
if (selected.size >= max) break;
const content = byKey.get(key);
if (content !== undefined && !selected.has(key)) { selected.set(key, content); usedRelevance = true; }
}
// 3. Recency backfill: top up to `max` (covers the no/thin-relevance case).
for (const m of byRecency) {
if (selected.size >= max) break;
if (!selected.has(m.key)) selected.set(m.key, m.content);
}

With MAX_EXISTING = 100 and RELEVANCE_QUOTA = 40, that’s a 60-slot recency backbone plus up to 40 relevance-selected memories, and the relevance search() only runs when you’re actually over budget. When it truncates, it logs the counts (total, kept, strategy), never the content. A silent cap that drops old-but-relevant facts is exactly the kind of bug that looks like “the model forgot”; making truncation observable is non-negotiable. (On tenants without embeddings, relevance selection degrades to pure recency, and the log says so.)

Three writes, in an order that is the correctness guarantee

Section titled “Three writes, in an order that is the correctness guarantee”

There is no distributed transaction across the vector store, the SQL profile, and the watermark. Instead, correctness comes from write order:

  1. applyEpisodicOps: write the episodic memories
  2. saveProfile: save the profile patch (optimistic-locked)
  3. watermarks.advance: record that this window is durably consolidated

The watermark (memory_consolidation_watermark, keyed (tenant_id, thread_id)) is an idempotency cursor: the pipeline only reads dialogue after processed_through, and advances it last. If an episodic write throws, the pipeline aborts before the profile save and before the watermark moves, so the window simply gets reprocessed on the next trigger. This is at-least-once with an idempotent replay, and it’s why there’s no dead-letter queue. Reprocessing converges because episodic keys are content hashes and the profile save is version-guarded.

That ordering is important enough that it’s pinned by a test, not just a comment:

'on an episodic write failure, saves no profile and leaves the watermark UNADVANCED (window replays)'

The test makes store.put reject, supplies a real profile change, and asserts that profiles.save was never called and the watermark never advanced. A future refactor can’t quietly reorder the writes past the watermark without turning that test red.

saveProfile itself is a single optimistic-locked upsert. It shallow-merges the new attributes over the existing jsonb bag in application code (re-emitted key overwrites, others preserved), then writes the whole bag under a version guard:

INSERT INTO customer_profile (…) VALUES (…, 1, NOW(), NOW())
ON CONFLICT (tenant_id, subject_key) DO UPDATE SET
…, version = customer_profile.version + 1, updated_at = NOW()
WHERE customer_profile.version = $8
RETURNING version

An empty RETURNING means the version moved under us (a concurrent consolidation won) so the save reports conflict, we leave the watermark unadvanced, and the window replays. No lost updates, no locks held across an LLM call.

One field earns special handling. lifecycle_stage (lead → comparing → active) is the field the model most often under-classifies: a customer states a purchase deadline and the next window still comes back “lead.” So the stage is monotonic by construction: ratchetLifecycle never lets it move backward:

const LIFECYCLE_RANK = { lead: 1, comparing: 2, active: 3 };
// null next → keep current; normalize case/whitespace before ranking;
// churned is orthogonal (settable/escapable any time);
// if both ranked and next < current → keep current (no regression).

Worth being precise about what this does and doesn’t fix: the ratchet mitigates the under-classification by guaranteeing the stored stage never regresses. It does not make the model classify better. That’s a policy problem, and the honest move is to measure it rather than claim it’s solved.

Scheduling: enqueue-and-forget, debounce, at-least-once

Section titled “Scheduling: enqueue-and-forget, debounce, at-least-once”

The engine must never be on the response path. The trigger, in the run-terminal handler of the agent engine, is fire-and-forget and double-gated:

if (this.consolidationScheduler && this.config.get("MEMORY_CONSOLIDATION_ENABLED", { infer: true })) {
void this.consolidationScheduler
.schedule({ tenantId, threadId, runId, subjectId, hashedUserId, assistantName, trigger: "run_close" })
.catch((err) => this.logger.warn({ err }, "memory consolidation schedule failed"));
}

It fires run_close on every turn, but the scheduler (InProcessConsolidationScheduler) coalesces per (tenant, thread) and debounces: a rapid back-and-forth collapses into one consolidation when the thread goes quiet (default 45s debounce, with a non-resettable max-wait ceiling and a max-turns valve so a long conversation can’t defer forever). Per-thread serialization means there are never two concurrent runs for the same thread. On failure it retries with bounded exponential backoff and then drops the job, safe, because the unadvanced watermark means the window is retried next session.

Crucially, the scheduler is a port (MemoryConsolidationScheduler.schedule(job)), and its contract is “enqueue and return; never throw into the caller.” The in-process adapter is the current implementation; a durable TemporalConsolidationScheduler with a real DLQ is a documented future behind the same interface. The durability substrate is swappable without touching the pipeline.

Identity is the memory key, and a security boundary

Section titled “Identity is the memory key, and a security boundary”

Here’s the question that decides whether cross-channel memory works: what do you key it on?

Key on the channel principal and you get siloed memory. A user’s web identity is their auth principal; their WhatsApp identity is a channel-address hash that doesn’t even contain the phone number. Same person, two keys, two disconnected memory buckets.

The fix is a channel-independent subject identity derived from a verified phone:

export function canonicalSubjectId(tenantId: string, e164: string): string {
return hashIdentity(tenantId, `phone:${e164}`); // sha256(`${tenantId}:phone:${e164}`)
}

Note there’s no channel prefix, so the same phone maps to the same subject on every channel. Threads still bind to the channel address (that’s what keeps conversations continuous), but memory keys on the subject. Everything downstream resolves subjectId ?? hashedUserId, in that order: the memory tools, the consolidation pipeline, the profile’s subject_key, and, as we’ll see, the injector.

And this is a security boundary, stated three times in the source because it matters:

Only derive this from a verified phone claim (BSP verification, or a trusted OTP-completed token), never a self-asserted request value. A caller who could assert an arbitrary phone could read another person’s memory.

In production (AUTH_MODE=jwt) the subject comes only from a verified phone claim in the token; an unparseable phone yields no subject rather than a wrong one. The header-based shortcut used in dev/QA is explicitly disabled in production. Personalization built on a self-asserted identity is an account-takeover path wearing a friendly name.

One small but real systems lesson lives here too. The three user-memory tools (save_user_memory, recall_user_memories, forget_user_memory) are a shared runtime capability, registered exactly once at boot by a UserMemoryToolRegistrar. Domains reference them by name via sharedToolNames rather than each building their own copies, because registering the same global tool name from two domains trips a DuplicateToolError at startup. The registry collision fails loud at boot on purpose; the fix is to register once, not to weaken the guard.

Reading it back: pre-turn injection, fail-open

Section titled “Reading it back: pre-turn injection, fail-open”

Storing memory is half the loop. The other half is a middleware that, before the model call, folds a compact persona block into the system prompt:

wrapModelCall: async (request, handler) => {
try {
const bag = readConfigurableBag(request.runtime);
const tenantId = readString(bag, "tenant_id");
const subjectKey = readString(bag, "subject_id") ?? readString(bag, "hashed_user_id");
if (!tenantId || !subjectKey) return handler(request);
const { block, cached } = await loadBlock(/* … */ () => profiles.get(tenantId, subjectKey));
if (!block) return handler(request);
const base = typeof request.systemPrompt === "string" ? request.systemPrompt : "";
const systemPrompt = base ? `${base}\n\n${block}` : block;
onApply?.({ blockChars: block.length, cached });
return handler({ ...request, systemPrompt });
} catch (err) {
onError?.(err); // Fail-open: any failure leaves the turn untouched.
return handler(request);
}
},

Everything about this is defensive. Fail-open at every branch: no subject key, no profile row, or a repository throw all fall through to the untouched turn: a personalization read must never break a run. A 30-second per-subject TTL cache (bounded, oldest-evicted) means a multi-turn run does one profile read, not one per turn. The block is bounded (a capped number of attributes, values truncated) and prefixed with a preamble telling the model to confirm rather than assume and not recite it back verbatim. And observability is counts-only: the hook logs blockChars and cached, never the block content or the subject key.

It composes in the middleware stack alongside skills (prompt-augmenting middleware, ahead of domain policy) and, like the write side, it’s independently gated (MEMORY_PROFILE_INJECTION_ENABLED, default off). Reading and writing memory are two flags, so you can turn on consolidation and watch what it stores for a while before you ever let it influence a response.

If you can’t measure memory, you don’t have memory

Section titled “If you can’t measure memory, you don’t have memory”

Every stage above makes probabilistic decisions. “It remembered something” and “it remembered the right thing, once, without duplicating it” look identical in a demo and completely different in production. So memory quality is measured by three evals, each a pure, deterministic scorer (unit-tested in CI) wrapped around a thin CLI that calls the real model/store on demand.

Extraction, eval:memory. Runs the real consolidation prompt cold (no existing memories, no writes) and scores what comes out. Memories are matched greedily (each gold fact claims the first unclaimed extracted fact it matches) into precision/recall/F1:

for (const gold of expected) {
const idx = extracted.findIndex((cand, i) => !claimed[i] && matches(cand, gold));
if (idx >= 0) { claimed[idx] = true; tp += 1; }
else { missed.push(gold); }
}
const fp = extracted.filter((_, i) => !claimed[i]).length; // spurious

The default matcher is token-set Jaccard; profile values use a digit-tolerant compare so "RM90,000" and "90000" count as equal. A --min-f1 flag turns it into a gate. On our (small) seed set the first run scored memory P/R/F1 of 100/88/93%, and, tellingly, lifecycle_stage at 1/2. The harness quantifies the exact residual the ratchet works around, instead of letting us pretend it’s fixed.

Recall, eval:recall. Seeds an isolated namespace, runs the real semantic search(), and scores the ranking with hit-rate, precision, recall, MRR and nDCG (macro-averaged across queries), cleaning up even if a seed throws:

topK.forEach((key, i) => {
if (!rel.has(key)) return;
hits += 1;
if (firstRelRank === 0) firstRelRank = i + 1;
dcg += 1 / Math.log2(i + 2);
});
// mrr = mean(1 / firstRelRank); ndcg = dcg / idcg

A --min-mrr flag gates it. Our first run scored MRR 1.00 / nDCG 100%, on a deliberately easy seed corpus. A perfect score there means the plumbing works, not that retrieval is solved; the honest next step is adversarial corpora with near-duplicates and distractors.

Merge, eval:merge. The one that measures the residual the pipeline is most exposed to: duplicate creates. It runs the prompt warm (seeded with existing memories and a current profile) and asks whether the model correctly reconciled: an evolved fact should be an UPDATE, a restated fact should be a no-op, and only a genuinely new fact should be a CREATE. The headline metric penalizes duplicates hardest:

const dupPenalty = createCount === 0 ? 0 : duplicateCount / createCount;
const reconciliation = (newFacts.f1 + updateRecall + deleteRecall) / 3;
const base = mean([reconciliation, opEconomy, ...(lifecycle == null ? [] : [lifecycle])]);
return Math.max(0, base * (1 - dupPenalty));

Because the whole risk is paraphrased duplicates, the default Jaccard matcher under-reports them, so the gate is meant to run with an embedding matcher (cosine ≥ 0.82), and the CLI warns you when it isn’t. --max-duplicate-rate and --min-value are the gates.

One honesty note that belongs in any writeup like this: these three evals are gate-ready but not yet wired into CI. The pure scorers run on every PR through unit tests; the live end-to-end numbers need the stack up and a vaulted key, so they run on demand. The exit-code gates exist; the workflow that calls them is the next step, not a claim I’ll make prematurely.

No architecture is free, and the honest failure modes here are worth as much as the design.

  • Replay can duplicate. applyEpisodicOps is a sequential loop, not an atomic batch. If op #1 commits and op #2 throws, the window replays, and because replay re-invokes a non-deterministic model, a CREATE whose content the model paraphrases the second time hashes to a different key and lands as a duplicate. UPDATE/DELETE converge (they target existing keys); CREATE converges only on byte-identical content. It’s mitigated by feeding the survivor back with a “prefer UPDATE, never duplicate” instruction, which rests on model compliance, which is exactly why eval:merge exists.
  • The real fix is deferred on purpose. An atomic all-ops batch write in one transaction would close the duplicate window, but it needs a new batch/transaction capability on the MemoryStore port. The sequenced design is honest about being at-least-once, and the durable version is queued behind the Temporal graduation: build the merge eval first, size the real duplicate rate, then pay for the fix.
  • No DLQ. A persistently failing job is dropped, not parked. That’s acceptable only because the unadvanced watermark retries the window next session. It’s a deliberate trade of a moving part for a convergent replay, and it’s the first thing that changes when this graduates to a durable scheduler.
  • Structured-attribute routing is probabilistic. Extraction is reliable; whether a fact lands in the typed profile bag versus episodic memory is model-dependent. Again: measured, not assumed.

Q: Why not just let the agent call a save_memory tool mid-conversation?

That tool exists, and it’s good for explicit “remember this.” But relying on it for everything means memory only gets written when the model remembers to write it, mid-task, competing for attention with the user’s actual request. The after-the-turn pipeline makes durability the default instead of an action the model has to choose.

Q: Graph database or table?

Table, for a user profile. The advanced multi-hop case is where graphs earn their complexity; a contact-card of continuously-updated facts isn’t it. The one thing a table gives up, temporal reasoning, comes back with valid_at/invalid_at columns. Keeping the store behind a port means a graph adapter could slot in later if a use case ever demands it.

Q: Every turn triggers consolidation, isn’t that expensive?

The enqueue is nearly free; the scheduler coalesces and debounces so the expensive part, the LLM call, happens once per quiet thread, not once per turn. Recall inside a session is served by the session, so the cadence only governs durability for the next contact.

Q: What happens to memory across two channels for the same person?

They share one bucket, keyed on the verified-phone subject, but only once the identity is verified (BSP-verified on WhatsApp, an OTP/JWT phone claim on web). Anonymous sessions write no user memory. Sharing memory is a feature; sharing it across an unverified identity is a vulnerability.

  • Wiring the evals into CI as real gates. The exit codes exist; a workflow that stands the stack up with a vaulted key and enforces --min-f1 / --max-duplicate-rate on a schedule is the next step.
  • Adversarial recall corpora. MRR 1.00 on an easy seed set proves plumbing, not quality. Near-duplicate and distractor corpora are where the retrieval numbers start to mean something.
  • Atomic batch writes, then a durable scheduler. Size the real duplicate rate with eval:merge, then decide whether an atomic MemoryStore batch write (and a Temporal-backed scheduler with a DLQ) is worth the new capability.
  • A first-class tier for direct user preferences. “Always do X” from the user is a different kind of memory than an inferred fact, and probably deserves its own top-priority tier separate from the consolidated bag.