Exploded View

Internals / Part 2 of 2 · 18 min read

The Log Is the Truth: Building Agent Memory From Scratch

Part 1 named the five gaps between persistence and real memory. Part 2 closes them: the one architectural decision that makes it work, how to measure whether retrieval is any good, and a step-by-step path to build a small memory layer you can start this week.

JOURNAL · APPEND-ONLY · THE TRUTH 2026-09-01 create deploy|branch = main 2026-09-02 create search|policy = abstain 2026-09-04 confirm deploy|branch 2026-09-08 supersede deploy|branch = release 2026-09-09 correct search|policy (v2) 2026-09-15 expire ci|owner (unconfirmed 14d) … never edited · every change has an actor and a reason project() latest event per fact_key INDEX · SQLITE · A PROJECTION memories (WHERE status = 'active' is UNIQUE per fact_key) deploy|branch release v2 1.0 green search|policy abstain v2 0.8 green ci|owner expired (queryable, not served) FTS5 · BM25 exact identifiers vectors · cosine paraphrase fuse · abstain if weak holds no truth of its own · delete it and replay NIGHTLY: rebuild index from journal, diff vs live · must match FIG. 01 The database is a pure function of the journal.
Fig. 01 The append-only journal is the truth; the queryable index is a projection of it, rebuilt from scratch every night and diffed against the live copy.

Part 1 ended on a claim: coding agents already persist and retrieve knowledge, but the pieces do not add up to a governed lifecycle. Capture is best-effort. Retrieval returns what is related, not what is current. The same fact lives in five places with no authority between them. A recorded decision is not applied where it matters. And nothing scopes who or what can see a memory.

This piece builds a layer that closes those five gaps. It is one version, a case study from a system I built to give a fleet of agents a shared memory. Treat the shapes as a reference, not a spec: the point is the reasoning, and the first rung of it is a single file you can start today.

A warning before the architecture. It is tempting to reach for the heavy machinery first, a graph database, a fine-tuned retriever, a vector store with its own server. Almost none of that earns its place at the start. The load-bearing decisions here are cheap, and most of the credibility of the result comes from measuring it, not from the parts list.

The one decision that makes it work

Start with the question every memory system has to answer: what is the source of truth?

Three properties are in tension, and a real memory needs all three at once. The queryable store has to be replaceable, because you will change your embeddings, your schema, your ranking, and you cannot treat a rebuild as data loss. Truth has to have durable history, because "this fact superseded that one, on this date" is the whole game, and editing a row in place destroys it. And agents have to be able to write, because a memory that only you can update is a document, not a memory.

Files as the source of truth give you replaceable, but not history or safe concurrent writes. A mutable database as the source of truth gives you writes, but the moment you UPDATE a row the old truth is gone. The design that satisfies all three is older than any of this: an append-only log is the truth, and the database is a projection built from it.

Every change is an event, appended, never edited:

{"ts":"2026-09-01T18:04:11Z","op":"create","key":"deploy|branch","value":"main","actor":"agent:codex","confidence":0.8,"source":"turn-4471"}
{"ts":"2026-09-08T09:12:33Z","op":"supersede","key":"deploy|branch","value":"release","prev":"...","actor":"human","confidence":1.0}

The database, the thing you actually query, is a materialized view of that log: the latest event per fact, plus whatever indexes make lookup fast. It holds no truth of its own. Delete it, replay the log from the beginning, and you get the same state back.

That single choice buys three things that are brutal to bolt on later. History is free: you can ask what the system believed on any past date. A rebuild path is free: change your schema or your embedding model, drop the derived store, replay. And an audit trail is free: every current value traces back to the event that set it, the actor, and the turn it came from. The rule worth taping above your desk: if you find yourself writing to the state table without first writing the event, stop.

One consequence is worth dwelling on. Once a night, rebuild the entire index from the log alone and diff it against the live one. If they match, the system is provably consistent with its own history. If they do not, something wrote to derived state without going through the log, and the nightly job finds it before a person does. The system proves, every night, that it can be destroyed and reconstructed from its own diary.

What a memory actually is

The unit is not a document or a chunk. It is a single, typed statement with one owner. The fields that matter:

memory
  id            stable across corrections
  type          fact | preference | decision | lesson | norm | profile | ...
  fact_key      canonical "subject|attribute"  ← the address
  statement     the text, one idea, short
  status        active | superseded | expired | proposed
  confidence    1.0 stated by a human · 0.8 asserted by a trusted agent · 0.6 inferred
  privacy       green (shareable) | amber (local only) | red (refused at the door)
  observed_at   when it became true
  recorded_at   when it was written
  last_confirmed_at
  superseded_by / version

Two of those fields do most of the work.

The fact_key is the address, and it is not the wording. "Ships from the release branch" and "the staging branch is called release" are the same fact and belong under the same key, deploy|branch. Two sentences that merely sound alike do not. Computing a canonical subject|attribute key, rather than trusting the sentence, is what makes "one current value per fact" enforceable instead of aspirational.

And you enforce it in the database, not in a script you hope runs. A partial unique index does it in one line:

CREATE UNIQUE INDEX one_active_per_key
  ON memories (tenant, fact_key)
  WHERE status = 'active';

Now the store physically refuses to hold two live values for the same fact. A conflicting write does not silently win; it raises, and the writer has to decide, explicitly, whether it is a correction (supersede the old, bump the version) or a genuinely different fact (a different key). Similar wording never merges two facts, and one fact never quietly grows two current answers. That is gap 3, source of truth, made structural.

Privacy sits in the same layer, as a hard pre-filter rather than a redaction step. Every memory carries a tier, and the tier is checked before retrieval runs, not after. Anything sensitive, financial or medical or personal, is refused at the door or kept strictly local; it never becomes a row a later bug can surface. Filtering after retrieval means the fact already left the vault. Filtering before means it never got that far. That is gap 5.

Retrieval: two searches, and the model was not what broke

Getting the right memory back is where most systems quietly fail, and where I learned the most surprising lesson.

Start with the honest baseline. Retrieval fuses two searches, because they fail in opposite ways. Keyword search (BM25, which SQLite gives you for free through FTS5) nails exact identifiers and rare strings but misses paraphrase. Vector search, nearest-neighbor over embeddings, catches meaning but misses exact tokens. You run both and combine the rankings. Neither alone is enough.

Then the surprise. I ran a spike to pick an embedding model, four local ones, against my own corpus. Every one of them failed the same way, on the project's own proper nouns. Ask "which model should the assistant use," and the correct memory scored a cosine similarity of 0.43. An unrelated sentence scored 0.44 and won. That is not a rounding error. That is the retrieval layer handing back the wrong fact, with a straight face, on exactly the terms the domain cares about most.

Bigger models did not fix it. What fixed it was a lookup table. I wrote a small alias registry, a handful of hand-written mappings from the informal name someone would actually type to the canonical term the memory was filed under, and expanded the query through it before embedding. The same correct memory jumped to about 0.72 on the first pass, and to about 0.82 once the aliases covered the right terms. Model size did not fix vocabulary. A cheap dictionary did. That is the single most useful thing I can hand you from this project: before you reach for a larger retriever, check whether your problem is capacity or vocabulary. More often than you would expect, it is vocabulary.

The rest of the stack is deliberately small, and it stays small because I measured it before believing it. Embeddings run in-process. The whole corpus, on the order of 600 active memories, lives in a single SQLite file of roughly 40 MB. Vector search is brute-force cosine over the lot, no approximate index, and it runs in about 20 milliseconds over 10,000 rows. No server, no daemon, nothing listening on a port. I rejected the heavier options on purpose: a hosted vector database (operations and a network hop for a personal-scale corpus), a model-serving daemon (a background dependency where an in-process library will do), and a graph database (a one-hop join table covers every relationship I actually have; promote to a graph only when a query proves it needs one). None of those were wrong in general. They were wrong for the size of the problem, and matching the machinery to the size of the problem is most of the craft.

How you know it works: measuring recall and precision

This is the part that separates a memory system you can trust from a pile of text that usually helps. If you build one thing from this piece, build the evaluation.

Two numbers carry most of the weight, and they pull against each other.

Recall@k asks: of the memories that should have come back for a query, how many appeared in the top k? It punishes forgetting.

Precision@1 asks: when the system returns a top answer, how often is it the right one? It punishes confident wrongness.

A third, MRR (mean reciprocal rank), rewards putting the right answer near the top of the list rather than merely somewhere in it.

You want both recall and precision, and you cannot have both for free, which is exactly the tension from Part 1's precision-over-recall example. A change that returns more candidates lifts recall and usually costs precision. So you measure both, on every change, or you are flying blind.

Measuring needs a golden set: a fixed list of real queries paired with the memories that are the correct answers. Build it from your own corpus, not a benchmark; a few dozen queries is enough to start and catches regressions immediately. Against mine, recall@5 has sat between 0.86 and 0.87 across recent runs, MRR around 0.75, and precision@1 around 0.67. Those are not state-of-the-art-benchmark numbers and they are not trying to be. They are stable numbers on my data that move when I break something, which is the property that matters most.

Averages hide the failure that actually hurts, so test for it directly with an adversarial matrix, a handful of cases the system must never get wrong:

Then two cheap, high-leverage checks. The nightly rebuild-and-diff from the section above is a determinism test: replay the log, diff against live, any mismatch is a bug. And an LLM-as-judge grades a sample of real queries against what was returned, every day, on live traffic. That last one is the real advantage of dogfooding: your own daily use generates more honest evaluation signal in a week than a synthetic suite does in a month, and it costs you nothing but the calls.

Every returned memory also carries its receipt: the source, the status, when it was recorded, when it was last confirmed. Without that, you cannot debug a wrong answer or decide how far to trust a right one.

Capture: let the worker write the memory

Retrieval is half the loop. The other half is getting good memories in without a human transcribing them, and this is where the naive instinct, a regex or a keyword trigger, produces noise.

The highest-precision judge of "did this turn produce something worth keeping" is the model that just did the turn. It has the full context; a separate pass re-reading the transcript later does not. So the capture loop is two halves, and it rides on hooks the agent harness already gives you. Claude Code exposes exactly the two you need: a SessionStart hook that fires when a session begins, and a Stop hook that fires at the end of every turn.

Did this turn produce a durable learning, decision, preference, or correction
that should outlive this session? If so, write at most five memories, each a single
short statement with a subject and attribute. If not, write nothing.

The cap matters as much as the question. An extractor with no ceiling turns every session into ten mediocre memories instead of one good one. And "write nothing" has to be an honored answer, or the store fills with the model performing helpfulness.

A trusted writer's correction supersedes in place. An untrusted writer's conflicting claim lands as proposed, linked to what it contradicts, and never flips a trusted value on its own. Near-duplicates with a different key get linked and flagged for review, never auto-merged; similar wording is not sameness. That is the whole learning loop, and it is gap 1, capture, turned from best-effort into something with rules.

Saying "I don't know"

One behavior does more for trust than any ranking improvement: when nothing clears the bar, the system returns insufficient_evidence, with the reason, instead of the nearest weak match dressed up as an answer.

This is the hardest instinct to build, because every retriever can always return something, and a padded answer looks like a helpful one right up until it is confidently wrong. Set an explicit floor, and return honest emptiness below it. An agent that says "I do not have that" is worth more than one that guesses, because you can trust the times it does answer. Disuse follows the same principle: a project fact nobody has confirmed in N days flips to expired, still queryable for history, no longer served as current. The system stops trusting stale facts on its own, and nobody has to notice.

Build your own, step by step

You do not need any of the above on day one. Here is the ladder I would climb again, each rung mapped to the gap from Part 1 it closes.

Day 1: one file. Create a single append-only log your agent reads at the start of every session and appends to at the end. One JSON line per fact: a timestamp, a subject|attribute key, a value, who wrote it. Nothing else. Reading it first closes a little of gap 2 (retrieval); appending to it closes a little of gap 1 (capture). This alone will change how it feels to work with the agent, and everything below is optimization on top of this habit.

Week 1: structure and the constraint. Move the log into events and project a memories table from it. Add the fact_key and the partial unique index so one fact has one current value (gap 3). Wire two hooks: a session-start recall brief, and an end-of-turn capture with the prompt above and a hard cap (gap 1, properly this time). Add the privacy tier as a column and filter on it before you return anything (gap 5).

Month 1: retrieval and self-checking. Add hybrid search: FTS5 for keywords, a brute-force vector scan for meaning, combined. Add the alias table the first time a query returns a confidently wrong neighbor, because it will (gap 2, properly). Add the nightly rebuild-and-diff and TTL expiry so staleness cannot accumulate silently. Build the golden set and the adversarial matrix now, not later; they are what let you change everything above without fear. Enforcement of real decisions, gap 4, is the last and hardest: retrieve the relevant rule at the decision point, and for an invariant that must hold, back it with a check that can fail, a test or a hook, because memory alone cannot stop an agent from acting against a fact it can see.

Notice that the whole thing degrades gracefully. Stop at day 1 and you have a real, if modest, memory. Stop at week 1 and you have governance. The month-1 work is what makes it fast and trustworthy at scale, and by then you will have measurements telling you exactly which rung to build next.

What you actually get

The model at the center of all this still forgets you between sessions. That was never the thing to fix. What you build around it, the log that holds the truth, the key that keeps one value per fact, the retrieval you can measure, the capture that runs on its own, and the honesty to say "I don't know," is the part that remembers.

None of it required a smarter model. It required treating memory as an engineering problem with a lifecycle: capture it reliably, retrieve the right thing, know which version is authoritative, apply the decisions that matter, and control what can be seen. Start with one file this week. Add a constraint when contradictions bite, an alias when retrieval lies, and a measurement before you trust any of it. Everything past that is a rung you climb when the numbers tell you to.

Missed Part 1?

Why Coding Agents Forget What You Already Taught Them

The diagnosis this piece builds on: how coding agents actually assemble what the model knows, why that falls short as a project grows, and the five gaps a real memory layer has to close.

Sources

Get Internals in your inbox

Deep technical teardowns and build guides.

One rigorous piece at a time.

Subscribe