Internals / Part 2 of 2 · 18 min read
The Log Is the Truth: Building Agent Memory From Scratch
Part 1 named the five gaps between persistence and real memory. Part 2 closes them: the one architectural decision that makes it work, how to measure whether retrieval is any good, and a step-by-step path to build a small memory layer you can start this week.
Part 1 ended on a claim: coding agents already persist and retrieve knowledge, but the pieces do not add up to a governed lifecycle. Capture is best-effort. Retrieval returns what is related, not what is current. The same fact lives in five places with no authority between them. A recorded decision is not applied where it matters. And nothing scopes who or what can see a memory.
This piece builds a layer that closes those five gaps. It is one version, a case study from a system I built to give a fleet of agents a shared memory. Treat the shapes as a reference, not a spec: the point is the reasoning, and the first rung of it is a single file you can start today.
A warning before the architecture. It is tempting to reach for the heavy machinery first, a graph database, a fine-tuned retriever, a vector store with its own server. Almost none of that earns its place at the start. The load-bearing decisions here are cheap, and most of the credibility of the result comes from measuring it, not from the parts list.
The one decision that makes it work
Start with the question every memory system has to answer: what is the source of truth?
Three properties are in tension, and a real memory needs all three at once. The queryable store has to be replaceable, because you will change your embeddings, your schema, your ranking, and you cannot treat a rebuild as data loss. Truth has to have durable history, because "this fact superseded that one, on this date" is the whole game, and editing a row in place destroys it. And agents have to be able to write, because a memory that only you can update is a document, not a memory.
Files as the source of truth give you replaceable, but not history or safe concurrent writes. A mutable database as the source of truth gives you writes, but the moment you UPDATE a row the old truth is gone. The design that satisfies all three is older than any of this: an append-only log is the truth, and the database is a projection built from it.
Every change is an event, appended, never edited:
{"ts":"2026-09-01T18:04:11Z","op":"create","key":"deploy|branch","value":"main","actor":"agent:codex","confidence":0.8,"source":"turn-4471"}
{"ts":"2026-09-08T09:12:33Z","op":"supersede","key":"deploy|branch","value":"release","prev":"...","actor":"human","confidence":1.0}
The database, the thing you actually query, is a materialized view of that log: the latest event per fact, plus whatever indexes make lookup fast. It holds no truth of its own. Delete it, replay the log from the beginning, and you get the same state back.
That single choice buys three things that are brutal to bolt on later. History is free: you can ask what the system believed on any past date. A rebuild path is free: change your schema or your embedding model, drop the derived store, replay. And an audit trail is free: every current value traces back to the event that set it, the actor, and the turn it came from. The rule worth taping above your desk: if you find yourself writing to the state table without first writing the event, stop.
One consequence is worth dwelling on. Once a night, rebuild the entire index from the log alone and diff it against the live one. If they match, the system is provably consistent with its own history. If they do not, something wrote to derived state without going through the log, and the nightly job finds it before a person does. The system proves, every night, that it can be destroyed and reconstructed from its own diary.
What a memory actually is
The unit is not a document or a chunk. It is a single, typed statement with one owner. The fields that matter:
memory
id stable across corrections
type fact | preference | decision | lesson | norm | profile | ...
fact_key canonical "subject|attribute" ← the address
statement the text, one idea, short
status active | superseded | expired | proposed
confidence 1.0 stated by a human · 0.8 asserted by a trusted agent · 0.6 inferred
privacy green (shareable) | amber (local only) | red (refused at the door)
observed_at when it became true
recorded_at when it was written
last_confirmed_at
superseded_by / version
Two of those fields do most of the work.
The fact_key is the address, and it is not the wording. "Ships from the release branch" and "the staging branch is called release" are the same fact and belong under the same key, deploy|branch. Two sentences that merely sound alike do not. Computing a canonical subject|attribute key, rather than trusting the sentence, is what makes "one current value per fact" enforceable instead of aspirational.
And you enforce it in the database, not in a script you hope runs. A partial unique index does it in one line:
CREATE UNIQUE INDEX one_active_per_key
ON memories (tenant, fact_key)
WHERE status = 'active';
Now the store physically refuses to hold two live values for the same fact. A conflicting write does not silently win; it raises, and the writer has to decide, explicitly, whether it is a correction (supersede the old, bump the version) or a genuinely different fact (a different key). Similar wording never merges two facts, and one fact never quietly grows two current answers. That is gap 3, source of truth, made structural.
Privacy sits in the same layer, as a hard pre-filter rather than a redaction step. Every memory carries a tier, and the tier is checked before retrieval runs, not after. Anything sensitive, financial or medical or personal, is refused at the door or kept strictly local; it never becomes a row a later bug can surface. Filtering after retrieval means the fact already left the vault. Filtering before means it never got that far. That is gap 5.
Retrieval: two searches, and the model was not what broke
Getting the right memory back is where most systems quietly fail, and where I learned the most surprising lesson.
Start with the honest baseline. Retrieval fuses two searches, because they fail in opposite ways. Keyword search (BM25, which SQLite gives you for free through FTS5) nails exact identifiers and rare strings but misses paraphrase. Vector search, nearest-neighbor over embeddings, catches meaning but misses exact tokens. You run both and combine the rankings. Neither alone is enough.
Then the surprise. I ran a spike to pick an embedding model, four local ones, against my own corpus. Every one of them failed the same way, on the project's own proper nouns. Ask "which model should the assistant use," and the correct memory scored a cosine similarity of 0.43. An unrelated sentence scored 0.44 and won. That is not a rounding error. That is the retrieval layer handing back the wrong fact, with a straight face, on exactly the terms the domain cares about most.
Bigger models did not fix it. What fixed it was a lookup table. I wrote a small alias registry, a handful of hand-written mappings from the informal name someone would actually type to the canonical term the memory was filed under, and expanded the query through it before embedding. The same correct memory jumped to about 0.72 on the first pass, and to about 0.82 once the aliases covered the right terms. Model size did not fix vocabulary. A cheap dictionary did. That is the single most useful thing I can hand you from this project: before you reach for a larger retriever, check whether your problem is capacity or vocabulary. More often than you would expect, it is vocabulary.
The rest of the stack is deliberately small, and it stays small because I measured it before believing it. Embeddings run in-process. The whole corpus, on the order of 600 active memories, lives in a single SQLite file of roughly 40 MB. Vector search is brute-force cosine over the lot, no approximate index, and it runs in about 20 milliseconds over 10,000 rows. No server, no daemon, nothing listening on a port. I rejected the heavier options on purpose: a hosted vector database (operations and a network hop for a personal-scale corpus), a model-serving daemon (a background dependency where an in-process library will do), and a graph database (a one-hop join table covers every relationship I actually have; promote to a graph only when a query proves it needs one). None of those were wrong in general. They were wrong for the size of the problem, and matching the machinery to the size of the problem is most of the craft.
How you know it works: measuring recall and precision
This is the part that separates a memory system you can trust from a pile of text that usually helps. If you build one thing from this piece, build the evaluation.
Two numbers carry most of the weight, and they pull against each other.
Recall@k asks: of the memories that should have come back for a query, how many appeared in the top k? It punishes forgetting.
Precision@1 asks: when the system returns a top answer, how often is it the right one? It punishes confident wrongness.
A third, MRR (mean reciprocal rank), rewards putting the right answer near the top of the list rather than merely somewhere in it.
You want both recall and precision, and you cannot have both for free, which is exactly the tension from Part 1's precision-over-recall example. A change that returns more candidates lifts recall and usually costs precision. So you measure both, on every change, or you are flying blind.
Measuring needs a golden set: a fixed list of real queries paired with the memories that are the correct answers. Build it from your own corpus, not a benchmark; a few dozen queries is enough to start and catches regressions immediately. Against mine, recall@5 has sat between 0.86 and 0.87 across recent runs, MRR around 0.75, and precision@1 around 0.67. Those are not state-of-the-art-benchmark numbers and they are not trying to be. They are stable numbers on my data that move when I break something, which is the property that matters most.
Averages hide the failure that actually hurts, so test for it directly with an adversarial matrix, a handful of cases the system must never get wrong:
- A fact was superseded last week. Query it. The current value must come back, never the stale one. In my runs this "stale-as-current" failure happens zero times, because supersession is a database constraint, not a ranking preference.
- Ask for something the corpus does not contain. The system must abstain, not return the nearest weak match.
- Ask as a consumer who is not cleared for a memory. It must not appear.
- Point the loader at a corrupt store. It must fail closed, not serve garbage.
Then two cheap, high-leverage checks. The nightly rebuild-and-diff from the section above is a determinism test: replay the log, diff against live, any mismatch is a bug. And an LLM-as-judge grades a sample of real queries against what was returned, every day, on live traffic. That last one is the real advantage of dogfooding: your own daily use generates more honest evaluation signal in a week than a synthetic suite does in a month, and it costs you nothing but the calls.
Every returned memory also carries its receipt: the source, the status, when it was recorded, when it was last confirmed. Without that, you cannot debug a wrong answer or decide how far to trust a right one.
Capture: let the worker write the memory
Retrieval is half the loop. The other half is getting good memories in without a human transcribing them, and this is where the naive instinct, a regex or a keyword trigger, produces noise.
The highest-precision judge of "did this turn produce something worth keeping" is the model that just did the turn. It has the full context; a separate pass re-reading the transcript later does not. So the capture loop is two halves, and it rides on hooks the agent harness already gives you. Claude Code exposes exactly the two you need: a SessionStart hook that fires when a session begins, and a Stop hook that fires at the end of every turn.
- Recall first. At the start of a session, inject a small brief of the durable memories relevant to the task. Keep it tight, a kilobyte or two, not a data dump; the point is to seed the desk, not bury it.
- Remember on learning. At the end of a turn, ask the model one question:
Did this turn produce a durable learning, decision, preference, or correction
that should outlive this session? If so, write at most five memories, each a single
short statement with a subject and attribute. If not, write nothing.
The cap matters as much as the question. An extractor with no ceiling turns every session into ten mediocre memories instead of one good one. And "write nothing" has to be an honored answer, or the store fills with the model performing helpfulness.
A trusted writer's correction supersedes in place. An untrusted writer's conflicting claim lands as proposed, linked to what it contradicts, and never flips a trusted value on its own. Near-duplicates with a different key get linked and flagged for review, never auto-merged; similar wording is not sameness. That is the whole learning loop, and it is gap 1, capture, turned from best-effort into something with rules.
Saying "I don't know"
One behavior does more for trust than any ranking improvement: when nothing clears the bar, the system returns insufficient_evidence, with the reason, instead of the nearest weak match dressed up as an answer.
This is the hardest instinct to build, because every retriever can always return something, and a padded answer looks like a helpful one right up until it is confidently wrong. Set an explicit floor, and return honest emptiness below it. An agent that says "I do not have that" is worth more than one that guesses, because you can trust the times it does answer. Disuse follows the same principle: a project fact nobody has confirmed in N days flips to expired, still queryable for history, no longer served as current. The system stops trusting stale facts on its own, and nobody has to notice.
Build your own, step by step
You do not need any of the above on day one. Here is the ladder I would climb again, each rung mapped to the gap from Part 1 it closes.
Day 1: one file. Create a single append-only log your agent reads at the start of every session and appends to at the end. One JSON line per fact: a timestamp, a subject|attribute key, a value, who wrote it. Nothing else. Reading it first closes a little of gap 2 (retrieval); appending to it closes a little of gap 1 (capture). This alone will change how it feels to work with the agent, and everything below is optimization on top of this habit.
Week 1: structure and the constraint. Move the log into events and project a memories table from it. Add the fact_key and the partial unique index so one fact has one current value (gap 3). Wire two hooks: a session-start recall brief, and an end-of-turn capture with the prompt above and a hard cap (gap 1, properly this time). Add the privacy tier as a column and filter on it before you return anything (gap 5).
Month 1: retrieval and self-checking. Add hybrid search: FTS5 for keywords, a brute-force vector scan for meaning, combined. Add the alias table the first time a query returns a confidently wrong neighbor, because it will (gap 2, properly). Add the nightly rebuild-and-diff and TTL expiry so staleness cannot accumulate silently. Build the golden set and the adversarial matrix now, not later; they are what let you change everything above without fear. Enforcement of real decisions, gap 4, is the last and hardest: retrieve the relevant rule at the decision point, and for an invariant that must hold, back it with a check that can fail, a test or a hook, because memory alone cannot stop an agent from acting against a fact it can see.
Notice that the whole thing degrades gracefully. Stop at day 1 and you have a real, if modest, memory. Stop at week 1 and you have governance. The month-1 work is what makes it fast and trustworthy at scale, and by then you will have measurements telling you exactly which rung to build next.
What you actually get
The model at the center of all this still forgets you between sessions. That was never the thing to fix. What you build around it, the log that holds the truth, the key that keeps one value per fact, the retrieval you can measure, the capture that runs on its own, and the honesty to say "I don't know," is the part that remembers.
None of it required a smarter model. It required treating memory as an engineering problem with a lifecycle: capture it reliably, retrieve the right thing, know which version is authoritative, apply the decisions that matter, and control what can be seen. Start with one file this week. Add a constraint when contradictions bite, an alias when retrieval lies, and a measurement before you trust any of it. Everything past that is a rung you climb when the numbers tell you to.
Missed Part 1?
Why Coding Agents Forget What You Already Taught Them
The diagnosis this piece builds on: how coding agents actually assemble what the model knows, why that falls short as a project grows, and the five gaps a real memory layer has to close.
Sources
- Event sourcing as a pattern (append-only log of events, state as a projection): Martin Fowler, "Event Sourcing"
- SQLite full-text search (BM25 ranking): SQLite FTS5
- Vector search in SQLite: sqlite-vec
- In-process embeddings: fastembed
- Retrieval evaluation, why hybrid beats dense-only out of domain: Thakur et al., "BEIR" (2021)
- Long-context "lost in the middle" (from Part 1): Liu et al., "Lost in the Middle" (2023)
- Claude Code hooks (SessionStart / Stop), used for recall and capture: code.claude.com/docs/en/hooks
Get Internals in your inbox
Deep technical teardowns and build guides.
One rigorous piece at a time.
Subscribe