Exploded View

Internals / Part 1 of 2 · 12 min read

Why Coding Agents Forget What You Already Taught Them

Your coding agent forgets the decision you made last week and repeats the mistake you already corrected. How these tools actually assemble what the model knows, why it falls short as a project grows, and the five gaps a real memory layer has to close.

An engineering schematic showing how a coding agent assembles the context for a single inference call from several stores.
Fig. 01 How a coding agent assembles the context for a single call, from several stores, by several rules at once.

You are three weeks into a build that spans four repositories. On day one the coding agent felt like a senior engineer who had somehow read all of it. This morning it proposed, for the third time, the approach your team ruled out in week one. Yesterday it "fixed" a module by putting back a bug the two of you had already removed. Every few days you retype the same context, because the agent turns up having forgotten the project, the decisions, and the corrections you assumed were settled.

The longer an agent works on a project, the more its usefulness depends on something that is not the model at all: how the surrounding tooling preserves, selects, updates, and applies what the project has learned. Add more agents and it gets harder, not easier. Run several in parallel without an explicit shared-state design and each one may see a different slice of the project, learn something locally, and leave the others unaware. You become the integration layer, carrying in your head what the system was supposed to carry for you.

That is the memory problem, and it is where a lot of real work on large language models now lives. This piece, the first of two, is about what "memory" actually is inside the tools you already run, and the specific places it falls short as the work grows. Part two takes a real system apart and shows how to build the layer that closes the gap.

What "memory" is inside your coding agent

Start with the model, because it holds less than it appears to.

Using a model does not normally write anything back into its weights. The parameters it learned in training are what answer you, and your conversation does not change them: the model cannot learn that your team renamed a service last week by being told once. For each inference, the model can condition only on what the tooling places into that one call's context. So persistence, the part that has to remember your project across calls and across days, lives entirely outside the model: in files, in conversation history, in summaries, in memory stores, in whatever the harness can read and pass forward.

Which raises the real question. When you ask a coding agent something, where does the context for that call actually come from? Not from one place, and not by one rule. The harness assembles it from several stores using several different policies at once:

Two tools make this concrete. Both start the same way: they gather the instruction files you maintain by hand, CLAUDE.md in Claude Code, AGENTS.md in Codex, layered by scope from a global file down through the project to the working directory, and feed them in as context. From there they diverge. Claude Code adds an auto-memory it writes for itself: a bounded index, its first 200 lines or 25KB, loaded at startup, with topic files it reads on demand and updates when you correct it, plus path-scoped rules that load only when a relevant file is touched. Codex concatenates its instruction files under a size budget, holds a long thread together with stateless requests, prompt caching, and auto-compaction, and adds an opt-in Memories feature, off by default, that summarizes past sessions into local notes it reads back in later ones.

Both tools, then, already do two of the things you would want from memory: they persist state, and they retrieve some of it by relevance. That is real, and if you build on these tools you should use it well. What neither one gives you is a single governed lifecycle over that state. The primitives exist, persistence, retrieval, scoping, even enforcement through hooks, but nothing connects them into a guarantee that the right thing was captured, one authoritative value when sources disagree, and a past decision applied at the moment it matters. That gap is the rest of this article.

Why it still is not enough: five gaps

The failures are not random. They fall along the path a fact has to travel, from getting learned, to getting retrieved, to being trusted, to being enforced, to being scoped. Five gaps. I am naming them as a framework on purpose, because in part two each one maps to a component you can build.

Keep one running example in mind: an engineer working across four repositories, with an agent that is supposed to remember how the whole system fits together.

1. Capture: does it actually learn, and keep it?

Auto-memory is a real start, but capture is model-decided and best-effort. The agent writes something down when it judges the moment worth remembering, which means there is no guarantee that a given correction becomes durable state. Nothing makes capture transactional: no "this decision must be recorded before we continue." Persistence exists; a promise that every consequential correction survives does not. The fact you most needed kept is often the one that was never written.

2. Retrieval: does the right thing reach the model at the right moment?

Getting a fact into storage is not the same as getting it into the next call, correctly, when it counts. Eager injection puts a fixed set of instructions in front of the model whether they are relevant or not, and being inside the context window is not the same as being used: research on long-context models finds they can miss information that is present, especially when it sits among a lot of competing text (the "lost in the middle" effect).

A common RAG baseline is dense-vector similarity search. It helps, but it is only one primitive: it ranks by semantic relatedness, and related is not the same as authoritative or correct. It struggles exactly where precision matters most: exact identifiers, versioned facts, negations, and competing near-duplicates. Across four repos, "the client" in one service is not "the client" in another, and a dense retriever can rank the wrong one first, because semantic proximity carries no notion of which source is authoritative. Serious retrieval mixes dense vectors with lexical search, metadata filters, reranking, and structure. A single similarity index is the naive baseline.

3. Source of truth: is there one current, authoritative value?

The same fact gets stated in a CLAUDE.md, in an auto-memory note, in a code comment, in a spec, each written at a different time. The underlying files may well have version history; git and note timestamps exist. What the harness does not give the model is a semantic rule that says "this decision supersedes that one" across those heterogeneous sources. So multiple representations stay simultaneously visible, and when they disagree the model sees all of them at once and picks, with no notion of which is current. Multiple repositories multiply it: each holds its own slice, and nothing reconciles them.

4. Honoring decisions: is the rule applied at the moment of action?

This is the one people underrate. Recall is not enough. A decision has operational value only if it reaches the agent at the point where it acts, and, when it must not be broken, is backed by a check that can fail.

Make it concrete. Your team decides that an ambiguous search must abstain rather than guess: precision over recall. It goes in a design note. Weeks later an agent is editing that search path and adds a fallback that returns the top-ranked match whenever the confidence check fails. Recall tests improve. The wrong-result rate climbs. The decision still exists in the note, but it never entered the agent's context while it wrote the fallback, and nothing checked the change against it.

A decision that is recorded but not applied at the point of action is a comment nobody read.

5. Scope and access: who sees what, and can you trust it?

Not everything should be visible to everyone, or to every agent. A flat instruction file, or a similarity index on its own, does not give you permission-aware retrieval: ask a related question and it surfaces whatever sits nearby, including what should have stayed private. The controls have to live in the retrieval path, identity, scope, metadata filtering, and post-retrieval authorization, not be hoped for afterward. This is also where the multi-agent problem lives: without an explicit design for what each agent may see and share, they diverge. And every fact that comes back should carry its provenance, where it came from, when, how sure, or the model cannot tell a current decision from a stale note or its own earlier guess.

Line the five up and they share one property. The default tooling persists information, but it does not manage that information's lifecycle.

GapHow you patch it todayWhy the patch breaks
CaptureHand-edit CLAUDE.md; hope auto-memory catches the restModel-decided, best-effort; no guaranteed record
RetrievalEager-load instructions, or add similarity searchIn-context is not reliably used; "related" is not "correct"
Source of truthA convention for where facts liveNo cross-source supersession; contradictions stay visible
Honoring decisionsA written tenet, a review checklistNot applied, or checked, at the moment of action
Scope and accessKeep sensitive things out of the filesNo permission-aware retrieval, no provenance

Every entry in that middle column is a manual norm: a habit, a convention, a document you promise to keep current. Norms hold until the project grows or the agents multiply, which is exactly when you lean on them, and exactly when they give.

The bridge: a memory operating system

The fix is not a bigger context window or a cleverer search box. It is to stop treating those five as habits and start treating them as a system's responsibility. Call that layer a memory operating system: the governed space between your agents and their knowledge, where each gap becomes an explicit, bounded mechanism.

Bounded is the key word. The goal is not magic, it is checkable behavior:

Notice these are engineering guarantees with edges, not promises of perfect recall. That is the difference between a memory you can reason about and a bigger pile of text.

Before you build anything: five questions for your own setup

You do not need to stand up a memory operating system this week to get value from this. You need to know where your current setup leaks. Run your own agent through the five gaps and answer honestly:

Wherever the honest answer is "nothing does," you have found the manual norm that will fail first as you scale. That is where a memory layer earns its place.

An agent that forgets your decisions every few days is not a model problem. It is a memory problem. And a memory problem is something you can engineer.

Part 2 is out

The Log Is the Truth: Building Agent Memory From Scratch

Part two builds one version of that layer, as a case study rather than the only shape it can take: an append-only event log as the source of truth, derived views for fast lookup, governed retrieval with abstention, and recall and precision measured rather than asserted. What the numbers were, where the naive version failed, the specific fix that moved it, and a step-by-step path to build a small one yourself, starting with a single file you can add to your workflow this week.

Sources

Get Internals in your inbox

Deep technical teardowns and build guides.

One rigorous piece at a time. Part 2 lands next.

Subscribe