1. Library
  2. Long-Horizon Agents, Part 2: Memory Is More than Recall

Long-Horizon Agents, Part 2: Memory Is More than Recall

20 mins
Light Mode
Why Memory Is A Key Piece of the Long-Horizon Puzzle
Memory as a Stateful Control Loop
Write Policy: What Becomes Memory?
Read Policy: What Reaches the Model?
Storage: Representation, Backend, and Indexes
How Existing Systems Combine These Choices
Letta and Letta Code
Mem0's Managed Platform
Graphiti and Zep
cognee
Mubit

Why Memory Is A Key Piece of the Long-Horizon Puzzle

In Part 1, I laid out the five systems I believe are missing before agents can graduate from completing tasks to owning outcomes: Memory, learning, goal orchestration, a durable execution layer, and a toolchain ecosystem. This post dives into agent memory.

To understand agent memory, we shouldn't think of it as a place to store the past. We need to think of it as a harness-level control system that decides which previous experiences should affect future behavior. The need for memory in the long-horizon context should be clear, but to provide a few examples:

  • An agent cannot pursue a goal over months if it forgets why the goal/sub-goal matters
  • Learning can't happen without memory, and an agent cannot learn if the outcome of one attempt disappears before the next
  • It cannot coordinate with other agents if each action begins with an isolated view of the organization

It is tempting to describe this as a storage problem. Save the agent's interactions, embed them, and retrieve the closest results before the next model call. That produces persistence, but not necessarily memory. A useful memory system must decide what to write, how to represent it, when to retrieve it, how to resolve conflicting information, and whether/how an outcome should modify future behavior. Current systems implement this functionality outside the model, in the harness layer that controls the agent loop.

Memory as a Stateful Control Loop

At time t, an agent receives an observation, assembles context and chooses an action. We can therefore describe a simplified agent without memory as:

action_t = model(instructions, recent_messages, observation_t)

A memory-augmented agent adds persistent state M_t, a write policy F, a retrieval policy R, and a context budget B:

context_t = R(query_t, M_t, B)
action_t = model(
instructions,
working_state_t,
context_t,
observation_t
)
M_t+1 = F(
M_t,
observation_t,
action_t,
outcome_t
)

F determines which observations become persistent state and how they revise existing records. R selects the state that reaches the next model call. Both can combine stochastic compute and deterministic code. Outcomes may arrive much later than actions. The write path must associate delayed evidence with the original action and revise any lessons derived from it.

The implementation of memory comes down to three connected choices: the write policy, the read policy, and the storage representation. Claude Code gives us a concrete example of how a harness combines them.

I came across himanshu's walkthrough of Claude Code's memory architecture while reading about memory systems. The diagram separates the paths that create and revise memory from the paths that bring it back into context:

Source: himanshu on X, March 31, 2026. This is a source-code analysis, not an official architecture diagram. Internal mechanisms such as autoDream and extractMemories reflect the version analyzed. The walkthrough below uses public documentation for shipped behavior.

To follow these choices through one task, imagine using Claude Code to investigate a database migration that broke an older worker. During the investigation, the developer explains that customers upgrade their workers independently of the backend, and that a rollout tracker records which versions are still running. What should survive it, and how should that information affect the next migration?

Write Policy: What Becomes Memory?

A write policy determines what gets retained, who interprets it, and when the result becomes available. These three are most common:

Automatically capture execution. Claude Code records messages, tool calls, and results. Recording means preserving a failed query + investigation + any dev explanation, but does not require immediate decisions on which details matter. Auto-capture moves the cost of the write policy to read, as someone will need to search for and interpret it later (which would mean that this approach also requires retention).

Let the active agent write a memory. Claude Code's auto memory selectively saves user info, feedback, context, and references, skipping facts recoverable from code or Git history. The difference is what is worth saving (such as which data column changed), while the operating constraint, and where to verify it, becomes durable. This approach requires less reconstruction on the next run. However, there’s a risk of misinterpreting a one-off issue, like a customer delay, as a hard-and-fast rule to implement (which can be mitigated with notes/references to supporting evidence).

Use a separate extraction pipeline. Another approach is processing recorded sessions independently of the agent with an extractor. Converting the resulting data from the job into appropriate data fields provides a consistent schema and lets us rerun extraction when the policy changes. However, the extra processing adds cost and potential risk of misinterpreting evidence. It also separates extraction (used to identify candidate facts) from consolidation (deciding whether new facts should update an existing memory). See Letta Code's context repositories for documented examples of using background consolidation to update agent memory.

Note: These write policies are not mutually exclusive. For example, you could retain an exchange, add a brief reference, and reconcile it with other experiences later. Of course, you would need to account for delays if subsequent processing happens in the background. Conversely, you would have to account for newly-discovered constraints if you ran immediate retries.

Read Policy: What Reaches the Model?

After noting changes from their write policy, agents need to understand whether written changes affect them and any upgrades they have received. These three read policies are most common:

Include designated records automatically. One approach is to keep a small set of instructions or pointers within context (Claude Code loads MEMORY.md at conversation start). Here, record storage does not count against total stored memory, as any reading of detailed topic files is done on demand. This approach connects read with write because a concise and descriptive entry makes the right notes easier to discover (and bad descriptions make it harder). This approach makes discovery more predictable but consumes context on unrelated tasks, and requires a choice about what is worth storing. (For reference, Claude Code distinguishes between persistent instructions vs. auto memory.)

Automatic inclusion makes discovery more predictable, but consumes context even on unrelated tasks. It also requires a choice about what deserves that space. CLAUDE.md supplies persistent instructions; auto memory contains Claude's own notes. Both guide behavior rather than enforce it. A deployment constraint that must be guaranteed needs an execution check as well as a reminder. Claude Code's documentation explains this distinction.

Have the harness retrieve and assemble context. Another approach runs retrieval before selected model calls, without waiting for the model to request it. This policy requires a migration-aware harness to proactively identify any affected services/schemas, retrieve any relevant deployment constraints, and attach them to the next call. (This approach would support exact filters with structured records and allow free-form notes to be text-searched.) The benefit here is that retrieval is not constrained by whether agents think to search. However, this policy carries the risk that the harness might introduce new issues (like constructing a poor query, returning stale records, or filling context with superficially similar stuff).

Let the model retrieve memory as needed. Another approach is to let agents use tools to investigate. For example, Claude Code combines upfront context with on-demand retrieval, so the model can discover a clue, refine its question, and perform additional retrieval. Agents using this policy might start by reading a rollout note, follow its reference to the tracker, and check current worker versions before proposing a migration (or, if the task was about why the original migration failed, it would search retained session records for more details). This approach adds flexibility but also adds costs for model calls and tool use. Having a compact index can help an investigation start better, but isn’t a guarantee it will end well, as agents may fail to search, choose the wrong note, or stop short of checking the source.

About retrieval and memory overhead: Who initiates retrieval is separate from how retrieval works. It’s possible for both automatic or agentic-requested read requests to use exact lookups, keyword search, vector similarity, or graph traversal. (Harnesses can also combine policies like including a small index while auto-retrieving known task constraints and letting agents investigate further.) The goal is to select enough evidence for decision-making without having to store the entire history into every call.

Compaction (condensing older conversation history, reasoning, and tool outputs to conserve context window) can also help here. Claude Code’s approach is to summarize a growing conversation within its context window, supporting continuity within a session while treating the selection of durable lessons for future sessions as a separate decision.

Storage: Representation, Backend, and Indexes

There are multiple ways to represent an agent’s experience, including an execution record, a written lesson, a structured fact, or a relationship. Your choice determines what aspects of an agentic job you can preserve and query directly. (Your backend determines how records persist and how to coordinate concurrent access.) We’ll cover four approaches below:

Raw logs and artifacts. Claude Code's session transcripts show that the model retains interactions. (Files or object storage can also hold command output, patches, etc..) Such artifacts help agents investigate and reprocess, but even a full history does not guarantee you can immediately identify which constraint applies to the current job. For that, you would need your agents to run a subsequent read or derived view.

Text files and summaries. Claude Code stores auto memory as Markdown under ~/.claude/projects/<project>/memory/, with a MEMORY.md index and individual topic files. Different memory types appear in frontmatter, but the default memory is local to the machine.

Here is an illustrative layout for our example:

memory/
├── MEMORY.md
├── feedback_migration_compatibility.md
└── reference_customer_rollout_tracker.md

The index helps discover the notes, which explain the constraint and point to its source. Using ordinary file tools makes such summaries easy to inspect and revise. Also, files can contain structured metadata, even if our system chooses to use Markdown. The challenge here is consistency, since notes can be readable while contradicting each other. Adding versioning makes changes reversible, but does not validate them. (See Letta Code, below, for an example.)

Structured records. If you need to answer questions across many customers and deployments, another approach is using rows or documents for important fields like customer_id, worker_version, observed_at, and source_event_id. Using fields lets agents filter and make explicit version checks. However, you would need a schema that tracks version changes among workers, which is a distinction that can be lost in a field overwrite unless you also retain earlier versions.

Entity-relationship graphs. Using graphs could connect customers to worker fleets, fleets to software versions, and versions to schema requirements. A key feature of graphs is temporal relationships, which could indicate when each dependency applied during the run. With graphs, multi-step questions (like which customers still depend on a particular column) become easier to express. However, as underlying deployments change, graphs require extraction, entity resolution, and updates to remain accurate.

Note: Vector indexes (specialized data structures that organize embeddings to speed up searches) can sit alongside different representations to find similar content. However, vector indexes do not determine whether facts are current, whether records refer to the same worker, or whether a proposed lesson is justified. For example, selection, organization, and retrieval seem like a huge part of memory design for Claude Code, even when the durable records are ordinary files. (For more on latent and parametric memory, see the report Memory in the Age of AI Agents.)

How Existing Systems Combine These Choices

Below are some real-world examples of how different platforms approach agentic memory, at least, according to publicly-available source code or architecture explainers:

Letta and Letta Code

  • Write: Agents edit persistent memory using tools, with background agents handling consolidation. Letta Code's context repositories let memory agents work in isolated Git worktrees and merge their changes.
  • Read: Letta uses a combination of automatically including records and using agents to access additional memory through tools. In context repositories, the file hierarchy helps it discover material to load on demand.
  • Store: Letta abstracts context window storage as memory blocks, which persist labeled, size-limited text. Context repositories represent memory as versioned files, with designated files loaded into the prompt.

Letta’s approach gives capable models flexibility over both representation and retrieval. Because Letta allows for persistent memory editing and on-demand retrieval, the correctness of a job depends on editing and search decisions. Version control provides rollback capability, but does not validate the meaning of a memory update.

Mem0's Managed Platform

  • Write: Mem0’s extraction pipeline processes conversations asynchronously, extracts ADD-only facts, deduplicates them, creates embeddings, and adds entity links and temporal metadata.
  • Read: Mem0 ranks information relevance based on semantic, BM25, entity, and temporal signals. Temporal relevance ranks information by freshness rather than immediately excluding stale candidates.
  • Store: Memory text, embeddings, and metadata live in a vector database. Mem0 uses an entity store to link related memories, and retains additional history and a rolling message window using SQL.

Mem0’s approach uses a standardized pipeline to interpret information during an agentic run. The platform takes the approach of preserving evidence by appending changed facts, but it leaves version resolution up to retrieval and reader. (Mem0 also publishes an open-source implementation that differs from the managed architecture described above.)

Graphiti and Zep

  • Write: Graphiti extracts entities and relationships from incoming episodes. The platform can preserve evidence history but allows new evidence to invalidate previous relationships.
  • Read: The platform uses a hybrid search approach that combines semantic and BM25 retrieval with rank fusion. It also offers configurable recipes, which add graph traversal and node-distance ranking/cross-encoder reranking.
  • Store: Graphiti uses a variety of artifacts for storage. Episodes retain source evidence, graph nodes represent entities, and edges represent relationships with temporal validity. (It should be noted that Graphiti is the open-source framework and Zep provides managed context-graph infrastructure.)

Graphiti’s approach to resolving changes during ingestion makes applicable state more explicit at read time. The trade-off here is that correct extraction becomes much more important. If the agent merges distinct workers or invalidates the wrong relationship, it may also corrupt subsequent retrieval.

cognee

  • Write: cognee uses configurable pipelines to ingest data, extract entities and relationships, and enrich existing memory. Later, the platform can also run improvement passes that incorporate useful session information into the permanent graph.
  • Read: cognee’s retrieval uses vectors and graph structure. The recall interface supports different modes according to the evidence/answer required.
  • Store: cognee’s architecture combines relational records (for source metadata and provenance), vectors (for semantic search), and a graph (for entities and relationships).

cognee’s approach lets teams customize the process of transforming source data to searchable memory. However, the approach also uses multiple derived representations, which creates extra maintenance overhead. Specifically, failed jobs can leave indexes incomplete which means not every dependent record instantly auto-updates with any corrections.

Mubit

  • Write: Mubit uses asynchronous ingestion to classify typed entries. The platform has a learning loop that attaches outcomes to memory identifiers and uses reflection to derive conditional lessons for validation and promotion.
  • Read: The platform uses recall queries to combines semantic, lexical and temporal signals with working state and lesson overlays. The platform ranks information based on outcome history, and also has a direct mode that returns evidence without model-based routing or answer synthesis.
  • Store: Mubit’s documented memory model includes facts, traces, lessons, rules, workflows, and exact archives, with a separate, mutable working state (rather than specifying a particular database backend).

Mubit’s system is able to prioritize procedural guidance over raw history using typing and outcome signals. This is a good benefit that depends on reliable attribution, since even a successful deployment doesn’t necessarily prove that every retrieved lesson contributed to it.