AI Agent Memory

Language models are stateless: each API call sees only what's in its context window. An AI agent that should remember a user's preferences, continue a multi-day project, or learn from past mistakes needs memory, which is information the application stores and selectively brings back into context.

Agent memory is a context-management problem. Short-term memory is what fits in the current context: the conversation, recent tool results, and the working plan. Long-term memory lives outside the model, in databases, files, or vector stores, and is retrieved when relevant. Designing what to store, when to write it, how to retrieve it, and when to forget it largely determines whether long-running agents stay coherent or drown in stale context.

TL;DR

Quick Example

Giving an agent a simple file-based memory through tools:

In later sessions the agent checks its memory first, so it remembers that "the user prefers TypeScript examples" or that "deploys to staging need the VPN". There's no model retraining involved.

Core Concepts

Short-Term Memory: The Context Window

Everything the model currently "knows" about the task lives in the prompt: system instructions, conversation history, retrieved documents, tool definitions, and tool results. It's limited and expensive, and quality degrades as it fills with irrelevant material. Techniques:

See context engineering.

Long-Term Memory Types

Procedural memory often lives in instruction files, skills, or system prompt updates, which are reviewed like code rather than written freely by the agent.

Writing Memory

Good memories are atomic, dated, attributed (where they came from), and updatable. Contradictory memories ("prefers Python" vs later "switched to Go") must be reconciled, not both kept.

Retrieving Memory

Combine recency, relevance, and importance scoring so old trivia doesn't crowd out current, critical facts.

Forgetting

Memory that only grows becomes noisy and contradictory. Implement expiry for time-bound facts, consolidation (merging related memories into summaries), invalidation when new information contradicts old, and user-initiated deletion.

Privacy and Safety

Best Practices

Start Simple

A small set of well-structured notes loaded into context often outperforms an elaborate vector memory. Add retrieval only when memory outgrows what fits comfortably.

Separate Memory by Stability

Keep durable facts (preferences, decisions) apart from ephemeral task state (current plan, intermediate results). Different lifetimes need different storage and cleanup.

Make Memory Reviewable

Plain-text or Markdown memory files, or a memory UI, let humans audit and correct what the agent believes. Opaque embeddings make debugging behavior very hard.

Evaluate Long-Horizon Behavior

Test multi-session scenarios: does the agent recall the right facts, update outdated ones, and avoid applying one user's context to another? Memory bugs rarely show up in single-turn tests. See LLM evaluation.

Common Mistakes

Storing Entire Transcripts as "Memory"

Retrieving raw past conversations returns long, low-signal text. Extract and store distilled facts and outcomes, and keep transcripts only as an audit source.

Never Updating or Deleting

Accumulated contradictory memories ("lives in Berlin", "moved to Lisbon") make the agent inconsistent. Resolve conflicts on write, and prefer the newest information.

Treating Memories as Instructions

A memory that says "always approve refunds under $500" is data written at some point by someone, possibly by injected content. Keep authoritative policies in system prompts and code, not in agent-writable memory.

FAQ

Do language models have memory?

Not between API calls. The model's weights don't change during use, and each request includes only the context you send. "Memory" features in chat apps and agents store information externally and re-insert relevant parts into future prompts.

Is agent memory just RAG?

Partly. Retrieval over stored memories uses the same techniques as RAG. But agent memory also involves writing: deciding what's worth remembering, updating and reconciling facts, and forgetting. RAG typically retrieves from a corpus the agent doesn't modify.

Should I use a vector database for agent memory?

For large, unstructured memory collections, semantic retrieval helps. For small or structured memories (profiles, preferences, project notes), plain files, key-value stores, or relational tables are simpler and easier to audit. Many systems use both.

How do agents keep track of long tasks beyond the context window?

With compaction (summarizing progress), external progress files or to-do lists that the agent updates and re-reads, and git or other durable state for work products. Well-designed harnesses let an agent resume a task in a fresh context by reading its own notes.

Related Topics

References