Agent memory
Agent memory is how an AI agent saves useful information outside the model and loads the right pieces back into its prompt later.
Agent memory is the part of an AI agent that remembers things between conversations. A language model is stateless: it keeps nothing from one call to the next. Each call knows only what is in its prompt, the text sent to it. So memory lives outside the model, in the app’s own storage. The agent’s code copies the useful parts into the prompt when they are needed.
Builders usually split memory in two. Short-term memory is the message history of one conversation, which LangGraph calls a thread. Long-term memory lasts across conversations. It is often filed under a user’s ID.
Think of a group chat about a school trip. What was said in the last hour is short-term memory: it helps you follow the talk right now. Knowing that one friend is vegetarian is long-term memory: it still matters next month, in a different chat. LangGraph saves short-term memory to a database after each step. That way a paused conversation can pick up again later. A long message list costs money, and in the end it won’t fit. So apps often trim or drop old messages. Long-term memory is saved in one of two ways: during the conversation, or later by a separate background job. The OpenAI Agents SDK loads the saved history before each run and stores the new messages afterwards. In its example, the first question asks which city the Golden Gate Bridge is in. A second question asks only for the state. The reply is California, because the bridge was already mentioned.
Researchers also borrow names from studies of human memory. Semantic memory holds facts, such as facts about a user. Episodic memory holds experiences, such as the agent’s past actions. Procedural memory holds instructions, such as the agent’s system prompt (its standing instructions).
Choosing what to load back is the hard part. Generative Agents, a simulated town of 25 characters, stored every experience in a plain-language memory stream. Retrieval scored each memory on recency, importance and relevance. Relevance measured how close a memory’s meaning was to the current situation. It did this by comparing embeddings (lists of numbers that encode meaning). The model rated importance from 1 to 10: tidying a room scored 2, while asking someone out on a date scored 8. The top-scoring memories went into the prompt, as many as would fit.
Products follow the same pattern. ChatGPT can hold on to details a person tells it to keep, and it can also draw on that person’s earlier chats. Anthropic’s memory tool works differently. Claude only asks for changes to files inside a /memories folder. The developer’s own code makes each change and chooses where the files are kept.
A language model keeps nothing between calls, and its prompt has a size limit.
Follow two facts through MemGPT, a research agent that manages its own memory.
- 1 · writeWhen the user mentions a birthday, the agent calls a function that adds Birthday is February 7 to its working context, a small notes block kept inside the prompt.
- 2 · evictAs the prompt nears its limit, older messages are pushed out and summarised, but every message stays in a searchable store called recall storage.
- 3 · searchWhen the user later mentions Six Flags, the agent searches recall storage for six flags and pulls the 3 matching old messages back into its prompt.
- 4 · replyOne of those messages says the user and James first met at Six Flags, so the reply can draw on it.
- 5 · updateWhen the user says they broke up, the agent replaces Boyfriend named James with Ex-boyfriend named James.
The model stores nothing itself: the application around it saves memories and puts the relevant ones back into the prompt.
| Who | What they ask | What it works with |
|---|---|---|
| ChatGPT user | “Write up today's meeting” | The layout they once said they want for meeting notes |
| Support team using Claude | “Draft a reply to this customer ticket” | A customer service guidelines file in its /memories folder |
| Chat app user | “Which state is that bridge in?” | Their earlier Golden Gate Bridge question, loaded from the saved session |
| Builder of a simulated town | “What should this character do next?” | Its memory stream, ranked by recency, importance and relevance |
- An agent can carry facts, such as a user's preferences, into later conversations instead of asking again.
- The prompt stays focused. The agent fetches a stored note only when a task calls for it, instead of stuffing every note in at the start.
- On questions about five earlier chat sessions, MemGPT answered far more accurately than the same models given only a summary.
- Stored memories can be inspected and edited. ChatGPT users can view and delete specific memories in settings.
- Retrieval can miss. Two of the most common mistakes in the Generative Agents study were failing to fetch the right memory and adding made-up details to it.
- An agent may give up too soon. MemGPT frequently quit before it had read every page of its search results.
- This style of memory depends on the model calling functions well. On GPT-3.5, which handles function calls less reliably, MemGPT's scores dropped sharply.
- Stored memory adds privacy and security work. Removing a ChatGPT chat leaves the memories taken from it in place, and a Claude memory handler has to refuse any file path that points outside /memories.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMemGPT: Towards LLMs as Operating Systems, Packer et al., UC Berkeley · read 28 Sept 2026
- paperGenerative Agents: Interactive Simulacra of Human Behavior, Park et al., UIST 2023 · read 28 Sept 2026
- paperCognitive Architectures for Language Agents, Sumers et al., Transactions on Machine Learning Research · read 28 Sept 2026
- docsMemory overview, LangChain · read 28 Sept 2026
- docsSessions, OpenAI Agents SDK · read 28 Sept 2026
- officialMemory and new controls for ChatGPT, OpenAI · read 28 Sept 2026
- docsMemory tool, Anthropic · read 28 Sept 2026