Memory and Context Management
The Memory Problem
LLMs have no persistent memory between conversations. Every API call starts fresh -- the model has no recollection of previous interactions unless you explicitly include that history in the prompt. For simple chatbots, passing the full conversation history works fine. For agents handling long, complex tasks, this approach quickly runs into problems.
Context windows -- the maximum amount of text an LLM can process at once -- are finite. Even the largest models (128K tokens for GPT-4o) can be overwhelmed by long agent sessions that accumulate many tool calls and observations. When the context fills up, the model either truncates early history or performance degrades.
Understanding memory management is what separates agents that work reliably on complex tasks from those that only work in demos.
The Four Types of Agent Memory
1. In-Context Memory (Working Memory)
The conversation history within the current API call. This is the simplest and most immediate form of memory.
What it contains: System prompt, user message, all previous Thought/Action/Observation steps in the current session.
Limit: Bounded by the model's context window (e.g., 128K tokens for GPT-4o).
Best for: Short to medium tasks where the full history fits comfortably in context.
2. External Storage (Database Memory)
Information stored in a database, file system, or key-value store that the agent can read and write using tools.
What it contains: Previous session summaries, user preferences, domain knowledge, cached results.
Best for: Persistent information that should survive across multiple agent sessions (user profiles, previous task outcomes, learned preferences).
3. Episodic Memory (Past Session Records)
Records of previous agent sessions -- what was attempted, what succeeded, what failed. The agent can learn from past experience by retrieving relevant records.
Best for: Agents that run repeatedly on similar tasks and should improve over time.
4. Semantic Memory (Vector Database)
Information stored as vector embeddings that enable meaning-based retrieval. The agent can search its knowledge base semantically: "Find what I know about customer complaints similar to this one" rather than exact keyword lookup.
Best for: Large knowledge bases, document retrieval, retrieval augmented generation (covered in Module 6).
Managing Context Window Overflow
As agent sessions grow longer, you must actively manage the context window:
Strategy 1: Summarisation
Periodically summarise the conversation history to compress earlier steps:
async function compressHistory(messages) {
// When messages exceed a threshold, summarise older ones
const MAX_MESSAGES = 20;
if (messages.length <= MAX_MESSAGES) {
return messages;
}
// Keep the system prompt and last 10 messages
const systemPrompt = messages[0];
const recentMessages = messages.slice(-10);
const olderMessages = messages.slice(1, -10);
// Summarise the older messages
const summaryResponse = await callLLM([
{ role: 'user', content: 'Summarise what has happened so far in this agent session. Be concise but capture key findings and decisions:\n\n' + olderMessages.map(m => m.role + ': ' + m.content).join('\n') }
]);
const summary = summaryResponse.choices[0].message.content;
return [
systemPrompt,
{ role: 'assistant', content: '[Summary of earlier steps] ' + summary },
...recentMessages
];
}
Strategy 2: Selective Retention
Not every tool result needs to be kept in full. For large tool outputs, summarise or truncate:
function truncateToolResult(result, maxChars = 2000) {
const resultStr = JSON.stringify(result);
if (resultStr.length <= maxChars) return result;
return {
truncated: true,
summary: 'Result was ' + resultStr.length + ' chars. Key data: ' + resultStr.substring(0, maxChars) + '...',
note: 'Use the tool again with more specific parameters to get the full result.'
};
}
Strategy 3: Hierarchical Memory
For very long tasks, use a hierarchical approach:
- Working memory: Last 5-10 steps in full detail
- Short-term summary: A paragraph summarising the last 30 minutes of work
- Task state: A structured object tracking what has been completed and what remains
Maintaining User State Across Sessions
For agents that interact with the same user over time, persist user state:
// Simple user state management
class UserStateManager {
constructor(database) {
this.db = database;
}
async loadUserContext(userId) {
const state = await this.db.get('user_state:' + userId);
return state || {
preferences: {},
previousSessions: [],
openTasks: [],
lastInteraction: null
};
}
async saveUserContext(userId, state) {
state.lastInteraction = new Date().toISOString();
await this.db.set('user_state:' + userId, state);
}
async buildContextPrompt(userId) {
const state = await this.loadUserContext(userId);
if (!state.previousSessions.length) return '';
const lastSession = state.previousSessions[state.previousSessions.length - 1];
return `Previous interaction summary: ${lastSession.summary}.
Open tasks from last session: ${state.openTasks.join(', ') || 'none'}.
User preferences: ${JSON.stringify(state.preferences)}.`;
}
}
Conversation Compression in Practice
Here is a practical implementation for a long-running agent:
// Token counting and context management
function estimateTokenCount(text) {
// Rough estimate: 1 token ≈ 4 characters
return Math.ceil(text.length / 4);
}
function buildAgentContext(systemPrompt, history, maxTokens = 100000) {
const systemTokens = estimateTokenCount(systemPrompt);
let availableTokens = maxTokens - systemTokens - 2000; // Reserve 2K for response
// Start from the end and work backwards
const includedMessages = [];
for (let i = history.length - 1; i >= 0; i--) {
const msgTokens = estimateTokenCount(JSON.stringify(history[i]));
if (availableTokens - msgTokens < 0) break;
includedMessages.unshift(history[i]);
availableTokens -= msgTokens;
}
if (includedMessages.length < history.length) {
// Add a note about truncated history
includedMessages.unshift({
role: 'assistant',
content: '[Note: Earlier conversation history has been omitted due to length. The agent has access to the most recent ' + includedMessages.length + ' messages.]'
});
}
return [{ role: 'system', content: systemPrompt }, ...includedMessages];
}
When to Use Each Memory Type
| Scenario | Memory Type |
|---|---|
| Short task within one session | In-context only |
| Long task within one session | In-context + summarisation |
| Returning user with preferences | External database |
| Agent that learns from past runs | Episodic memory |
| Large knowledge base queries | Vector database (RAG) |
Key Takeaways
- LLMs have no persistent memory -- every API call starts fresh unless you include history in the prompt.
- The four types of agent memory are: in-context (working), external storage (database), episodic (past sessions), and semantic (vector database).
- Context window management strategies include summarisation, selective retention, and hierarchical memory.
- Persist user state in a database for agents that interact with the same user across multiple sessions.
- Token counting and context budgeting are practical skills every agent builder needs to master.
Try it yourself
Key Takeaways
- LLMs have no persistent memory between API calls -- all context must be explicitly provided in each request.
- The four memory types are: in-context (working), external storage (database), episodic (past sessions), and semantic (vector database).
- Context window management strategies include summarisation (compressing older steps), selective retention, and hierarchical memory.
- Persist user state in a database for agents that interact with the same users across multiple sessions.
- Estimate token counts and budget your context window proactively -- reactive handling of overflow causes reliability problems.
Quick Quiz
1.Why do AI agents need special memory management for long tasks?
2.What is the summarisation strategy for context management?
3.What is episodic memory in the context of AI agents?
4.Which memory type is best for an agent that needs to search a large knowledge base to answer questions?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx