Chapter 4: Memory & Context ​
If the LLM is the processor (CPU), then context is the RAM. By default, LLMs have amnesia. Every API call is a brand-new event unless you deliberately pass memory forward.
This chapter shows how to give your agent a working memory, long-term memory, and retrieval using LangChain-first examples with clear visuals.
What You Will Learn ​
- How short-term memory actually works (the pass-through trick)
- How to manage the context window (sliding window + summarization)
- How long-term memory is built with RAG
- When to use RAG vs MCP for context
- How to build a PDF Q&A agent with LangChain
The Core Mental Model ​
- Context = what the model sees right now
- Short-term memory = conversation buffer
- Long-term memory = retrieval from a knowledge store
1. Short-Term Memory (The Conversation Buffer) ​
Short-term memory is how the agent remembers what you said two turns ago. It is not magic. It is a pass-through technique.
Every turn, you send the entire conversation so far.
Pass-Through Example ​
Turn 1
- Input:
User: Hi - Output:
AI: Hello
Turn 2
- Input:
User: Hi, AI: Hello, User: My name is Sarah - Output:
AI: Nice to meet you, Sarah.
Turn 3
- Input:
User: Hi, AI: Hello, User: My name is Sarah, AI: Nice to meet you, User: Who am I? - Output:
AI: You are Sarah.
The Context Window Problem ​
LLMs cannot accept infinite history. Each model has a context window (e.g., 16k, 128k tokens). Past that limit, the model either fails or the request becomes too expensive.
Two Classic Solutions ​
A. Sliding Window Keep only the last N turns and drop the rest.
def sliding_window(messages, max_turns=10):
return messages[-max_turns:]function slidingWindow(messages, maxTurns = 10) {
return messages.slice(-maxTurns);
}B. Summarization Summarize older turns into a compact memory note.
def summarize_history(llm, messages):
prompt = (
"Summarize the conversation in 5 bullet points. "
"Preserve user goals and preferences. Omit small talk."
)
text = "\n".join([f"{m['role']}: {m['content']}" for m in messages])
return llm.invoke(prompt + "\n\n" + text).contentimport OpenAI from "openai";
const openai = new OpenAI();
async function summarizeHistory(messages) {
const prompt =
"Summarize the conversation in 5 bullet points. " +
"Preserve user goals and preferences. Omit small talk.";
const text = messages.map((m) => `${m.role}: ${m.content}`).join("\n");
const response = await openai.chat.completions.create({
model: "gpt-4o",
messages: [{ role: "user", content: prompt + "\n\n" + text }],
});
return response.choices[0].message.content;
}LangChain Example: Conversation Buffer ​
This is the simplest short-term memory. It keeps all turns in memory and passes them through automatically.
from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationChain
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(temperature=0)
memory = ConversationBufferMemory()
conversation = ConversationChain(
llm=llm,
memory=memory,
verbose=True
)
conversation.predict(input="My name is Sarah.")
conversation.predict(input="What is my name?")import OpenAI from "openai";
const openai = new OpenAI();
// Maintain conversation history manually (equivalent to ConversationBufferMemory)
const messages = [];
async function chat(userInput) {
messages.push({ role: "user", content: userInput });
const response = await openai.chat.completions.create({
model: "gpt-4o",
temperature: 0,
messages,
});
const reply = response.choices[0].message.content;
messages.push({ role: "assistant", content: reply });
return reply;
}
await chat("My name is Sarah.");
console.log(await chat("What is my name?"));2. Long-Term Memory (RAG) ​
Short-term memory disappears when your script ends. Long-term memory persists across sessions.
RAG (Retrieval-Augmented Generation) is the standard way to build long-term memory.
Think of RAG as an open-book exam:
- Standard LLM: answer from its own memory
- RAG agent: search the library, find the right page, then answer
The RAG Pipeline ​
- Ingest: Load documents (PDFs, text files, Notion pages)
- Chunk: Split into smaller pieces (500-1000 words)
- Embed: Convert chunks into vectors (numbers)
- Store: Put vectors in a vector database
- Retrieve: Fetch nearest chunks for a query
- Generate: Answer based on retrieved context
Vector Databases (Quick Start) ​
- Pinecone (cloud, managed)
- Chroma (local, easy dev)
- FAISS (local, fast)
- PGVector (PostgreSQL users)
3. The New Standard: MCP (Model Context Protocol) ​
In 2024-2025, a new standard emerged: MCP (Model Context Protocol). Think of it as USB-C for AI context.
Before MCP, every integration was custom. If you wanted Google Drive + Slack + GitHub, you wrote custom code for each. That was integration hell.
What MCP Does ​
- MCP Server: a connector that exposes data in a standard format
- MCP Client: your agent, which plugs into the server
When to Use RAG vs MCP ​
- Use RAG for large, mostly static knowledge bases (manuals, policies, wikis)
- Use MCP for live systems and tools (databases, filesystems, APIs)
Use Cases: RAG vs MCP ​
RAG Use Cases (Static Knowledge)
- Employee handbook Q&A
- Product manuals and troubleshooting
- Internal wiki and SOP lookup
- Compliance policy search
- Research paper summarization
- Customer support knowledge base
MCP Use Cases (Live Systems)
- Read and summarize files from a shared drive
- Query a database for real-time metrics
- Pull issues and PRs from GitHub
- Fetch tickets from a helpdesk system
- Read Slack or Teams threads for context
- Update CRM notes or create tasks
4. Project: PDF Chat Agent (LangChain) ​
We will build a RAG agent that answers questions about a PDF.
Install Dependencies ​
pip install langchain-community langchain-openai faiss-cpu pypdfFlow Overview ​
Code: PDF Chat Agent ​
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import FAISS
from langchain.chains import RetrievalQA
from langchain_openai import ChatOpenAI
# 1. Load the PDF
loader = PyPDFLoader("policy.pdf")
documents = loader.load()
# 2. Chunk the text
splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=100
)
chunks = splitter.split_documents(documents)
# 3. Embed the chunks
embeddings = OpenAIEmbeddings()
# 4. Store in FAISS (local vector DB)
vectorstore = FAISS.from_documents(chunks, embeddings)
# 5. Create the retriever
retriever = vectorstore.as_retriever()
# 6. Connect to LLM
llm = ChatOpenAI(temperature=0)
qa_chain = RetrievalQA.from_chain_type(llm, retriever=retriever)
# 7. Ask a question
response = qa_chain.run("What is the vacation policy in this document?")
print(response)import fs from "fs";
import path from "path";
import OpenAI from "openai";
import { PDFExtract } from "pdf.js-extract";
// npm install openai pdf.js-extract
const openai = new OpenAI();
// 1. Load and extract text from PDF
async function loadPDF(filePath) {
const pdfExtract = new PDFExtract();
const data = await pdfExtract.extract(filePath, {});
return data.pages
.map((p) => p.content.map((c) => c.str).join(" "))
.join("\n");
}
// 2. Chunk the text
function chunkText(text, chunkSize = 1000, overlap = 100) {
const chunks = [];
let start = 0;
while (start < text.length) {
chunks.push(text.slice(start, start + chunkSize));
start += chunkSize - overlap;
}
return chunks;
}
// 3 & 4. Embed and store chunks in memory
async function embedChunks(chunks) {
const response = await openai.embeddings.create({
model: "text-embedding-3-small",
input: chunks,
});
return response.data.map((d) => d.embedding);
}
// 5. Retrieve top-k chunks via cosine similarity
function cosineSimilarity(a, b) {
const dot = a.reduce((sum, v, i) => sum + v * b[i], 0);
const normA = Math.sqrt(a.reduce((s, v) => s + v * v, 0));
const normB = Math.sqrt(b.reduce((s, v) => s + v * v, 0));
return dot / (normA * normB);
}
async function retrieve(query, chunks, embeddings, k = 3) {
const qEmbRes = await openai.embeddings.create({
model: "text-embedding-3-small",
input: [query],
});
const qEmb = qEmbRes.data[0].embedding;
return chunks
.map((c, i) => ({ chunk: c, score: cosineSimilarity(qEmb, embeddings[i]) }))
.sort((a, b) => b.score - a.score)
.slice(0, k)
.map((r) => r.chunk);
}
// 6 & 7. Answer question using retrieved context
async function askPDF(question, pdfPath) {
const text = await loadPDF(pdfPath);
const chunks = chunkText(text);
const embeddings = await embedChunks(chunks);
const context = await retrieve(question, chunks, embeddings);
const response = await openai.chat.completions.create({
model: "gpt-4o",
temperature: 0,
messages: [
{
role: "system",
content: "Answer questions based only on the provided context.",
},
{
role: "user",
content: `Context:\n${context.join("\n\n")}\n\nQuestion: ${question}`,
},
],
});
return response.choices[0].message.content;
}
const answer = await askPDF(
"What is the vacation policy in this document?",
"policy.pdf",
);
console.log(answer);What Just Happened? ​
The agent did not read the whole PDF. It found the most relevant chunk, inserted it into the prompt, and answered based on that evidence.
5. Putting It Together: Memory + RAG Agent (LangChain) ​
Now combine short-term memory with long-term memory.
Code: Memory + RAG ​
from langchain_openai import ChatOpenAI
from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationalRetrievalChain
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import PyPDFLoader
# Build vector store (same as before)
loader = PyPDFLoader("policy.pdf")
documents = loader.load()
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=100)
chunks = splitter.split_documents(documents)
vectorstore = FAISS.from_documents(chunks, OpenAIEmbeddings())
# Memory + Retrieval
memory = ConversationBufferMemory(
memory_key="chat_history",
return_messages=True
)
llm = ChatOpenAI(temperature=0)
chain = ConversationalRetrievalChain.from_llm(
llm=llm,
retriever=vectorstore.as_retriever(),
memory=memory
)
chain.invoke({"question": "What is the vacation policy?"})
chain.invoke({"question": "How many days does it allow?"})import OpenAI from "openai";
import { PDFExtract } from "pdf.js-extract";
// npm install openai pdf.js-extract
const openai = new OpenAI();
async function buildVectorStore(pdfPath) {
const pdfExtract = new PDFExtract();
const data = await pdfExtract.extract(pdfPath, {});
const text = data.pages
.map((p) => p.content.map((c) => c.str).join(" "))
.join("\n");
const chunks = [];
let start = 0;
while (start < text.length) {
chunks.push(text.slice(start, start + 1000));
start += 900;
}
const embRes = await openai.embeddings.create({
model: "text-embedding-3-small",
input: chunks,
});
return { chunks, embeddings: embRes.data.map((d) => d.embedding) };
}
function cosineSim(a, b) {
const dot = a.reduce((s, v, i) => s + v * b[i], 0);
return (
dot /
(Math.sqrt(a.reduce((s, v) => s + v * v, 0)) *
Math.sqrt(b.reduce((s, v) => s + v * v, 0)))
);
}
async function retrieveChunks(query, store, k = 3) {
const qRes = await openai.embeddings.create({
model: "text-embedding-3-small",
input: [query],
});
const qEmb = qRes.data[0].embedding;
return store.chunks
.map((c, i) => ({ c, score: cosineSim(qEmb, store.embeddings[i]) }))
.sort((a, b) => b.score - a.score)
.slice(0, k)
.map((r) => r.c);
}
// Memory + RAG chat
const store = await buildVectorStore("policy.pdf");
const chatHistory = [];
async function chat(question) {
const context = await retrieveChunks(question, store);
const messages = [
{
role: "system",
content:
"You are a helpful assistant. Answer based on the provided context and conversation history.",
},
...chatHistory,
{
role: "user",
content: `Context:\n${context.join("\n\n")}\n\nQuestion: ${question}`,
},
];
const res = await openai.chat.completions.create({
model: "gpt-4o",
temperature: 0,
messages,
});
const reply = res.choices[0].message.content;
chatHistory.push({ role: "user", content: question });
chatHistory.push({ role: "assistant", content: reply });
return reply;
}
console.log(await chat("What is the vacation policy?"));
console.log(await chat("How many days does it allow?"));Common Pitfalls ​
- Dumping too much history into the prompt
- Forgetting to summarize old turns
- Building RAG without citing sources
- Using memory for factual data instead of retrieval
Checklist ​
- My agent keeps short-term context clean
- My agent summarizes or trims old history
- My agent retrieves documents before answering
- My responses are grounded in sources
What Comes Next ​
In Chapter 5, you will build multi-agent workflows and learn how to coordinate agents for larger tasks.