Skip to content

Chapter 4: Memory & Context ​

If the LLM is the processor (CPU), then context is the RAM. By default, LLMs have amnesia. Every API call is a brand-new event unless you deliberately pass memory forward.

This chapter shows how to give your agent a working memory, long-term memory, and retrieval using LangChain-first examples with clear visuals.

What You Will Learn ​

  • How short-term memory actually works (the pass-through trick)
  • How to manage the context window (sliding window + summarization)
  • How long-term memory is built with RAG
  • When to use RAG vs MCP for context
  • How to build a PDF Q&A agent with LangChain

The Core Mental Model ​

  • Context = what the model sees right now
  • Short-term memory = conversation buffer
  • Long-term memory = retrieval from a knowledge store

1. Short-Term Memory (The Conversation Buffer) ​

Short-term memory is how the agent remembers what you said two turns ago. It is not magic. It is a pass-through technique.

Every turn, you send the entire conversation so far.

Pass-Through Example ​

Turn 1

  • Input: User: Hi
  • Output: AI: Hello

Turn 2

  • Input: User: Hi, AI: Hello, User: My name is Sarah
  • Output: AI: Nice to meet you, Sarah.

Turn 3

  • Input: User: Hi, AI: Hello, User: My name is Sarah, AI: Nice to meet you, User: Who am I?
  • Output: AI: You are Sarah.

The Context Window Problem ​

LLMs cannot accept infinite history. Each model has a context window (e.g., 16k, 128k tokens). Past that limit, the model either fails or the request becomes too expensive.

Two Classic Solutions ​

A. Sliding Window Keep only the last N turns and drop the rest.

python
def sliding_window(messages, max_turns=10):
    return messages[-max_turns:]
javascript
function slidingWindow(messages, maxTurns = 10) {
  return messages.slice(-maxTurns);
}

B. Summarization Summarize older turns into a compact memory note.

python
def summarize_history(llm, messages):
    prompt = (
        "Summarize the conversation in 5 bullet points. "
        "Preserve user goals and preferences. Omit small talk."
    )
    text = "\n".join([f"{m['role']}: {m['content']}" for m in messages])
    return llm.invoke(prompt + "\n\n" + text).content
javascript
import OpenAI from "openai";

const openai = new OpenAI();

async function summarizeHistory(messages) {
  const prompt =
    "Summarize the conversation in 5 bullet points. " +
    "Preserve user goals and preferences. Omit small talk.";
  const text = messages.map((m) => `${m.role}: ${m.content}`).join("\n");
  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    messages: [{ role: "user", content: prompt + "\n\n" + text }],
  });
  return response.choices[0].message.content;
}

LangChain Example: Conversation Buffer ​

This is the simplest short-term memory. It keeps all turns in memory and passes them through automatically.

python
from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationChain
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(temperature=0)
memory = ConversationBufferMemory()

conversation = ConversationChain(
    llm=llm,
    memory=memory,
    verbose=True
)

conversation.predict(input="My name is Sarah.")
conversation.predict(input="What is my name?")
javascript
import OpenAI from "openai";

const openai = new OpenAI();

// Maintain conversation history manually (equivalent to ConversationBufferMemory)
const messages = [];

async function chat(userInput) {
  messages.push({ role: "user", content: userInput });
  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    temperature: 0,
    messages,
  });
  const reply = response.choices[0].message.content;
  messages.push({ role: "assistant", content: reply });
  return reply;
}

await chat("My name is Sarah.");
console.log(await chat("What is my name?"));

2. Long-Term Memory (RAG) ​

Short-term memory disappears when your script ends. Long-term memory persists across sessions.

RAG (Retrieval-Augmented Generation) is the standard way to build long-term memory.

Think of RAG as an open-book exam:

  • Standard LLM: answer from its own memory
  • RAG agent: search the library, find the right page, then answer

The RAG Pipeline ​

  1. Ingest: Load documents (PDFs, text files, Notion pages)
  2. Chunk: Split into smaller pieces (500-1000 words)
  3. Embed: Convert chunks into vectors (numbers)
  4. Store: Put vectors in a vector database
  5. Retrieve: Fetch nearest chunks for a query
  6. Generate: Answer based on retrieved context

Vector Databases (Quick Start) ​

  • Pinecone (cloud, managed)
  • Chroma (local, easy dev)
  • FAISS (local, fast)
  • PGVector (PostgreSQL users)

3. The New Standard: MCP (Model Context Protocol) ​

In 2024-2025, a new standard emerged: MCP (Model Context Protocol). Think of it as USB-C for AI context.

Before MCP, every integration was custom. If you wanted Google Drive + Slack + GitHub, you wrote custom code for each. That was integration hell.

What MCP Does ​

  • MCP Server: a connector that exposes data in a standard format
  • MCP Client: your agent, which plugs into the server

When to Use RAG vs MCP ​

  • Use RAG for large, mostly static knowledge bases (manuals, policies, wikis)
  • Use MCP for live systems and tools (databases, filesystems, APIs)

Use Cases: RAG vs MCP ​

RAG Use Cases (Static Knowledge)

  • Employee handbook Q&A
  • Product manuals and troubleshooting
  • Internal wiki and SOP lookup
  • Compliance policy search
  • Research paper summarization
  • Customer support knowledge base

MCP Use Cases (Live Systems)

  • Read and summarize files from a shared drive
  • Query a database for real-time metrics
  • Pull issues and PRs from GitHub
  • Fetch tickets from a helpdesk system
  • Read Slack or Teams threads for context
  • Update CRM notes or create tasks

4. Project: PDF Chat Agent (LangChain) ​

We will build a RAG agent that answers questions about a PDF.

Install Dependencies ​

bash
pip install langchain-community langchain-openai faiss-cpu pypdf

Flow Overview ​

Code: PDF Chat Agent ​

python
from langchain_community.document_loaders import PyPDFLoader
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import FAISS
from langchain.chains import RetrievalQA
from langchain_openai import ChatOpenAI

# 1. Load the PDF
loader = PyPDFLoader("policy.pdf")
documents = loader.load()

# 2. Chunk the text
splitter = RecursiveCharacterTextSplitter(
    chunk_size=1000,
    chunk_overlap=100
)
chunks = splitter.split_documents(documents)

# 3. Embed the chunks
embeddings = OpenAIEmbeddings()

# 4. Store in FAISS (local vector DB)
vectorstore = FAISS.from_documents(chunks, embeddings)

# 5. Create the retriever
retriever = vectorstore.as_retriever()

# 6. Connect to LLM
llm = ChatOpenAI(temperature=0)
qa_chain = RetrievalQA.from_chain_type(llm, retriever=retriever)

# 7. Ask a question
response = qa_chain.run("What is the vacation policy in this document?")
print(response)
javascript
import fs from "fs";
import path from "path";
import OpenAI from "openai";
import { PDFExtract } from "pdf.js-extract";

// npm install openai pdf.js-extract

const openai = new OpenAI();

// 1. Load and extract text from PDF
async function loadPDF(filePath) {
  const pdfExtract = new PDFExtract();
  const data = await pdfExtract.extract(filePath, {});
  return data.pages
    .map((p) => p.content.map((c) => c.str).join(" "))
    .join("\n");
}

// 2. Chunk the text
function chunkText(text, chunkSize = 1000, overlap = 100) {
  const chunks = [];
  let start = 0;
  while (start < text.length) {
    chunks.push(text.slice(start, start + chunkSize));
    start += chunkSize - overlap;
  }
  return chunks;
}

// 3 & 4. Embed and store chunks in memory
async function embedChunks(chunks) {
  const response = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: chunks,
  });
  return response.data.map((d) => d.embedding);
}

// 5. Retrieve top-k chunks via cosine similarity
function cosineSimilarity(a, b) {
  const dot = a.reduce((sum, v, i) => sum + v * b[i], 0);
  const normA = Math.sqrt(a.reduce((s, v) => s + v * v, 0));
  const normB = Math.sqrt(b.reduce((s, v) => s + v * v, 0));
  return dot / (normA * normB);
}

async function retrieve(query, chunks, embeddings, k = 3) {
  const qEmbRes = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: [query],
  });
  const qEmb = qEmbRes.data[0].embedding;
  return chunks
    .map((c, i) => ({ chunk: c, score: cosineSimilarity(qEmb, embeddings[i]) }))
    .sort((a, b) => b.score - a.score)
    .slice(0, k)
    .map((r) => r.chunk);
}

// 6 & 7. Answer question using retrieved context
async function askPDF(question, pdfPath) {
  const text = await loadPDF(pdfPath);
  const chunks = chunkText(text);
  const embeddings = await embedChunks(chunks);
  const context = await retrieve(question, chunks, embeddings);

  const response = await openai.chat.completions.create({
    model: "gpt-4o",
    temperature: 0,
    messages: [
      {
        role: "system",
        content: "Answer questions based only on the provided context.",
      },
      {
        role: "user",
        content: `Context:\n${context.join("\n\n")}\n\nQuestion: ${question}`,
      },
    ],
  });
  return response.choices[0].message.content;
}

const answer = await askPDF(
  "What is the vacation policy in this document?",
  "policy.pdf",
);
console.log(answer);

What Just Happened? ​

The agent did not read the whole PDF. It found the most relevant chunk, inserted it into the prompt, and answered based on that evidence.


5. Putting It Together: Memory + RAG Agent (LangChain) ​

Now combine short-term memory with long-term memory.

Code: Memory + RAG ​

python
from langchain_openai import ChatOpenAI
from langchain.memory import ConversationBufferMemory
from langchain.chains import ConversationalRetrievalChain
from langchain_community.vectorstores import FAISS
from langchain_openai import OpenAIEmbeddings
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain_community.document_loaders import PyPDFLoader

# Build vector store (same as before)
loader = PyPDFLoader("policy.pdf")
documents = loader.load()
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=100)
chunks = splitter.split_documents(documents)
vectorstore = FAISS.from_documents(chunks, OpenAIEmbeddings())

# Memory + Retrieval
memory = ConversationBufferMemory(
    memory_key="chat_history",
    return_messages=True
)
llm = ChatOpenAI(temperature=0)

chain = ConversationalRetrievalChain.from_llm(
    llm=llm,
    retriever=vectorstore.as_retriever(),
    memory=memory
)

chain.invoke({"question": "What is the vacation policy?"})
chain.invoke({"question": "How many days does it allow?"})
javascript
import OpenAI from "openai";
import { PDFExtract } from "pdf.js-extract";

// npm install openai pdf.js-extract

const openai = new OpenAI();

async function buildVectorStore(pdfPath) {
  const pdfExtract = new PDFExtract();
  const data = await pdfExtract.extract(pdfPath, {});
  const text = data.pages
    .map((p) => p.content.map((c) => c.str).join(" "))
    .join("\n");

  const chunks = [];
  let start = 0;
  while (start < text.length) {
    chunks.push(text.slice(start, start + 1000));
    start += 900;
  }

  const embRes = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: chunks,
  });
  return { chunks, embeddings: embRes.data.map((d) => d.embedding) };
}

function cosineSim(a, b) {
  const dot = a.reduce((s, v, i) => s + v * b[i], 0);
  return (
    dot /
    (Math.sqrt(a.reduce((s, v) => s + v * v, 0)) *
      Math.sqrt(b.reduce((s, v) => s + v * v, 0)))
  );
}

async function retrieveChunks(query, store, k = 3) {
  const qRes = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: [query],
  });
  const qEmb = qRes.data[0].embedding;
  return store.chunks
    .map((c, i) => ({ c, score: cosineSim(qEmb, store.embeddings[i]) }))
    .sort((a, b) => b.score - a.score)
    .slice(0, k)
    .map((r) => r.c);
}

// Memory + RAG chat
const store = await buildVectorStore("policy.pdf");
const chatHistory = [];

async function chat(question) {
  const context = await retrieveChunks(question, store);
  const messages = [
    {
      role: "system",
      content:
        "You are a helpful assistant. Answer based on the provided context and conversation history.",
    },
    ...chatHistory,
    {
      role: "user",
      content: `Context:\n${context.join("\n\n")}\n\nQuestion: ${question}`,
    },
  ];
  const res = await openai.chat.completions.create({
    model: "gpt-4o",
    temperature: 0,
    messages,
  });
  const reply = res.choices[0].message.content;
  chatHistory.push({ role: "user", content: question });
  chatHistory.push({ role: "assistant", content: reply });
  return reply;
}

console.log(await chat("What is the vacation policy?"));
console.log(await chat("How many days does it allow?"));

Common Pitfalls ​

  • Dumping too much history into the prompt
  • Forgetting to summarize old turns
  • Building RAG without citing sources
  • Using memory for factual data instead of retrieval

Checklist ​

  • My agent keeps short-term context clean
  • My agent summarizes or trims old history
  • My agent retrieves documents before answering
  • My responses are grounded in sources

What Comes Next ​

In Chapter 5, you will build multi-agent workflows and learn how to coordinate agents for larger tasks.

Released under the MIT License.