Questions
19 of 32
1What is a Message in LangChain and how does it differ from a plain string prompt?
2What are the core message types in LangChain — HumanMessage, AIMessage, SystemMessage, ToolMessage, FunctionMessage — and when do you use each?
3What is the difference between SystemMessage and HumanMessage — how does the LLM treat them differently under the hood?
4What is a BaseMessage and why does LangChain model all messages as objects instead of raw strings?
5What is the content field in a message and why can it be either a string or an array of content blocks?
6What is a multimodal message and how do you pass images or file data inside a message content block?
7How do you construct a conversation history as a BaseMessage[] array and pass it correctly to a Chat Model?
8What is MessagePlaceholder in a ChatPromptTemplate and how does it let you inject dynamic message history into a prompt?
9How does AIMessage carry tool call requests and how does ToolMessage carry the result back — walk through the full round trip?
10What is the difference between AIMessage.tool_calls and AIMessage.additional_kwargs.function_call — why do both exist?
11How do you trim messages to stay within the LLM's context window without losing important conversation context?
12How do you filter messages by type (e.g. only keep HumanMessages) using LangChain's built-in message utilities?
13What is mergeMessageRuns() and when would you use it to preprocess a message list?
14How do you convert LangChain messages to OpenAI's raw API format and back — when would you need to do this?
15How does LangGraph's MessagesAnnotation work and why is it the recommended state shape for agent graphs?
16How does the messages reducer in LangGraph handle message updates — why can you append, replace, or delete messages by ID?
17How do you implement message deduplication in a LangGraph state to avoid the same message being added twice?
18How do you implement a sliding window memory — keeping only the last N messages — without losing the system prompt?
19How do you implement conversation summarization — replacing old messages with a summary message to save tokens?
20How do you persist and restore a full message history across sessions using a checkpoint saver in LangGraph?
21What is the difference between storing messages in MemorySaver vs an external store like Redis or PostgreSQL via a custom BaseCheckpointSaver?
22How do you stream individual message chunks using AIMessageChunk and how do you aggregate them into a complete AIMessage?
23What is RemoveMessage in LangGraph and how do you use it to surgically delete specific messages from agent state?
24How do you attach custom metadata to a message (e.g. timestamps, user IDs, trace IDs) without breaking LLM compatibility?
25How do you handle token counting per message — accounting for role overhead, tool schemas, and system prompt tokens — to accurately predict context usage?
26How do you design a multi-tenant message store where conversation histories are isolated per user and per session?
27How do you use LangSmith to inspect the exact message array sent to the LLM at every step of an agent run?
28What are the security implications of injecting user-supplied content directly into a SystemMessage — how do you prevent prompt injection attacks?
29Your agent is hitting the context window limit after 20 turns — what is your strategy to manage message history without losing critical context?
30A user's message contains both text and an image — how do you construct the correct multimodal HumanMessage content block for a vision model?
31You need to replay a past conversation from a database and continue it — how do you reconstruct the message state correctly in LangGraph?
32Your LLM is returning inconsistent tool call formatting across providers (OpenAI vs Anthropic vs Gemini) — how do LangChain messages abstract this away?
19 / 32

How do you implement conversation summarization — replacing old messages with a summary message to save tokens?

Implement conversation summarization using a node that triggers when the message count exceeds a threshold, compresses older messages into a summary using an LLM, and replaces them with a single SystemMessage containing the summary.

Summarization is a common technique for managing long conversations without losing critical context. You can add a conditional edge in your LangGraph that checks if the number of messages has exceeded a limit (e.g., 20 messages). If so, a summarization node is invoked. This node takes all messages except the most recent few (e.g., last 5), sends them to an LLM with a summarization prompt, and replaces those messages with a new SystemMessage containing the summary. The remaining recent messages are kept intact. This reduces token usage while preserving high-level context.

Summarization Node Example
Difficulty: 6/10
Topics: conversation summarization, token budgeting, LangChain memory

Scenario Questions

0-2 years experience
  1. 1

    We have a simple chatbot built with LangChain that stores every user and assistant message. How would you add a step that collapses the first five messages into a single summary to keep the token count low?

  2. 2

    If you forget to update the memory after summarizing, what would you see in the next model call?

2-5 years experience
  1. 1

    Imagine the conversation is running in production and you notice occasional token‑limit errors after long sessions. Walk me through how you’d modify the LangChain pipeline to automatically summarize older turns, and what parameters you’d expose for tuning.

  2. 2

    During testing the summarizer sometimes drops important user intent. How would you debug that and improve the summarization prompt?

5-8 years experience
  1. 1

    Design a scalable summarization component for a multi‑tenant chatbot service using LangChain. Explain how you’d handle per‑user token budgets, concurrency, and fallback when the LLM summarizer fails.

  2. 2

    What are the performance implications of summarizing on every turn versus on a token‑threshold trigger, and how would you benchmark the trade‑off?

8+ years experience
  1. 1

    Our product roadmap includes migrating from a single‑LLM summarizer to a hybrid approach that uses a cheap extractor model first. How would you architect this change in a LangChain codebase while keeping backward compatibility for existing customers?

  2. 2

    Discuss the long‑term maintenance considerations of storing summarized conversation history in a database versus keeping it only in memory. How does this affect data privacy, auditability, and token budgeting?

Follow-up Questions

  • How would you measure the token savings versus added latency?
  • What would you do if the generated summary itself approaches the token limit?
  • Can you walk me through how you’d expose a configuration to tune the summarization frequency?