These 60 LangChain interview questions reflect how LangChain interviews run in 2026, now that the framework is on its 1.x line and agents are built with create_agent plus middleware rather than the old chain classes. Interviewers rarely ask you to recite what a chain is any more; they want to know whether you can compose runnables, force structured output, wire tools safely, ground answers with retrieval, trace everything in LangSmith and say clearly when LangGraph is the better tool. Each answer starts with the direct answer and then the trade-offs a senior engineer would mention. If your role leans towards stateful orchestration, pair this page with our LangGraph interview questions and answers, because most 2026 loops ask about both.
How to use this guide
- Freshers and early-career engineers: focus on fundamentals, models and prompts, LCEL and structured output. Interviewers check that you can read a short LangChain snippet and explain what each piece does, not that you have memorised class names.
- Developers with 2 to 6 years of experience: expect RAG, tool calling, agents, memory and LangSmith questions, plus at least one scenario. You are tested on judgement: chain or agent, which retriever, how you would know quality dropped.
- Senior and architect roles: expect LangChain vs LangGraph, production concerns, migration from legacy code and the scenario section, often on a whiteboard. Interviewers listen for failure modes, approvals, security boundaries and evaluation ownership.
- Code samples target LangChain 1.x on Python 3.10 or newer and are deliberately short. Provider model IDs are read from configuration rather than hardcoded, because model names change faster than interview guides. These are high-value, commonly asked questions, not a script.
Contents
- LangChain fundamentals (Q1 to Q5)
- Models and prompts (Q6 to Q10)
- LCEL and runnables (Q11 to Q15)
- Structured output and tools (Q16 to Q20)
- RAG with LangChain (Q21 to Q26)
- Agents in LangChain 1.x (Q27 to Q32)
- Memory and history patterns (Q33 to Q35)
- LangChain vs LangGraph (Q36 to Q38)
- LangSmith tracing and evaluation (Q39 to Q41)
- Production concerns (Q42 to Q45)
- Migration from legacy LangChain (Q46 to Q48)
- Real-world scenarios (Q49 to Q60)
- Key takeaways
- Interview preparation checklist
- FAQ
LangChain fundamentals
1. What is LangChain, and when would you use it instead of calling a model API directly?
Answer: LangChain is an open-source framework for building LLM applications and agents. It gives you a standard interface over many model providers, plus prompt templates, message types, tool calling, structured output, retrievers, vector store integrations and a prebuilt agent loop. Use it when you need to compose several steps, swap providers without rewriting code, call tools, ground answers in your own data, or run an agent with guardrails. For a single prompt to a single provider, the provider SDK alone is often simpler, and saying so in an interview is a sign of maturity rather than a weakness.
Interview tip: Mention that LangChain 1.x is deliberately smaller than older versions: its core job is now the agent harness and the integrations, with low-level orchestration handled by LangGraph underneath.
2. How is the LangChain ecosystem split into packages in 2026?
Answer: langchain-core holds the base abstractions: messages, prompts, runnables, tools, output parsers, documents, embeddings and vector store interfaces. langchain (1.x) is the slimmed-down main package with create_agent, middleware, init_chat_model and re-exports of messages and tools. Provider integrations ship as separate partner packages such as langchain-openai, langchain-anthropic, langchain-aws and langchain-google-genai. langchain-text-splitters holds chunking utilities. langchain-classic contains legacy chains, older retrievers, the indexing API and the hub module. langchain-community was sunset by the LangChain team in 2026 after a long period of accepting no new integrations; new work should use a dedicated integration package or, increasingly, MCP tools.
Real-world example: A requirements file that pins langchain-core, langchain, one partner package and langchain-text-splitters is easy to audit. One that pulls in the whole community package drags in a long tail of optional dependencies your security team will ask about.
3. What is the difference between a chain and an agent?
Answer: A chain is a fixed sequence of steps you define in code: prompt, model, parser, maybe a retriever. The control flow never changes at runtime. An agent lets the model decide, in a loop, which tool to call next and when to stop. Chains are predictable, cheaper and easier to test; agents handle open-ended tasks but need step limits, approvals, and evaluation of the path they took, not just the final answer. In 1.x, chains are built with LCEL and agents with create_agent.
Interview tip: Say you default to a chain and promote to an agent only when the task genuinely requires choosing among tools. Interviewers like engineers who do not reach for autonomy by reflex.
4. What are the core building blocks you combine in a typical LangChain application?
Answer: Chat models (the LLM behind a standard interface), messages (system, human, AI and tool messages), prompt templates, output parsers or structured output schemas, tools, retrievers backed by a vector store, and either an LCEL pipeline or an agent created with create_agent. Around these you add LangSmith for tracing and evaluation and, for agents, a checkpointer for persistence. Everything that can be invoked implements the Runnable interface, which is why they compose cleanly.
5. What does "provider-agnostic" really mean in LangChain, and where does it break down?
Answer: It means the same invoke, stream, bind_tools and with_structured_output calls work across providers, and 1.x adds standard content blocks (message.content_blocks) so text, reasoning, citations and tool calls are read the same way regardless of vendor. It breaks down at the edges: providers differ in native structured output support, context window sizes, tool-call quirks, server-side tools, rate limits and pricing. Swapping a provider is a configuration change plus a re-run of your evaluation set, never only a configuration change.
Models and prompts
6. How do you initialise a chat model in LangChain 1.x?
Answer: Either import the provider class from its partner package (for example ChatOpenAI from langchain_openai) or use init_chat_model, which takes a "provider:model" string and returns the right class. create_agent also accepts that string directly. Keep the model identifier in configuration so you can change it per environment without a code change.
import os
from langchain.chat_models import init_chat_model
# CHAT_MODEL looks like "provider:model-name"
model = init_chat_model(
os.environ["CHAT_MODEL"], temperature=0
)
reply = model.invoke("Summarise our VPN policy.")
print(reply.text)
Interview tip: Note that .text is a property in 1.x; the old .text() method form still works but emits a deprecation warning.
7. What message types exist, and why does the distinction matter?
Answer: SystemMessage sets behaviour and rules, HumanMessage carries user input, AIMessage carries model output (including any tool_calls and usage metadata), and ToolMessage returns a tool result tied to a specific tool_call_id. The distinction matters because providers treat roles differently for safety and instruction-following, because tool results must be matched to the call that requested them, and because trimming or summarising history must never orphan a tool message from its AI message.
8. How do ChatPromptTemplate and MessagesPlaceholder work?
Answer: ChatPromptTemplate builds a list of messages from role and template pairs with named variables. MessagesPlaceholder reserves a slot where a whole list of messages, usually conversation history, is injected at runtime. Templates keep instructions in version-controlled code and make variables explicit, which is what lets you test and trace prompts.
from langchain_core.prompts import (
ChatPromptTemplate, MessagesPlaceholder)
prompt = ChatPromptTemplate.from_messages([
("system", "You are a concise IT helpdesk assistant."),
MessagesPlaceholder("history", optional=True),
("human", "{question}"),
])
9. How do you keep prompts maintainable across a team?
Answer: Store templates in code or a prompt registry with versioning, keep the system prompt short and specific, separate stable instructions from per-request context, and tie every prompt version to an evaluation run. LangSmith lets you manage and compare prompt versions alongside traces; the older hub module now lives in langchain-classic. Review prompt changes like code changes, because a one-line wording change can shift tool selection or refusal behaviour.
Interview tip: For deeper prompt design questions, our prompt engineering interview questions cover few-shot design, instruction hierarchy and prompt injection in more depth.
10. What are standard content blocks and why were they added?
Answer: Modern models return more than plain text: reasoning traces, citations, server-side tool calls, images. Each provider encodes these differently. Standard content blocks, exposed through message.content_blocks, normalise them into typed blocks such as text, reasoning and tool call, so your application logic does not branch per provider. Practically, it means a citation renderer or a reasoning logger works the same whether the backend is a hosted API, a cloud platform model or a local model.
LCEL and runnables
11. What is LCEL, and what do you get from it?
Answer: LCEL (LangChain Expression Language) is the declarative way to compose runnables with the pipe operator: prompt | model | parser. Because every component implements the Runnable interface, the composed chain automatically supports invoke, batch, stream and their async variants, plus retries, fallbacks, configuration and tracing. It replaced the subclass-based chains such as LLMChain.
from langchain_core.output_parsers import StrOutputParser
chain = prompt | model | StrOutputParser()
answer = chain.invoke({"question": "Reset my VPN?"})
answers = chain.batch([{"question": "a"},
{"question": "b"}])
12. Explain RunnableParallel, RunnablePassthrough and RunnableLambda.
Answer: RunnableParallel runs several runnables on the same input and returns a dict of their outputs; a plain dict in a chain is coerced into one. RunnablePassthrough passes the input through unchanged, often to carry the original question next to retrieved context; its assign method adds new keys while keeping existing ones. RunnableLambda wraps an ordinary Python function so it can sit in a pipeline, which is how you insert formatting, validation or redaction steps.
Interview tip: Keep lambdas small and pure. A lambda that calls a database is a hidden side effect that will not retry or trace the way you expect; make it a tool or a proper runnable instead.
13. How does streaming work in LCEL, and what is astream_events for?
Answer: Calling stream or astream on a chain yields output chunks as soon as the final step produces them, provided each step can process input incrementally. Parsers like StrOutputParser stream; a step that needs the whole input blocks streaming until it finishes. astream_events exposes intermediate events from every step (model tokens, retriever results, tool start and end), which is what you use to show "searching policies..." in a UI. For agents, 1.x renamed the streamed model node from agent to model, a detail that breaks old event filters.
14. How do you add retries and fallbacks to a chain?
Answer: Every runnable has with_retry (retry on chosen exception types with backoff) and with_fallbacks (try alternative runnables if the primary fails). A common pattern is a primary model with a fallback to a second provider or a smaller model, wrapped with retries for rate-limit errors only. Do not retry on validation errors blindly; that just repeats the same mistake and costs tokens.
primary = model.with_retry(stop_after_attempt=3)
safe_model = primary.with_fallbacks([backup_model])
chain = prompt | safe_model | StrOutputParser()
15. What is RunnableConfig used for?
Answer: A config dict passed at invoke time carries cross-cutting settings without changing chain code: tags and metadata for tracing, callbacks, max_concurrency for batch calls, run_name, and configurable values such as a thread_id for agent persistence. Attaching a tenant ID or request ID as metadata is the cheapest way to make production traces searchable.
Structured output and tools
16. How do you get reliable structured output from a model?
Answer: Call with_structured_output on a chat model with a Pydantic model, TypedDict or JSON schema. LangChain uses the provider's native structured output or tool calling to enforce the schema and returns a validated object. Prefer this to asking for JSON in the prompt and parsing it, which fails silently on malformed output. Use include_raw=True when you need the raw message for debugging or token accounting.
from pydantic import BaseModel, Field
class Ticket(BaseModel):
category: str = Field(
description="network, access or hardware")
priority: int = Field(ge=1, le=4)
summary: str
triage = model.with_structured_output(Ticket)
t = triage.invoke("VPN drops every 10 min, urgent")
Real-world example: An insurer extracting claim type, policy number and incident date from emails should still validate business rules after parsing; a schema enforces shape, not truth. Our guide to function calling and structured outputs covers the validation layer.
17. How do you define a tool, and what makes a tool definition good?
Answer: Decorate a typed Python function with @tool; LangChain derives the name, the argument schema from type hints and the description from the docstring. A good tool has a specific name, a docstring that says when to use it and what the inputs look like, narrow typed arguments, and a result that is short and useful to the model. The model chooses tools almost entirely from names and descriptions, so vague descriptions are the most common cause of wrong tool selection.
from langchain.tools import tool
@tool
def get_ticket(ticket_id: str) -> str:
"""Fetch an incident by ID, e.g. INC0012345."""
return itsm_client.get(ticket_id).summary()
18. Walk through the tool-calling loop when you use bind_tools directly.
Answer: You call model.bind_tools([...]), invoke it, and inspect response.tool_calls. For each call you run the matching tool with the given arguments and append a ToolMessage with the same tool_call_id. You then invoke the model again with the full message list, repeating until it returns a final answer without tool calls. The model never executes anything itself; your code does. create_agent runs this loop for you, which is why hand-rolled loops are mostly an interview exercise now.
19. How do you stop a model from passing dangerous arguments to a tool?
Answer: Treat model-produced arguments as untrusted input. Validate them with the schema plus business rules inside the tool, enforce authorisation with the end user's identity rather than a shared super-user account, scope tools to the minimum operation (read-only where possible), and require human approval for writes with real-world impact. Values such as the tenant ID or user ID should be injected from runtime context, not chosen by the model; LangChain supports injected arguments and runtime context so the model never sees or sets them.
Interview tip: Link this to prompt injection: a malicious document retrieved by RAG can instruct the model to call a tool. Approval gates and least-privilege tools limit the blast radius.
20. How do MCP tools fit into a LangChain application?
Answer: MCP (Model Context Protocol) is an open protocol, introduced by Anthropic in late 2024, for connecting AI applications to tools and data. The langchain-mcp-adapters package loads tools exposed by MCP servers and converts them into LangChain tools, so a create_agent agent can use a ServiceNow, GitHub or database MCP server alongside local tools. The benefit is reuse across clients; the cost is that you inherit the server's authentication model and must still apply approvals and allow-lists. See how to build an MCP server in Python for the server side.
RAG with LangChain
21. How does retrieval-augmented generation work in LangChain, end to end?
Answer: At ingestion time, document loaders read sources into Document objects, text splitters chunk them, an embeddings model converts chunks to vectors and a vector store saves them with metadata. At query time a retriever embeds the question, fetches the top matching chunks, optionally reranks them, and the chain passes them as context to the model with instructions to answer only from that context and cite sources. For the concept itself, see what RAG is and why it works.
ingest: load -> split -> embed -> vector store
query: question -> retriever -> top-k chunks
-> (rerank) -> prompt + context
-> model -> answer with citations
22. Write a minimal RAG chain with LCEL.
Answer: Retrieve, format the documents with their source metadata, and pass both context and question into the prompt. Keeping the source in the formatted context is what makes citations possible.
from langchain_core.runnables import (
RunnableParallel, RunnablePassthrough)
retriever = store.as_retriever(search_kwargs={"k": 4})
def format_docs(docs):
return "\n\n".join(
f"[{d.metadata['source']}] {d.page_content}"
for d in docs)
rag_chain = (
RunnableParallel(
context=retriever | format_docs,
question=RunnablePassthrough(),
)
| rag_prompt
| model
| StrOutputParser()
)
Interview tip: Mention that the older RetrievalQA and create_retrieval_chain helpers now live in langchain-classic; new code uses LCEL like this or a retrieval tool inside an agent.
23. How do you choose a text splitter and chunk size?
Answer: RecursiveCharacterTextSplitter from langchain-text-splitters is the usual default: it splits on paragraphs, then sentences, then words, to keep chunks coherent. Use structure-aware splitters for Markdown, HTML or code, and token-based splitting when you need to respect a model's limits precisely. Chunk size is a trade-off between precision (small chunks) and enough context to answer (larger chunks); overlap reduces the chance a fact is cut in half. Decide by measuring retrieval quality on real questions, not by copying a default. Our RAG chunking strategies guide compares approaches.
24. Document loaders: what goes wrong with them in enterprise projects?
Answer: Loaders are the easy part to write and the hard part to get right. Scanned PDFs need OCR, tables lose structure when flattened to text, headers and footers repeat on every page, and SharePoint or Confluence exports carry permissions that a naive loader discards. The loader must preserve metadata you will later need: source URL, page, section, owner, access groups and last-modified date. Many loaders lived in the community package, so in 2026 teams increasingly use a dedicated integration package or their own parsing code. See parsing PDFs, tables and scans for RAG.
25. Which vector store would you pick, and how does LangChain abstract it?
Answer: LangChain defines a VectorStore interface (add documents, similarity search, metadata filters, as_retriever) implemented by integration packages for stores such as pgvector via langchain-postgres, Chroma, Qdrant, Elasticsearch, OpenSearch, Pinecone and cloud-native options. InMemoryVectorStore in langchain-core is for tests and demos. Choose based on what the organisation already operates, metadata filtering needs, hybrid search support, scale and data residency. A team that already runs PostgreSQL often starts with pgvector because backups, access control and monitoring already exist. More in our vector database interview questions.
26. How do you improve retrieval when plain similarity search is not good enough?
Answer: In order of usual payoff: fix chunking and metadata, add metadata filters (department, product, date), add hybrid search combining keyword and vector scores, add a reranker on the top candidates, then consider query rewriting or multi-query retrieval. Parent-document approaches retrieve a small chunk but return its larger parent for context. Several of the classic retriever helpers (multi-query, ensemble, contextual compression, parent document) are in langchain-classic now; the techniques are still valid, but be ready to implement them with LCEL or your vector store's native features. Measure each change on the same evaluation set.
Agents in LangChain 1.x
27. What is create_agent and what does it replace?
Answer: create_agent, imported from langchain.agents, is the standard way to build an agent in 1.x. It takes a model, tools, a system_prompt, optional middleware, an optional response_format, and an optional checkpointer and store. It returns a compiled LangGraph graph, so you get streaming, persistence and human-in-the-loop support without writing the graph yourself. It replaces create_react_agent from langgraph.prebuilt (note the parameter rename from prompt to system_prompt) and the legacy AgentExecutor, which now lives in langchain-classic.
from langchain.agents import create_agent
from langgraph.checkpoint.memory import InMemorySaver
agent = create_agent(
model=os.environ["CHAT_MODEL"],
tools=[get_ticket, close_ticket],
system_prompt="You are an IT helpdesk agent.",
checkpointer=InMemorySaver(),
)
cfg = {"configurable": {"thread_id": "user-42"}}
result = agent.invoke(
{"messages": [{"role": "user",
"content": "Status of INC0012345?"}]},
cfg)
28. What is middleware in LangChain 1.x?
Answer: Middleware is the extension mechanism of create_agent. It hooks into the agent loop at defined points: before_agent, before_model, after_model, after_agent, and wrappers around each model call (wrap_model_call) and tool call (wrap_tool_call). You can write middleware as a class extending AgentMiddleware or with decorators such as @before_model and @dynamic_prompt. It replaces the old pre-model and post-model hooks and is how you do context engineering: trimming history, changing the prompt per user, filtering tools, redacting data, enforcing limits.
Interview tip: Describe middleware as "policy around the loop". The loop itself (model, tools, repeat) stays simple; everything an enterprise needs is layered on top.
29. Name the built-in middleware you would use in production and why.
Answer: The ones that come up most:
| Middleware | What it does | When to use |
|---|---|---|
HumanInTheLoopMiddleware | Pauses before chosen tool calls for approve, edit, reject or respond decisions | Any write action: refunds, ticket closure, emails |
SummarizationMiddleware | Summarises older messages when a token, message or fraction trigger is hit | Long-running conversations |
PIIMiddleware | Detects email, card numbers, IPs and custom patterns; redact, mask, hash or block | Regulated data, logs and prompts |
ModelCallLimitMiddleware, ToolCallLimitMiddleware | Cap model or tool calls per run or per thread | Runaway loops and cost control |
ModelRetryMiddleware, ToolRetryMiddleware, ModelFallbackMiddleware | Backoff retries and fallback models | Provider outages and flaky APIs |
LLMToolSelectorMiddleware | Uses a model to pick relevant tools before the main call | Agents with many tools |
Check the current documentation for the full list, because the set grows between minor releases.
30. How does human-in-the-loop approval work with create_agent?
Answer: Add HumanInTheLoopMiddleware(interrupt_on={"close_ticket": True}) and a checkpointer. When the model proposes that tool call, the run stops and returns an __interrupt__ describing the requested action and the allowed decisions. Your UI shows it to an approver, then you resume the same thread with Command(resume={"decisions": [{"type": "approve"}]}), or an edit, reject or respond decision. Because state is checkpointed, the approval can arrive minutes or hours later. Our human-in-the-loop AI guide covers approval design.
from langgraph.types import Command
agent.invoke(
Command(resume={"decisions": [{"type": "approve"}]}),
cfg)
31. How do you get structured output from an agent rather than a plain model?
Answer: Pass response_format to create_agent. Passing a schema type lets LangChain choose: ProviderStrategy when the provider supports native structured output, otherwise ToolStrategy, which exposes the schema as a special tool. The validated result appears under result["structured_response"]. Structured output is produced inside the main loop, so no extra model call is needed, and ToolStrategy can feed validation errors back so the model corrects itself (controlled by handle_errors). The prompted-JSON fallback from earlier versions was removed in 1.x.
32. How do you pass per-request context, such as user identity, into an agent and its tools?
Answer: Define a context_schema on the agent and pass context= to invoke or stream. Tools and middleware read it through the runtime (for example a ToolRuntime parameter in the tool), so the model never sees or fabricates it. This replaced the older habit of stuffing values into config["configurable"]. Use it for user ID, tenant, locale, entitlements and the downstream access token your tools need.
Interview tip: This is the answer to "how do you stop user A seeing user B's data through an agent": identity flows through runtime context into tools that enforce authorisation, never through the prompt.
Memory and history patterns
33. What memory options does LangChain offer in 2026?
Answer: The old memory classes (ConversationBufferMemory, window, summary and vector-backed memory) are legacy and live in langchain-classic. The current pattern for agents is short-term memory through a LangGraph checkpointer keyed by thread_id, which persists the message state of a conversation, plus long-term memory through a store for facts that should survive across threads, such as user preferences. Context-window pressure is handled with trimming in a before_model hook or SummarizationMiddleware. The underlying ideas (buffer, window, summary, retrieval) still apply; they are just implemented as state plus middleware. See AI agent memory explained.
34. InMemorySaver works in development. What do you use in production?
Answer: A durable checkpointer such as PostgresSaver from langgraph-checkpoint-postgres, so conversations survive restarts and scale across replicas. Plan retention: threads contain user data, so define deletion on request, time-based expiry and encryption at rest. Index thread IDs by user and tenant so you can enforce access, and never derive a thread ID from something guessable.
35. How do you handle chat history in a simple LCEL chain without an agent?
Answer: Pass prior messages into a MessagesPlaceholder from your own store, trimming with trim_messages by token count. RunnableWithMessageHistory in langchain-core still exists to wrap a chain with a history factory, but for new conversational work the LangChain docs steer you towards an agent or LangGraph graph with a checkpointer, which handles persistence, interrupts and streaming consistently.
LangChain vs LangGraph
36. What is the difference between LangChain and LangGraph?
Answer: LangChain is the higher-level framework: model integrations, prompts, tools, retrievers and the create_agent harness with middleware. LangGraph is the lower-level orchestration runtime: you define explicit state, nodes and edges, with durable execution, checkpointing, interrupts and time travel. create_agent is itself built on LangGraph. LangChain's docs frame it as: use LangChain for a customisable agent harness, LangGraph for advanced needs that combine deterministic and agentic workflows, and Deep Agents (the deepagents package) for a batteries-included harness with planning, a virtual filesystem and subagents. Our LangGraph interview guide goes deep on state, reducers and checkpointing.
| Need | LangChain 1.x | LangGraph |
|---|---|---|
| Single agent with tools | create_agent | Possible, more code |
| Fixed multi-step workflow with branches | LCEL for simple cases | Explicit graph |
| Approvals and persistence | Middleware plus checkpointer | Native interrupts and checkpointers |
| Multi-agent with handoffs | Limited | Supervisor or custom graphs |
| Long-running durable jobs | Not the main focus | Durable execution |
37. What moved from LangChain to LangGraph, and what moved the other way?
Answer: Over the 0.x years, agent execution, memory and persistence shifted to LangGraph: checkpointers, threads, interrupts and the prebuilt ReAct agent all came from there. In 1.x, the recommended agent entry point moved back up into LangChain as create_agent, while still running on LangGraph underneath. So the clean mental model is: LangGraph owns runtime and state; LangChain owns the developer-facing harness and integrations; legacy chain-era code lives in langchain-classic.
38. When would you start with LangChain and later move to LangGraph?
Answer: Start with create_agent when one agent, a handful of tools and some middleware solve the problem. Move to a custom LangGraph graph when you need deterministic steps that must always run (a compliance check before any reply), parallel branches, several specialised agents with explicit handoffs, or a long-running workflow with multiple approval points. Because create_agent returns a graph, you can embed it as a node in a larger LangGraph workflow rather than rewriting it.
Real-world example: Consider a bank's dispute-handling assistant. Version one is a create_agent agent that looks up transactions. Version two must always run a fraud-rule check, branch to a human for amounts above a threshold, and log a regulator-ready audit record; that becomes a LangGraph workflow with the agent as one node. Our LangGraph for enterprise AI article covers that design.
LangSmith tracing and evaluation
39. How do you enable LangSmith tracing, and what do you look at in a trace?
Answer: Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY, optionally LANGSMITH_PROJECT, and LangChain and LangGraph runs are traced automatically. Use tracing_context to control tracing selectively, and add tags and metadata through the run config. In a trace you read the exact prompt sent, retrieved documents, each tool call with arguments and results, latency per step, token usage and errors. Most "the model is wrong" bugs turn out to be retrieval or tool bugs once you open the trace.
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=...
export LANGSMITH_PROJECT=helpdesk-agent-uat
40. How do you run offline evaluation in LangSmith?
Answer: Create a dataset of inputs with reference outputs, write evaluators (functions that receive inputs, outputs and reference outputs and return a score), and run client.evaluate(target, data=..., evaluators=[...]) to produce an experiment you can compare against earlier ones. Evaluators can be deterministic checks (valid JSON, correct tool called, citation present) or LLM-as-judge evaluators, for example from the openevals package. Run it in CI on every prompt, model or retrieval change. For metric choice, see our LLM evaluation interview questions.
41. How do you evaluate an agent, as opposed to a RAG chain?
Answer: For a RAG chain you measure retrieval (did the right chunks come back) and answer quality (faithfulness to context, relevance, citation correctness). For an agent you also evaluate the trajectory: did it call the right tools, in an acceptable order, with valid arguments, without unnecessary steps, and did it stop for approval where required. Final-answer scoring alone misses an agent that got the right answer by calling a forbidden tool. Combine offline datasets with online evaluation on sampled production traces and human feedback. See AI agent evaluation.
Production concerns
42. What are the most common production issues with LangChain applications?
Answer: Ungrounded answers (fix with better retrieval, citations and faithfulness evaluation), latency and cost growth (caching, smaller models for simple steps, fewer tokens in context, streaming), prompt injection through user input or retrieved documents (input checks, least-privilege tools, approvals), non-determinism breaking tests (low temperature, structured output, evaluation thresholds instead of exact matches), dependency drift across fast-moving packages (pin versions, run evaluations on upgrade) and runaway agent loops (call limits). The pattern behind all of them: you cannot fix what you do not trace.
43. How do you control cost and latency?
Answer: Measure first in LangSmith by step. Then: route simple classification or extraction to a smaller model; cap retrieved context; cache deterministic prompt results with an LLM cache (for example set_llm_cache) and use provider prompt caching for long stable prefixes; stream to improve perceived latency; run independent steps in parallel with RunnableParallel; limit agent iterations with call-limit middleware; and use InMemoryRateLimiter or a gateway to stay inside provider quotas. Our LLM latency optimisation guide goes further.
44. How do you secure a LangChain agent that touches enterprise systems?
Answer: Keep secrets out of prompts and code (use a secrets manager), give each tool the narrowest credential and act on behalf of the end user where the downstream system supports it, enforce authorisation inside tools, redact PII with middleware before it reaches the model or logs, require approvals on write actions, cap calls, and keep an audit trail linking user, thread, tool call and approver. Treat retrieved content as untrusted data, never as instructions. Red-team the agent with injection attempts before go-live.
45. How do you deploy and version a LangChain service?
Answer: Wrap the chain or agent in an API service (FastAPI is common), containerise it, and deploy it like any other service with health checks, autoscaling and secrets from the platform. Pin langchain-core, langchain and partner package versions, and treat prompts, model IDs, retrieval settings and evaluation datasets as versioned artefacts. Promote a change only when the evaluation suite passes in CI. Use a durable checkpointer and external vector store so replicas stay stateless. LangChain also offers a managed deployment option for LangGraph-based agents; check current documentation for its features and naming.
Migration from legacy LangChain
46. You inherit code using LLMChain and ConversationChain. How do you migrate it?
Answer: LLMChain, ConversationChain, RetrievalQA and similar classes are legacy; in 1.x they are importable only from langchain-classic. A quick, low-risk step is to switch imports to langchain_classic so the code runs, then migrate each chain: LLMChain becomes prompt | model | StrOutputParser(); ConversationChain becomes an agent or chain with a checkpointer or explicit history; RetrievalQA becomes an LCEL RAG chain or a retrieval tool in an agent. Run the evaluation set before and after each change.
# legacy (langchain-classic)
# chain = LLMChain(llm=llm, prompt=prompt)
# chain.run(question="...")
# 1.x
chain = prompt | model | StrOutputParser()
chain.invoke({"question": "..."})
47. What are the breaking changes to watch when moving to LangChain 1.x?
Answer: Python 3.10 or newer for all packages; legacy chains, older retrievers, the indexing API and hub moved to langchain-classic; create_react_agent replaced by create_agent with prompt renamed to system_prompt; pre and post model hooks replaced by middleware; custom agent state must be a TypedDict; runtime context passed via context=; .text() became the .text property; prompted structured output removed in favour of ToolStrategy or ProviderStrategy; and the streamed node name changed from agent to model.
48. How would you migrate an AgentExecutor-based agent safely?
Answer: Inventory the tools, prompt, memory and any custom output parsing. Rebuild it with create_agent, moving the prompt to system_prompt, memory to a checkpointer, iteration limits to call-limit middleware, and error handling to wrap_tool_call or tool retry middleware. Capture a set of real conversations from production logs, replay them against both versions in LangSmith and compare tool trajectories and answers. Roll out behind a feature flag to a small group before switching everyone.
Real-world scenarios
49. Your RAG chatbot for HR policies gives confident answers that are wrong. What do you do?
Answer: Separate retrieval failure from generation failure before touching the prompt. Open failing traces, check whether the right policy chunk was retrieved, and only then look at how the model used it.
What I would check:
- Whether the correct document exists in the index and is the current version, not a superseded policy.
- Whether chunking split the relevant clause or lost table structure.
- Whether retrieval returns it in the top results; if not, add metadata filters, hybrid search or a reranker.
- Whether the prompt tells the model to answer only from context, cite sources and say "not found" otherwise.
- Faithfulness scores on an evaluation set built from real employee questions.
Production consideration: Add a "document last updated" field to citations and a scheduled re-index job; stale content is the most common silent cause of wrong answers.
50. An agent is stuck in a loop, calling the same tool repeatedly. How do you fix it?
Answer: Stop the bleeding with call limits, then fix the cause, which is usually a tool result the model cannot interpret.
What I would check:
- The trace: does the tool return an error, an empty result or something ambiguous that invites a retry?
- Whether the tool description makes clear what a "not found" result means.
- Whether errors are returned as clear messages to the model rather than swallowed.
- Whether two tools overlap so the model alternates between them.
Production consideration: Keep ModelCallLimitMiddleware and ToolCallLimitMiddleware on every agent permanently, and alert when runs hit the limit.
51. A retailer wants an order-support agent that can issue refunds. How do you design it?
Answer: A create_agent agent with read tools (order lookup, shipment status, return policy retrieval) and one write tool (issue refund), with human approval on refunds above a business-defined threshold and hard validation inside the refund tool.
What I would check:
- Customer identity flows through runtime context, so the agent can only query that customer's orders.
- The refund tool re-validates order status, amount and refund eligibility server-side.
- Approval decisions and approver identity are logged with the thread ID.
- Evaluation covers abusive requests ("refund all my orders") and prompt injection in order notes.
Production consideration: Make the refund tool idempotent with a request key, because retries and resumed threads can otherwise issue a refund twice.
52. Latency for a document Q&A endpoint is too high for users. Where do you start?
Answer: Break down latency per step in LangSmith, then attack the largest contributor.
What I would check:
- Time to first token versus total time; enable streaming if the UI waits for the full answer.
- Retrieval latency and whether the vector store needs an index or is under-provisioned.
- How many chunks go into context; fewer, better chunks after reranking cut model time.
- Sequential steps that could run in parallel, such as query rewriting and a metadata lookup.
- Whether a smaller model passes the evaluation set for this task.
Production consideration: Track latency percentiles, not averages, and set a budget per step so regressions are visible in CI.
53. Consider a hospital that wants a discharge-summary assistant. What changes because the data is clinical?
Answer: The architecture is a structured-output chain over the patient record, not an autonomous agent, and every output is reviewed by a clinician before it is used.
What I would check:
- Data residency and whether the model endpoint is approved for patient data.
- PII handling in traces: redact or restrict access to LangSmith projects containing clinical text.
- Structured output schema for medications and follow-ups, validated against the source record.
- An evaluation set reviewed by clinicians covering omissions, not just wrong statements.
Production consideration: Log which source fields each summary sentence came from, so reviewers can verify quickly and auditors can trace decisions.
54. The agent has forty tools and keeps picking the wrong one. What do you do?
Answer: Reduce what the model sees at each step and sharpen what it does see.
What I would check:
- Overlapping or vaguely named tools that should be merged or renamed.
- Docstrings that state when to use, and when not to use, each tool.
- Whether
LLMToolSelectorMiddlewareor a custom middleware can filter tools by intent or user role. - Whether the problem is really several agents' worth of scope, better split in LangGraph.
Production consideration: Add tool-selection accuracy to the evaluation suite so a new tool cannot silently degrade selection of existing ones.
55. Security review finds customer emails and phone numbers in LangSmith traces. How do you respond?
Answer: Treat it as a data-handling incident: restrict access, clean up, then prevent recurrence.
What I would check:
- Who has access to the affected projects and the retention setting on those traces.
- Whether PII enters at user input, retrieved documents or tool results.
PIIMiddlewareon input, output and tool results with redact or mask strategies, plus custom detectors for local identifiers.- Masking options for inputs and outputs in the tracing configuration; check current LangSmith documentation for the exact settings.
Production consideration: For Indian deployments, align trace retention and access with your obligations under the Digital Personal Data Protection (DPDP) Act.
56. A provider changed behaviour and answer quality dropped overnight. How do you catch and handle this?
Answer: Pin the model version where the provider allows it, and run scheduled evaluations against production-like datasets so drift shows up as a failing experiment rather than user complaints.
What I would check:
- Experiment comparison in LangSmith between yesterday's and today's runs.
- Online evaluator scores and thumbs-down rates on sampled traces.
- Whether a fallback model passes the evaluation set and can be switched by configuration.
Production consideration: Keep the model ID in configuration and test a second provider periodically; provider-agnostic code is only useful if the alternative is already evaluated.
57. A GCC team in Hyderabad must ship an internal IT helpdesk agent integrated with ServiceNow. Outline your approach.
Answer: Discovery first: which ticket types are high-volume and low-risk, which actions require approval, and what "resolved" means to the service desk. Then a create_agent agent with a knowledge-base retrieval tool, read tools for ticket status, and write tools for creating or updating tickets behind approval middleware, possibly through a ServiceNow MCP server.
user (Teams/portal)
-> API (FastAPI, Entra ID auth)
-> create_agent + middleware
|-> KB retriever (pgvector)
|-> ticket tools / MCP server
`-> approval (HITL)
-> LangSmith traces + evals
What I would check:
- Service account scopes in ServiceNow and whether actions run as the user.
- Evaluation set from past tickets with known resolutions.
- Deflection and escalation metrics agreed with the service desk owner.
Production consideration: Launch with read-only tools to one department, then add write actions once trajectories are clean. Our ServiceNow AI agent project walks through a similar build.
58. An insurer wants claims emails classified and routed. Chain or agent?
Answer: A chain. Classification and extraction with a fixed output schema is a deterministic pipeline: with_structured_output for claim type, urgency, policy number and missing documents, then rule-based routing in code. An agent adds cost and unpredictability without benefit here.
What I would check:
- Confusion between similar claim categories on a labelled sample.
- Validation of extracted policy numbers against the policy system.
- A confidence or "needs review" path for ambiguous emails.
Production consideration: Use batch with max_concurrency for backlogs, and keep a human review queue for low-confidence items.
59. A long-running support conversation starts failing with context-length errors. What is your fix?
Answer: Manage context explicitly instead of sending the full history every turn.
What I would check:
- Token growth per turn in traces, especially large tool results being kept in history.
SummarizationMiddlewarewith a token or fraction trigger and a sensible number of recent messages kept.- Trimming that never separates a tool call from its tool message.
- Moving durable facts (customer tier, open case ID) to long-term memory or runtime context rather than history.
Production consideration: Summaries can drop critical details; evaluate long conversations specifically, and keep the summary prompt focused on facts the task needs.
60. Your team has a working LangChain demo. What must happen before it goes to production?
Answer: Move from demo to engineered system: evaluation, security, observability and ownership.
What I would check:
- An evaluation dataset from real users, with agreed thresholds, running in CI.
- LangSmith tracing with tenant and request metadata, plus PII controls.
- Authentication, per-user authorisation in tools, approvals on writes and call limits.
- Pinned dependencies, durable checkpointer and vector store, containerised deployment.
- Cost per conversation estimated and monitored; fallback model evaluated.
- A named owner for the knowledge base and for reviewing failing traces.
Production consideration: Most failed launches skip the evaluation and ownership items, not the code. Our article on why AI demos fail in enterprise production lists the usual gaps.
Want to build these patterns hands-on rather than only read about them? Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers LangChain, LangGraph, RAG, evaluation and deployment with labs, in our Ameerpet classroom or live online.
Key takeaways
- LangChain 1.x is a smaller framework:
create_agent, middleware, standard messages and content blocks, and integrations, running on LangGraph underneath. - Use LCEL for fixed pipelines and
create_agentonly when the model must choose tools; default to the simpler option. - Structured output (
with_structured_outputorresponse_format) and well-described, least-privilege tools are what make LLM output usable by other systems. - Legacy chains, memory classes, older retrievers and
AgentExecutorlive inlangchain-classic;langchain-communityhas been sunset, so prefer dedicated integration packages. - Memory is now state: a checkpointer with thread IDs for conversations, a store for long-term facts, and middleware for trimming and summarisation.
- LangSmith tracing and evaluation are part of the build, not an afterthought; evaluate agent trajectories, not only final answers.
- Reach for LangGraph when you need deterministic steps, multiple agents, parallel branches or long-running durable workflows.
Interview preparation checklist
- Install LangChain 1.x in a fresh virtual environment and know which package each import comes from.
- Write an LCEL chain from memory: prompt, model, parser, with
invoke,batchandstream. - Build a small RAG chain with a splitter, a vector store and citations, and explain your chunking choice.
- Build a
create_agentagent with two tools, a checkpointer, a call limit and human approval on one tool. - Return a Pydantic object from both a model (
with_structured_output) and an agent (response_format). - Turn on LangSmith tracing, read a trace end to end, and run one offline evaluation with a dataset.
- Migrate one legacy snippet (
LLMChainorRetrievalQA) to 1.x and explain the change. - Prepare a two-minute explanation of LangChain vs LangGraph with a concrete example.
- Rehearse two scenario answers aloud: a wrong-answer RAG bot and an agent with a write action.
- Be ready to discuss security: identity in runtime context, PII redaction, prompt injection and approvals.
FAQ
Is LangChain still worth learning in 2026?
Yes. LangChain remains a widely used framework for LLM applications and agents, and its 1.x release made it simpler. Even teams that use other frameworks expect engineers to understand the patterns it popularised: runnables, tool calling, retrieval and agent loops.
Should I learn LangChain or LangGraph first?
Start with LangChain for models, prompts, tools, RAG and create_agent, then learn LangGraph for explicit state, branching workflows and multi-agent systems. Since create_agent runs on LangGraph, the second step builds directly on the first.
Do I need to know legacy chains like LLMChain for interviews?
You should recognise them and know they now live in langchain-classic, because many companies still maintain older code. Interviewers care more that you can migrate them to LCEL or create_agent than that you remember their parameters.
What Python skills do LangChain interviews expect?
Comfort with type hints, Pydantic models, async and await, decorators, virtual environments and dependency pinning, plus building a small FastAPI service. Most LangChain bugs in interviews are really Python or data-shape bugs.
How important is LangSmith in LangChain interviews?
Increasingly important for mid-level and senior roles. Expect questions on reading traces, building evaluation datasets and comparing experiments, because that is how teams prove an LLM application works.
Which projects should I build to prepare for a LangChain interview?
A RAG assistant over real documents with citations and an evaluation set, and a tool-calling agent that performs at least one approved write action against a mock enterprise system. Deploy one of them and keep the traces to discuss.
Is LangChain used in Java or only Python?
LangChain's official libraries are Python and JavaScript or TypeScript. Java teams commonly use LangChain4j, a separate community project, or Spring AI, which follow similar concepts.
How long does it take to prepare for a LangChain interview?
It depends on your Python and LLM background. Developers who already build APIs usually need focused hands-on practice on RAG, agents and evaluation rather than more reading; freshers should first get comfortable with Python and LLM fundamentals.
Do I need cloud skills for LangChain roles?
For most production roles, yes. You will be asked how you deploy, secure and observe the service, and how you call models through platforms such as Amazon Bedrock, Azure OpenAI or Google Cloud.
If you want to take LangChain agents from a notebook to an enterprise deployment, with integrations, approvals, observability and a simulated customer engagement, Cloudsoft's FDE PRO Forward Deployed Engineer program builds exactly that over 12 weeks, with placement support until you're placed. For a broader foundation across AI, ML, cloud and security, explore the APEX program. Classroom in Ameerpet or live online; book a free demo on +91 96660 19191.



