New batches starting this week Β· Limited seats

LangGraph Interview Questions and Answers 2026 (60 Questions)

60 LangGraph interview questions with practical answers, from StateGraph and reducers to checkpointers, human-in-the-loop interrupts, memory, multi-agent design, LangSmith deployment and real production scenarios.

LangGraph interview questions 2026: 60 questions on state, routing, checkpointers, interrupts, memory and deployment
Last updated Β· 44 min read Β· 9,708 words

These LangGraph interview questions cover what hiring panels probe in 2026: how you model state, route control flow, persist and resume runs, pause for human approval, and ship a graph to production with tracing. LangGraph is a low-level orchestration framework and runtime for long-running, stateful agents, and interviewers want to see that you can reason about its state, checkpoints and interrupts, not just call an agent helper. The 60 questions below run from fundamentals to production scenarios, with answers checked against the LangGraph 1.x documentation.

How to use this guide

  • Freshers and juniors are usually tested on the fundamentals and state-design sections: what a node, edge and reducer are, and why a graph beats a plain chain for agent loops.
  • Mid-level engineers get control flow, persistence, human-in-the-loop and memory: Command, Send, checkpointers, threads and interrupt().
  • Senior and architect roles spend most of the interview on multi-agent design, testing, deployment, cost and the scenario section, where there is no single right answer and the panel is grading your trade-offs.
  • Code snippets are short and illustrative. They were written against the LangGraph 1.x Python API; build each one yourself in a scratch project rather than memorising it.

LangGraph sits on top of LangChain concepts, so if your interview also covers chains, retrievers and LCEL, pair this page with the LangChain interview questions and answers.

Contents

Fundamentals

1. What is LangGraph, and what problem does it solve?

Answer: LangGraph is an open-source (MIT-licensed) orchestration framework and runtime for building stateful, long-running agents as graphs. You define a shared state, a set of nodes (functions that read state and return updates) and edges (what runs next). The runtime runs the graph step by step, can save state after each step, and can pause and resume.

The problem it solves is control. A free-running agent loop ("call the model, run whatever tool it picks, repeat") works in a demo but is hard to bound, audit or resume. LangGraph lets you mix deterministic steps (validate input, check policy, write to a database) with model-driven steps (choose a tool, draft a reply) in one explicit flow, with persistence, streaming and human approval built in.

Interview tip: Say "orchestration runtime", then name the three things it gives you that a loop does not: durable state, explicit control flow and interrupts.

2. LangChain vs LangGraph: how do they relate in the 1.x releases?

Answer: They are complementary layers. LangChain is the agent framework: model integrations, tools, messages, and the create_agent helper with middleware. LangGraph is the lower-level runtime underneath. In the 1.x releases, LangChain's create_agent is built on LangGraph, and the older langgraph.prebuilt.create_react_agent is deprecated in favour of from langchain.agents import create_agent.

In practice: start with create_agent when a standard tool-calling loop fits, and add middleware for things like summarisation, retries or approvals. Drop down to a hand-built StateGraph when you need custom branches, parallel fan-out, several cooperating agents, or a workflow that is mostly deterministic with a few LLM steps.

Interview tip: Answering "LangGraph replaced LangChain" is a common mistake. The current docs position them as two layers of one stack.

3. What are nodes, edges and conditional edges?

Answer: A node is a Python function (or runnable) that receives the current state and returns a partial update. A normal edge is a fixed transition: after node A, always run node B. A conditional edge calls a routing function on the state and returns the name of the next node (or several, or END). Conditional edges are how you build branches and loops, for example "if the last AI message has tool calls, go to tools, otherwise finish".

Two special markers bound the graph: START is the entry point and END is the exit. You wire them with add_edge(START, "first_node") and by returning END from a router.

Interview tip: Mention that nodes should return only the keys they change. Returning the whole state is a classic way to clobber another branch's update.

4. Walk me through the smallest useful LangGraph program.

Answer: Define a state schema, add a node, wire START to it, add a conditional edge that loops or ends, compile, invoke. Illustrative example:

import operator
from typing import Annotated, Literal, TypedDict
from langgraph.graph import StateGraph, START, END

class State(TypedDict):
    question: str
    notes: Annotated[list[str], operator.add]
    attempts: int

def research(state: State) -> dict:
    n = state.get("attempts", 0) + 1
    return {"notes": [f"pass {n}"], "attempts": n}

def route(state: State) -> Literal["research", "__end__"]:
    return END if state["attempts"] >= 3 else "research"

builder = StateGraph(State)
builder.add_node("research", research)
builder.add_edge(START, "research")
builder.add_conditional_edges("research", route)
graph = builder.compile()

compile() validates the graph and returns a runnable with invoke, stream and state methods. Calling graph.invoke({"question": "q", "notes": [], "attempts": 0}) runs three passes, and notes accumulates because of the operator.add reducer.

5. When is LangGraph overkill?

Answer: When the flow is a single model call or a fixed linear pipeline with no branching, no loops, no pause points and nothing to resume. A one-shot classification, a straightforward retrieve-then-answer chain or a batch summarisation job does not need a graph runtime. Use a plain function, an LCEL chain or create_agent.

Reach for LangGraph when at least one of these is true: the agent loops over tools; a human must approve something mid-run; the run must survive a restart or wait hours for input; several specialised steps or agents share state; or you need deterministic guardrail steps around non-deterministic model steps.

Interview tip: Panels respect a candidate who says "I would not use it here". It shows you choose tools by requirement, not habit.

6. What is the difference between the Graph API and the Functional API?

Answer: Both run on the same runtime and get the same persistence, streaming and interrupts. The Graph API (StateGraph) is declarative: you list nodes and edges, and the topology is visible and can be rendered as a diagram. The Functional API uses two decorators: @entrypoint marks the workflow function and @task marks units of work whose results are checkpointed. You write normal Python control flow (if, for) instead of edges.

Choose the Graph API when the flow has many branches, you want to visualise it, or several people maintain it. Choose the Functional API to add durability and interrupts to existing procedural code with minimal restructuring.

State and graph design

7. What is "state" in LangGraph?

Answer: State is the shared data structure that every node reads and updates. It is the single source of truth for one run. You define it with a TypedDict, a dataclass or a Pydantic model. Each key is a channel. Nodes return partial updates, and the runtime merges each update into the state using that key's reducer. With a checkpointer attached, the state after every step is saved, which makes it inspectable and resumable.

The prebuilt MessagesState is a common starting point: one key, messages, with the add_messages reducer. Most real graphs extend it with business fields such as ticket_id, risk_score or approved.

Real-world example: In a claims workflow, state might hold the claim ID, extracted fields, policy-check results, a list of documents still missing and the reviewer's decision. That is far easier to audit than a long message list.

8. What are reducers, and how does add_messages differ from operator.add?

Answer: A reducer is the function that combines the current value of a key with a node's update. Without one, the update overwrites the value. You attach a reducer with Annotated, for example Annotated[list[str], operator.add], which concatenates lists.

add_messages is smarter for chat history. It appends new messages, but if an incoming message has the same ID as an existing one it replaces it, and it handles RemoveMessage for deletions. It also converts shorthand tuples like ("user", "hi") into message objects. operator.add would blindly duplicate.

When you need to bypass a reducer once, for example to reset an error list, return the value wrapped in Overwrite(...) from langgraph.types.

Interview tip: Reducers also matter for parallel branches: two nodes writing the same key in one step need a reducer, or the update conflicts.

9. How do input, output and private state schemas help in graph design?

Answer: StateGraph accepts separate input_schema and output_schema alongside the internal state. The input schema limits what callers can send, and the output schema limits what the graph returns. Internal working fields such as raw retrieval results, scratch notes or intermediate scores stay inside. Nodes can also declare private state types for channels only some nodes use.

This gives you an API contract. The front end or calling service sees {"question"} in and {"answer", "citations"} out, and you can refactor internals without breaking callers. It also keeps sensitive intermediate data out of API responses.

10. What goes in state, and what goes in runtime context or config?

Answer: State holds data that changes during the run and should be checkpointed: messages, extracted fields, decisions. Runtime context holds static, per-run dependencies that are not part of the evolving state: user ID, tenant ID, role, feature flags, a database connection. In 1.x you declare it with StateGraph(State, context_schema=Ctx), pass context=Ctx(...) at invoke time, and read it in a node through a runtime: Runtime[Ctx] parameter. Config (config["configurable"]) carries runtime plumbing such as thread_id.

Putting a user's role in state is a security smell: a buggy node could overwrite it. Keeping it in context makes it read-only for the run.

11. How do you decide node granularity?

Answer: Split nodes at the points where you need one of the runtime's features: a checkpoint boundary, a retry boundary, a place to pause, a separate trace span, or a branch point. Keep together steps that always run together and have no reason to be resumed separately.

Too coarse, and a failure in step four of a giant node re-runs steps one to three on resume. Too fine, and you pay checkpoint overhead and get a graph nobody can read. A practical rule: one node per side effect (each external write gets its own node, so its retry and idempotency are easy to reason about) and one node per LLM call that you want to evaluate separately.

Interview tip: Tie your answer to resumability. "A node is the unit of re-execution" is the sentence panels listen for.

Control flow

12. What is Command, and when would you use it instead of a conditional edge?

Answer: Command lets a node update state and choose the next node in one return value. Use it when the routing decision is made inside the node's own logic and you would otherwise duplicate that logic in a separate router. Annotate the return type so the graph knows the possible destinations. Illustrative:

from typing import Literal
from langgraph.types import Command

def check_limit(state) -> Command[
        Literal["approve", "escalate"]]:
    if state["amount"] > 50_000:
        return Command(update={"status": "review"},
                       goto="escalate")
    return Command(update={"status": "auto"},
                   goto="approve")

Command has other uses too. Command(resume=...) resumes an interrupt, and Command(graph=Command.PARENT, goto=...) lets a node inside a subgraph hand control to a node in the parent graph, which is the basis of agent handoffs.

Interview tip: Conditional edges keep routing visible in the graph definition. Command keeps it next to the business logic. Say which you would choose and why.

13. What is the Send API, and how do you implement map-reduce with it?

Answer: Send(node_name, payload) schedules a node with its own input, separate from the main state. Return a list of Send objects from a conditional edge, and LangGraph runs that node once per item in parallel in the same step. A reducer on the output key gathers the results. Illustrative:

from langgraph.types import Send

def fan_out(state):
    return [Send("summarise", {"doc": d})
            for d in state["docs"]]

def summarise(payload: dict) -> dict:
    return {"summaries": [payload["doc"][:200]]}

builder.add_conditional_edges(START, fan_out,
                              ["summarise"])

Here summaries would be Annotated[list[str], operator.add]. The number of branches is decided at runtime, which static edges cannot do.

Real-world example: Checking each clause of a contract against a policy, or each attachment in a claim, in parallel and then combining the results in one review node.

14. How do cycles work, and how do you prevent infinite loops?

Answer: A cycle is just an edge back to an earlier node, such as agent β†’ tools β†’ agent. LangGraph supports it natively, which is the main reason it exists. To bound loops, use layers:

  • An explicit exit in the router (no tool calls, answer found, max attempts reached in state).
  • The recursion_limit in config. Exceeding it raises GraphRecursionError. The default has changed between releases, so set it explicitly per graph rather than relying on it.
  • The RemainingSteps managed value in state, so the agent can wrap up gracefully before the hard limit.
  • Business budgets: a tool-call counter, token or cost ceilings, and wall-clock timeouts.

Interview tip: Mention that the recursion limit is a safety net, not a design. The router should end the loop on purpose.

15. What is a super-step, and how do parallel branches behave?

Answer: LangGraph runs in discrete super-steps, based on the Pregel / bulk-synchronous model. In each super-step, every node scheduled for that step runs, possibly in parallel. Their updates are applied together through the reducers, and then the next set of nodes is chosen. A checkpoint is written per super-step.

Two consequences come up in interviews. First, parallel nodes in one step cannot see each other's updates until the next step. Second, if two parallel nodes write the same key without a reducer, the update fails as invalid. If one branch in a step raises an error, the step fails, but with a checkpointer the successful branches' writes are kept as pending writes, so a retry does not redo them.

16. How do you handle errors and retries inside a graph?

Answer: Use the right tool at each level:

  • Transient failures (timeouts, rate limits): attach a RetryPolicy to the node, for example add_node("call_api", fn, retry_policy=RetryPolicy(max_attempts=3)), with backoff.
  • Expected business failures (record not found, validation failed): catch them in the node, write an error field to state and route to a recovery or escalation node.
  • Tool errors in agent loops: return the error as a tool message so the model can correct its arguments, with a cap on attempts.
  • Crashes: rely on the checkpointer and resume the thread from the last saved step.

Interview tip: Never wrap interrupt() in a bare try/except. It works by raising a special exception, and catching it breaks the pause.

Persistence

17. What does a checkpointer do, and what are threads?

Answer: A checkpointer saves a snapshot of the graph state after each super-step. A thread is the ID under which those snapshots are grouped, passed as config={"configurable": {"thread_id": "..."}}. Reuse the same thread ID and the graph continues from its last checkpoint. Use a new one and it starts fresh.

Checkpointing enables conversation memory within a thread, human-in-the-loop pauses, fault-tolerant resume after a crash, time-travel debugging, and inspection of any past step with get_state and get_state_history.

from langgraph.checkpoint.memory import InMemorySaver

app = builder.compile(checkpointer=InMemorySaver())
cfg = {"configurable": {"thread_id": "ticket-4711"}}
app.invoke({"question": "Reset my VPN token"}, cfg)
print(app.get_state(cfg).values)

18. Which checkpointer would you use in development and in production?

Answer: InMemorySaver for unit tests and notebooks. It lives in RAM, and everything is lost when the process restarts. SqliteSaver (package langgraph-checkpoint-sqlite) for local development and single-process tools. PostgresSaver or AsyncPostgresSaver (package langgraph-checkpoint-postgres) for production services. Community and partner integrations exist for other databases; check current documentation for their status.

Production concerns to mention: connection pooling, using the async saver with async graphs, running the saver's setup() for migrations, retention and cleanup of old threads, encryption of serialized state, and keeping the checkpoint database in the same region as the app for latency and data residency.

Interview tip: If you are asked "why did all my conversations vanish after deploy?", the answer is almost always an in-memory checkpointer in production.

19. What is time travel, and how do you use it for debugging?

Answer: Because every step is checkpointed, you can list a thread's history with get_state_history(config), pick an earlier checkpoint by its checkpoint_id, and either replay from it (invoke with None as input and a config containing that checkpoint ID) or fork it by calling update_state with corrected values and resuming. The original history remains, and the fork becomes a new branch.

Real-world example: A support agent picked the wrong tool at step 6 of a 10-step run. Instead of reproducing the whole conversation, you fork at step 5, patch the state (or swap the prompt), and re-run from there to confirm the fix.

20. What are durability modes?

Answer: The durability argument to invoke/stream controls when checkpoints are written. "sync" persists before the next step starts, the safest choice. "async", the default, persists while the next step runs, so a crash at the wrong moment can lose the latest step. "exit" persists only when the graph finishes or pauses, the fastest choice, but intermediate progress is lost on a crash.

Choose per workflow. A payment or provisioning workflow with external side effects should use "sync". A short chat turn can live with "async". A high-volume batch job where re-running from scratch is cheap might use "exit".

21. How do you make side effects safe under resume and replay?

Answer: Assume any node can run more than once: on retry, on resume after a crash, and on resume after an interrupt (the whole node re-runs from the top). So:

  • Put each external write in its own node, after any interrupt in that node, never before it.
  • Make writes idempotent with a key derived from the thread and step (for example thread_id + action) that the downstream API or your own table de-duplicates.
  • In the Functional API, wrap side effects and non-deterministic calls in @task so their results are checkpointed and reused on resume.
  • Record the outcome in state ("refund_issued": true) and check it before acting.

Interview tip: Mentioning idempotency keys separates engineers who have run agents in production from those who have only run notebooks.

Human-in-the-loop

22. How do you implement human-in-the-loop in LangGraph?

Answer: Call interrupt(payload) inside a node. The runtime saves state, stops, and surfaces the payload to the caller (in the result's __interrupt__ field, or as a stream event). Later, invoke the same thread with Command(resume=value). That value becomes the return value of interrupt(), and the node continues. It requires a checkpointer and a thread ID. Illustrative:

from langgraph.types import interrupt, Command

def human_review(state):
    answer = interrupt({"refund": state["refund"],
                        "ask": "Approve this refund?"})
    return {"decision": answer}

# caller, later, same thread_id:
app.invoke(Command(resume="approved"), cfg)

The pause can last seconds or days. Nothing runs while it waits. The state sits in the checkpoint store.

For the product side of approval design, such as who approves, SLAs and audit, see human-in-the-loop AI patterns.

23. What happens to code before interrupt() when you resume?

Answer: It runs again. On resume, the runtime restarts the whole node from the beginning. It does not continue from the line where interrupt() was called. Everything above the interrupt executes a second time. Three rules follow from the docs:

  • Code before an interrupt must be idempotent, or it should be moved into an earlier node.
  • Do not wrap interrupt() in a try/except that swallows exceptions.
  • If a node has several interrupts, call them in the same order every time. Resume values are matched by index, so conditionally skipping one breaks the mapping.

When parallel branches each interrupt, you resume with a mapping from each interrupt's ID to its value.

24. Static breakpoints vs dynamic interrupts: what is the difference?

Answer: Static breakpoints (interrupt_before=[...] or interrupt_after=[...] at compile or invoke time) always pause at a named node. They are mainly for debugging and tests: step through a graph or stop after node 3 so node 4 does not run. Dynamic interrupts (interrupt() in a node) pause conditionally, carry a payload explaining what needs a decision, and receive a typed answer. They are the right tool for production approvals.

Interview tip: Say you would only interrupt when it is needed, for example when the amount exceeds a threshold or the action is irreversible. Approving every step trains reviewers to click "yes" without reading.

25. How do you add approvals to a create_agent agent without hand-building a graph?

Answer: Use LangChain's HumanInTheLoopMiddleware. You configure interrupt_on per tool: True to require review, False to auto-approve, or a config with allowed_decisions. When the model proposes a guarded tool call, the agent pauses with the pending action. The reviewer resumes with a decision: approve, edit (changed arguments), reject (with feedback) or respond (a direct reply in place of the tool result). Illustrative:

from langchain.agents import create_agent
from langchain.agents.middleware import (
    HumanInTheLoopMiddleware)

agent = create_agent(
    model=model, tools=[issue_refund, lookup_order],
    middleware=[HumanInTheLoopMiddleware(
        interrupt_on={"issue_refund": True,
                      "lookup_order": False})],
    checkpointer=checkpointer)

agent.invoke(Command(resume={"decisions": [
    {"type": "approve"}]}), cfg)

Under the hood it is the same interrupt() mechanism, since create_agent runs on LangGraph.

Memory

26. Short-term vs long-term memory in LangGraph: what is the difference?

Answer: Short-term memory is thread-scoped. It is the graph state (usually the message history) saved by the checkpointer for one conversation or one workflow run. Long-term memory is cross-thread. It lives in a store (BaseStore), a key-value document store organised by namespaces, and survives across conversations: user preferences, learned facts, past resolutions.

You attach both at compile time: builder.compile(checkpointer=..., store=...). Nodes access the store through the runtime (runtime.store). For the wider design space (episodic, semantic and procedural memory), see how AI agent memory works.

27. How does the long-term memory store work?

Answer: Items are JSON documents addressed by a namespace tuple plus a key, for example ("users", user_id, "prefs") and "language". The core methods are put, get, search, delete and list_namespaces, with async equivalents. InMemoryStore is for development, and PostgresStore is the common production choice. Illustrative:

def remember(state, runtime: Runtime[Ctx]):
    ns = ("users", runtime.context.user_id, "prefs")
    runtime.store.put(ns, "language",
                      {"value": "Telugu"})
    hits = runtime.store.search(ns, limit=5)
    return {"messages": [...]}

Configure an index with an embedding function and dimensions, and search(ns, query="...") does semantic similarity search over stored items. Stores can also support TTLs, so stale memories expire.

28. How do you keep a long conversation within the context window?

Answer: Manage the short-term memory explicitly instead of sending the whole thread every turn:

  • Trim: keep the last N tokens of messages with trim_messages, always keeping the system message and complete tool-call/tool-result pairs.
  • Delete: return RemoveMessage(id=...) updates. The add_messages reducer removes them from state.
  • Summarise: a node (or LangChain's summarisation middleware) folds older turns into a running summary field.
  • Offload: move durable facts to the store and retrieve them when relevant.

Trimming only what you send to the model keeps the full history in the checkpoint for audit. Deleting from state shrinks the checkpoint too. Choose based on retention policy.

29. How do you stop one user's memories leaking into another user's session?

Answer: Make tenant and user identity part of every namespace, and take that identity from trusted runtime context, never from model output or user text. A namespace like (tenant_id, user_id, "memories"), built in code from an authenticated context, means a search cannot cross users even if the prompt asks it to.

Also restrict which nodes can write to the store, validate what gets saved (no secrets, no raw PII unless policy allows it), support deletion requests, and log store reads alongside the trace. Under India's DPDP Act and similar regimes, a user's ability to see and erase stored memories is a design requirement, not an extra.

Multi-agent patterns in LangGraph

30. How would you build a multi-agent system in LangGraph?

Answer: Start by asking whether you need more than one agent. Often one agent with good tools, or one agent loading specialised "skills" (prompts and knowledge) on demand, is enough. If you do need several, the LangChain docs describe these patterns:

PatternHow it worksGood for
SubagentsA main agent calls specialist agents as tools and keeps controlMulti-domain tasks, parallel work, teams owning separate agents
HandoffsThe active agent transfers control (and state) to anotherMulti-hop conversations where the user talks to whoever is active
RouterA classification step sends input to one or more specialistsClear-cut domains, input triage
SkillsOne agent loads specialised instructions on demandFocused tasks without multi-agent overhead
Custom workflowA hand-built StateGraph mixing fixed steps and agentsRegulated, mostly deterministic processes

In LangGraph terms, each agent is a node or a subgraph. Routing is done with conditional edges, Command or Send, and handoffs use Command(graph=Command.PARENT, goto=...). For a worked build, see the multi-agent enterprise workflow project.

31. Supervisor vs handoff (swarm-style) designs: what are the trade-offs?

Answer: A supervisor (subagents pattern) centralises routing. Every turn goes through one coordinator, which makes control flow easy to reason about, audit and limit. The cost is an extra model call per hop and a coordinator that can become a bottleneck or a single point of confusion. Handoffs let the active agent pass control directly, which cuts latency and lets a specialist hold a long conversation. But flow is harder to predict, and you need guards against ping-pong between agents.

Earlier helper libraries (langgraph-supervisor, langgraph-swarm) packaged these patterns. The current docs focus on building them with core primitives, or using the higher-level Deep Agents harness. Check current documentation before depending on a helper library.

32. How do subgraphs work, and how do they share state with the parent?

Answer: A subgraph is a compiled graph used inside another graph. There are two ways to connect it. If parent and child share state keys, add the compiled subgraph directly as a node and it reads and writes those channels. If the schemas differ, call the subgraph inside a wrapper node that maps parent state to child input and the child output back.

Persistence is controlled by the subgraph's compile(checkpointer=...). The default (None) inherits the parent's checkpointer for one call, so interrupts work but each call starts fresh. True keeps per-thread state that accumulates across calls. False runs it like a plain function. Stream with subgraphs=True to see nested events with their namespace.

Interview tip: Subgraphs are also a team boundary. The fraud team owns one subgraph, the KYC team another, each with its own tests.

33. How do you stop agents from looping between each other or duplicating work?

Answer: Put shared, structured progress into state that every agent reads: a task list with statuses, a record of which agent did what, and a hop counter. Then enforce limits in code: a maximum number of handoffs per run, no handing back to the agent that just handed off without new information, and a coordinator rule that a finished task is never reassigned.

Give each agent a narrow description and tool set, so routing decisions have less overlap. Evaluate routing as its own metric (was the right agent chosen?) rather than only grading final answers. Many "multi-agent loops" are really two agents with overlapping descriptions.

Tools and MCP integration

34. How do tool calls work in a LangGraph agent loop?

Answer: The model node calls a chat model bound to tool schemas. If the returned AI message contains tool_calls, a conditional edge (the prebuilt tools_condition, or your own) routes to a tool-execution node. ToolNode from langgraph.prebuilt runs each requested tool, can run several in parallel, and appends ToolMessage results to state. The loop returns to the model until it answers without tool calls.

Tools that need graph data can declare injected parameters (InjectedState, InjectedStore, or a ToolRuntime argument) that the runtime fills in and the model never sees. That is how you pass a user ID into a tool without letting the model choose it.

35. How do you connect a LangGraph agent to MCP servers?

Answer: MCP (Model Context Protocol) is an open protocol, introduced by Anthropic in late 2024, for connecting AI applications to tools and data. LangGraph does not need special MCP support. You load MCP tools as standard LangChain tools and use them in ToolNode or create_agent. The widely used route is the langchain-mcp-adapters package, whose MultiServerMCPClient connects to several servers (stdio for local processes, streamable HTTP for remote ones) and returns their tools. Recent LangChain releases also include a built-in MCP adapter, marked beta at the time of writing. Check current documentation for which to use.

Enterprise points to raise: authenticate the MCP connection per user or service identity, allow-list which server tools the agent may see, and log every MCP call in the trace. To build the server side yourself, see the MCP server Python tutorial.

36. How do you keep a large tool catalogue manageable?

Answer: Fewer, sharper tools beat a giant list. Every tool schema costs prompt tokens and adds a way for the model to pick wrongly. Options:

  • Split by domain into subagents, each with its own small tool set.
  • Select tools dynamically per request (LangChain has tool-selection middleware that uses a cheaper model to pre-filter).
  • Merge chatty low-level tools into task-level tools ("create_ticket_with_assignment" instead of three calls).
  • Write descriptions that say when not to use the tool.

Measure tool-selection accuracy on an eval set before and after any change.

37. How do you secure tools that write to enterprise systems?

Answer: Treat the model as an untrusted caller. Concretely: run tools with the end user's delegated permissions (not a shared admin account), validate arguments with a strict schema plus business rules in code, gate irreversible actions behind an interrupt or the HITL middleware, use idempotency keys, rate-limit and cap tool calls per run, and never let model output choose identities, tenants or URLs. Defend against prompt injection that arrives through tool results (retrieved documents, emails, tickets) by not giving those results the authority to trigger new privileged calls without a check. For identity patterns, see identity and access for AI agents.

Testing and evaluation

38. How do you unit test a LangGraph graph?

Answer: Build the graph in a factory function, compile it in each test with a fresh InMemorySaver, and test at three levels:

  • Single node: graph.nodes["research"].invoke(state) runs one node in isolation.
  • Routing: call router functions directly with crafted states and assert the destination.
  • Partial runs: seed state with update_state(cfg, values, as_node="step_2") to pretend earlier steps ran, and use interrupt_after=["step_4"] to stop where you want to assert.

Replace the model with a fake chat model that returns scripted messages (including tool calls), so tests are fast, free and deterministic. Use pytest.

def test_router_stops_after_three():
    s = {"question": "q", "notes": [], "attempts": 3}
    assert route(s) == END

39. Unit tests pass, but the agent still gives bad answers. How do you evaluate it?

Answer: Unit tests check wiring. Evaluations check behaviour. Build a dataset of realistic inputs with expected outcomes, run the graph on it, and score:

  • Final output: correctness against a reference, groundedness for RAG answers, format and policy compliance.
  • Trajectory: did it call the right tools in a sensible order, with valid arguments, without unnecessary steps?
  • Single steps: routing accuracy, retrieval quality, the arguments of one critical tool call.
  • Operational: steps, tokens, latency and interrupts per run.

LangSmith datasets and evaluators, Ragas for RAG metrics, and LLM-as-judge with human calibration are common choices. See how to evaluate AI agents for metric design.

40. How do you test human-in-the-loop flows?

Answer: With a checkpointer and a fixed thread ID, invoke until the interrupt. Assert that __interrupt__ is present and its payload contains exactly what a reviewer needs. Then resume with each decision type (approve, edit, reject) and assert the downstream path and side effects. Also test the uncomfortable cases: resume twice, resume after a long gap with changed data, resume a thread that already finished, and confirm the code before the interrupt is safe to re-run.

41. How do you run evaluations in CI without huge cost or flakiness?

Answer: Tier them. On every pull request, run deterministic unit tests plus a small, fixed "smoke" eval set with pinned prompts, low temperature and cached model responses where possible. Run the full eval set nightly or before release, and compare against a stored baseline with thresholds instead of exact-match assertions. Track per-example regressions, not just averages, because one broken critical path can hide inside a stable mean. When a production trace shows a failure, add it to the dataset.

Deployment and observability

42. What are the options for deploying a LangGraph application?

Answer: Two broad routes.

  • Self-managed: LangGraph is a library, so wrap the compiled graph in your own service (for example FastAPI), run it in containers on Kubernetes or a serverless platform, and point it at a Postgres checkpointer and store. You own queueing, scaling, auth and long-running background runs.
  • LangSmith Deployment (previously called LangGraph Platform; renamed in late 2025, alongside LangGraph Studio becoming LangSmith Studio): a managed runtime for agent workloads built around an Agent Server with assistants (configurations), threads (state) and runs (executions), plus task queues, persistence and streaming endpoints. Documented environments include Cloud, Hybrid (vendor-managed control plane, your data plane), Self-hosted with a control plane, and Standalone Agent Server via Docker or Kubernetes. Plan requirements differ by environment.

Locally, the LangGraph CLI with a langgraph.json config runs a development server you can attach Studio to.

Interview tip: Use the current name, and mention the old one once. It shows you keep up without being confused by older blog posts.

43. What are the streaming modes, and which would you use for a chat UI?

Answer: stream(..., stream_mode=...) supports: values (full state after each step), updates (only what each node changed), messages (LLM tokens with metadata, as they are generated), custom (anything a node emits via get_stream_writer()), checkpoints and tasks (both need a checkpointer), and debug (everything). You can pass a list to combine modes. Passing version="v2" gives every chunk a uniform shape, {"type", "ns", "data"}, so code does not have to branch on single vs multiple modes or subgraph namespaces.

For a chat UI: messages for tokens, plus updates or custom for progress such as "searching policy documents…". Use values for debugging or small states only. Newer releases also add an event-centric stream_events API whose latest version is marked experimental.

for part in app.stream(inputs, cfg,
        stream_mode=["messages", "updates"],
        version="v2"):
    if part["type"] == "messages":
        token, meta = part["data"]

44. How do you observe a LangGraph app in production?

Answer: Trace every run, with spans per node, model call and tool call, plus inputs and outputs, token counts, latency and errors, tagged with the thread ID, user, tenant and graph version. LangSmith traces LangGraph natively once its environment variables are set. Langfuse and OpenTelemetry-based backends are common alternatives, especially where data must stay in your own cloud. On top of traces, build dashboards for: runs per outcome (completed, interrupted, failed, hit limit), steps per run, cost per run, interrupt wait time, and tool error rates.

Redact or hash PII before it reaches the tracing backend, and set retention rules. See AI observability for the full signal list.

45. How do you version and roll out graph changes safely?

Answer: Treat a graph like an API with stored data. Version the graph and its prompts together, and record the version in run metadata. Keep state schema changes backward-compatible: add optional fields, avoid renaming or removing keys that paused threads still hold. Avoid deleting or renaming nodes that interrupted threads may resume into. Roll out with a canary or by assistant/config version, compare eval and production metrics, and keep the previous version deployable. If a breaking change is unavoidable, drain or migrate in-flight threads first.

Performance and cost

46. Where does latency come from in a LangGraph agent, and how do you reduce it?

Answer: Mostly from sequential model calls, then tool latency, then checkpoint writes. Tactics:

  • Cut hops: replace an LLM router with a rule where the decision is obvious, and merge steps that always run together.
  • Parallelise independent work with multiple outgoing edges or Send.
  • Use a smaller model for routing and extraction, and save the large model for the final answer.
  • Stream tokens and progress so perceived latency drops even when total time does not.
  • Cache deterministic nodes with a CachePolicy (TTL), and use provider prompt caching for long, stable system prompts.
  • Choose a durability mode that matches the risk, and keep the checkpoint database close to the app.

More detail is in LLM latency optimization.

47. How do you control cost per run?

Answer: Measure cost per run and per outcome first, from traces. Then bound it in code: recursion limits, tool-call and model-call caps (LangChain middleware offers limit middlewares for agents), and token budgets stored in state that force a wrap-up. Reduce tokens per call by trimming history, summarising, sending only relevant state to each node rather than the whole message list, and keeping tool catalogues small. Route easy requests to cheaper models, and cache repeat answers. An LLM gateway in front of providers adds central budgets, fallbacks and per-team chargeback.

48. How do you scale a LangGraph service horizontally?

Answer: Keep workers stateless. All durable state lives in the checkpointer and store, so any replica can continue any thread. Run long or background runs through a task queue rather than inside HTTP requests, and make sure two workers never execute the same thread at the same time (a per-thread lock or the platform's run queue handles this). Size the Postgres connection pool to the replica count, prefer async graphs for I/O-heavy nodes, and monitor checkpoint table growth. Rate limits at the model provider usually become the real ceiling before CPU does, so plan quotas and backoff.

Scenario-based questions

49. A bank wants an agent that drafts refund decisions, but anything above a threshold needs a human. Design it.

Answer: Consider a retail bank's card-disputes desk. I would build a custom StateGraph, not a free agent: intake (validate and fetch the transaction through a read-only tool) β†’ assess (LLM drafts a decision with reasons and policy citations, as structured output) β†’ a conditional edge on amount and risk β†’ human_review (dynamic interrupt() with the draft, evidence and policy excerpt) β†’ execute (idempotent refund call) β†’ notify. Below the threshold, a deterministic policy check can auto-approve. Above it, the case always pauses.

intake -> assess -> [amount/risk?]
                     |-- low  -> policy_check -> execute
                     '-- high -> human_review
                                   |-- approve -> execute
                                   '-- reject  -> notify
execute -> notify -> END

What I would check:

  1. That the threshold and risk rules live in code or config, not in the prompt.
  2. That the reviewer's identity, decision and edits are written to state and an audit log.
  3. That execute uses an idempotency key, so a resume cannot refund twice.
  4. That a Postgres checkpointer with durability="sync" is used for this flow.
  5. Reviewer SLAs and what happens to cases paused past the SLA.

Production consideration: Regulators and internal audit will ask "who decided?". The graph must make the human decision explicit and traceable, never implied by a prompt. For domain controls, see generative AI in banking.

50. Token spend tripled overnight and traces show some runs taking hundreds of steps. What do you do?

Answer: Contain first, then diagnose. Lower the recursion limit and add or tighten tool-call caps through a config change, so runaway runs end quickly. Then open the longest traces and look for the repeating pattern.

What I would check:

  1. Whether a tool started failing or returning an unhelpful error the model keeps retrying.
  2. Whether a prompt, model or tool-description change shipped yesterday.
  3. The router's exit condition: does it still detect "done" with the new model's output format?
  4. Two agents handing a task back and forth (a hop counter will show it).
  5. History growth: each loop re-sends a longer context, so cost grows faster than step count.

Production consideration: Add an alert on steps per run and cost per run, not just total spend, and add the failing traces to the regression eval set before shipping the fix.

51. A hospital discharge-summary workflow crashed mid-run. After resume, the EHR got two draft notes. Why, and how do you fix it?

Answer: The node that posted the draft was re-executed on resume. Either the write happened in a node whose checkpoint had not been persisted when the pod died (async or exit durability), or the write sat before an interrupt() in the same node, so the whole node ran again on resume.

What I would check:

  1. The durability mode for this graph, and whether the write node completed its checkpoint.
  2. Whether the EHR write and an interrupt share a node.
  3. Whether the EHR API accepts an idempotency key, or whether we de-duplicate on our side using thread ID plus document type.
  4. That state records "draft_posted" with the returned document ID, and that the node checks it before posting.

Production consideration: In clinical systems a duplicate is a safety issue, not a cosmetic one. Use sync durability, isolate each external write in its own node, make it idempotent, and keep a clinician review step before anything is filed as final.

52. You inherit an app built on the old AgentExecutor or create_react_agent. How do you migrate it to the 1.x stack?

Answer: Migrate in two moves. First, move the standard tool-calling agent to langchain.agents.create_agent. Map the system prompt to system_prompt, the pre/post-model hooks and custom logic to middleware, and the memory to a checkpointer with thread IDs. Second, if the old code had bespoke control flow hacked into callbacks or prompt instructions ("always ask the manager before..."), move that into an explicit StateGraph with the agent as one node.

What I would check:

  1. A golden eval set that runs on old and new versions before the switch.
  2. Deprecated imports (for example from langgraph.prebuilt) and changed message or state formats.
  3. How existing conversation history will be loaded into the new threads, or whether to start fresh.
  4. Streaming consumers in the front end, since chunk formats may differ.

Production consideration: Run both behind a flag, shadow traffic where policy allows, and switch per tenant. Do not combine a framework migration with a model change in the same release.

53. An insurer's claims graph fans out document checks with Send. One check times out, and the whole claim is stuck. What do you change?

Answer: Consider an insurer checking each claim attachment (invoice, discharge summary, ID proof) in parallel. A super-step fails if any branch raises, so one slow OCR call blocks every claim it touches. I would give the check node a RetryPolicy for transient errors and a timeout, and on final failure return a structured "check_failed" result instead of raising. The reduce node can then route the claim to manual review with the failed item marked.

What I would check:

  1. Whether successful branches' writes were kept (with a checkpointer they are saved as pending writes, so a retry does not redo them).
  2. Timeout settings on the external OCR or extraction service.
  3. That the reducer can merge partial results and record which documents are missing.
  4. Concurrency limits, so a 40-attachment claim does not exhaust the provider's rate limit.

Production consideration: Degrade to "needs human" rather than "stuck". A stuck claim is invisible, while a routed one has an owner and an SLA.

54. In a GCC IT-ops multi-agent platform, the supervisor keeps routing network incidents to the database agent. How do you debug it?

Answer: Consider an IT-ops team in a Hyderabad GCC running a supervisor with network, database and identity specialists. Mis-routing is usually a description and evidence problem, not a model problem. I would pull the misrouted traces and compare what the supervisor saw against the agent descriptions.

What I would check:

  1. Overlapping agent descriptions (for example both mention "connection timeout").
  2. Whether the supervisor receives the alert's structured fields (source system, CMDB class) or only free text.
  3. Whether a deterministic pre-router on CMDB class can handle the obvious cases before any LLM decision.
  4. Routing accuracy on a labelled set of past incidents, measured before and after changes.
  5. Whether the database agent "accepts" tasks it should hand back. Add an explicit "not mine" path.

Production consideration: Make routing an evaluated component with its own metric and dataset. Teams that only grade final resolutions cannot see where errors start.

55. The Postgres checkpoint tables have grown to a size the DBA is unhappy about. What is your plan?

Answer: Checkpoints accumulate per step per thread forever unless you manage them. Large states (full documents or long message histories in state) multiply the problem.

What I would check:

  1. Which graphs and threads dominate the storage, and the average checkpoint size.
  2. Whether big payloads (retrieved documents, file contents) are stored in state when a reference such as an object-store key would do.
  3. Retention requirements from compliance: how long must finished threads be kept, and in what form?
  4. Whether short-lived graphs could use exit durability or no checkpointer at all.

Production consideration: Implement a scheduled cleanup that deletes or archives finished threads past retention (the checkpointer API supports deleting a thread), shrink state to references, and add checkpoint-size monitoring so growth is caught early.

56. A retailer's support assistant suddenly greets a customer with another customer's order history. What went wrong?

Answer: Either two users shared a thread ID, or the long-term memory namespace did not include the user (or tenant). Both are identity bugs, not model bugs. Consider a retailer whose web and WhatsApp channels generate thread IDs differently, where one channel falls back to a constant ID when the session cookie is missing.

What I would check:

  1. How thread IDs are generated, and whether any path can produce a shared or default value.
  2. Store namespaces: are they built from authenticated context, and do they include tenant and user?
  3. Any tool that queries orders: does it take the customer ID from injected context or from model-supplied arguments?
  4. Caches (node caching, response caching) keyed without the user.

Production consideration: Treat this as a security incident with a data-exposure review. Add tests that assert cross-user isolation, and reject requests that lack an authenticated identity instead of defaulting.

57. The chat UI shows nothing for 15 seconds, then the whole answer at once. How do you fix it?

Answer: The front end is probably consuming invoke or stream_mode="values", which only produce output at step boundaries. Switch to stream_mode=["messages", "updates"] (or add custom progress events from long nodes with get_stream_writer()), use the v2 chunk format, and send chunks over server-sent events or WebSockets.

What I would check:

  1. That the model call inside the node supports streaming and is not wrapped in code that collects the full response first.
  2. Proxies or load balancers buffering the response (a common culprit in nginx or API gateway setups).
  3. Filtering by node metadata so internal router tokens do not leak into the chat bubble.
  4. Time to first token vs total time, to see if the delay is upstream of generation (retrieval, routing).

Production consideration: Progress messages ("checking your policy…") during tool steps matter as much as token streaming for perceived speed.

58. Hundreds of procurement approvals are paused for days. You need to deploy a graph change. What could break?

Answer: Paused threads hold state shaped by the old schema and point at nodes in the old graph. On resume they run your new code. Renamed state keys, removed nodes, a changed interrupt payload contract, or an added interrupt before an existing one (index-based matching) can all break resumes or corrupt decisions.

What I would check:

  1. A diff of state schema, node names, edges and interrupt order between versions.
  2. That new state fields are optional with safe defaults.
  3. A test that resumes a thread created by the old version against the new code, using a copy of real checkpoints.
  4. Whether the approval UI renders both old and new interrupt payloads.

Production consideration: Keep changes additive, version graphs (with LangSmith Deployment, as separate assistants or revisions), and if needed let old threads finish on the old version while new threads start on the new one.

59. Your architect asks: LangGraph, plain create_agent, or a workflow engine such as Temporal or AWS Step Functions? How do you decide?

Answer: Decide by where the uncertainty is. If the process is mostly fixed business steps with long waits, compensation logic and strict SLAs, and the LLM is one step among many, a general workflow engine is a strong fit, with the LLM call as an activity. If the flow is a standard "model plus tools" loop, create_agent with middleware is the least code. LangGraph is the middle ground: agent-native state (messages, tool calls), interrupts, streaming of tokens, and checkpointed loops where the model influences the path.

NeedLeans towards
Standard tool-calling assistantcreate_agent
Custom agent loops, HITL, token streamingLangGraph
Long business process, compensation, many non-AI stepsWorkflow engine (LLM as a step)

What I would check:

  1. What the platform team already runs and supports on call.
  2. How much of the flow is model-decided vs fixed.
  3. Requirements for audit, retention and data residency.

Production consideration: Combinations are normal. A workflow engine can call a LangGraph service for the agentic step. For a wider framework comparison, see AI agent frameworks compared.

60. Whiteboard an enterprise knowledge assistant on LangGraph, end to end.

Answer: Consider an internal assistant for a services company's HR and IT policies, used by employees in Hyderabad and Bengaluru through Teams. The graph: guard_input (PII and injection checks) β†’ classify (cheap model or rules: policy question, action request, out of scope) β†’ for questions, retrieve (hybrid search filtered by the user's entitlements from runtime context) β†’ grade (are the documents relevant? if not, rewrite the query once) β†’ answer (with citations) β†’ guard_output. For action requests, an agent subgraph with MCP tools for the ticketing system, with writes gated by the HITL middleware.

guard_in -> classify -> question -> retrieve -> grade
                |                     ^          |
                |                     '-rewrite--'
                |                                v
                |                     answer -> guard_out
                '-> action -> agent subgraph (MCP)
                                -> approval -> done

What I would check:

  1. Postgres checkpointer for threads, Postgres store for user preferences, namespaced by tenant and user.
  2. Retrieval permissions enforced in the retrieve node, not by the prompt.
  3. Streaming to Teams with progress events, and tracing (LangSmith or Langfuse) tagged by graph version.
  4. Eval sets for retrieval quality, answer groundedness, routing and tool-call correctness, run in CI.
  5. Budgets: recursion limit, a single rewrite loop, model tiering.

Production consideration: Most of the effort is in entitlements, document freshness and evaluation, not graph wiring. For the RAG layer, see the RAG interview questions. For the agent theory, see the agentic AI interview questions.

If you want to build graphs like this with mentors reviewing your design decisions, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers LangGraph agents alongside cloud and security fundamentals, in Ameerpet classrooms or live online.

Key takeaways

  • LangGraph is the orchestration runtime. LangChain 1.x's create_agent runs on it. Know when to use each.
  • State design (keys, reducers, context vs state) decides how easy the rest of the graph is to build, test and resume.
  • Command combines update and routing; Send gives dynamic fan-out; recursion limits are a safety net, not an exit condition.
  • Checkpointers plus thread IDs give memory, resume, time travel and the foundation for interrupts. Use a database-backed saver in production.
  • On resume, a node re-runs from the top. Make side effects idempotent and isolated.
  • Long-term memory lives in a namespaced store. Build namespaces from trusted identity.
  • Production readiness means evaluations, tracing, versioning of paused threads and cost limits, not just a working graph. The deployment product is now called LangSmith Deployment.

Interview preparation checklist

  • Build a StateGraph from memory: custom state, a reducer, a conditional edge loop, compile and invoke.
  • Add an InMemorySaver, run two turns on one thread, then inspect get_state_history and fork from an earlier checkpoint.
  • Implement an approval node with interrupt() and resume it with Command(resume=...), and explain why code above the interrupt re-runs.
  • Write a map-reduce with Send and a list reducer.
  • Build the same assistant twice: once with create_agent plus HumanInTheLoopMiddleware, once as a custom graph. Be ready to compare them.
  • Connect at least one MCP server's tools to an agent and explain how you would restrict them.
  • Write pytest tests for one node, one router and one partial run with update_state(as_node=...).
  • Stream with messages and updates modes into a simple UI.
  • Trace a run in LangSmith or Langfuse and be able to explain one failure from the trace.
  • Prepare two stories from your own project: one design trade-off and one production bug, told as problem, diagnosis, fix and prevention.

FAQ

What skills are required for a LangGraph developer role?

Strong Python, a clear grasp of LLM tool calling and structured outputs, graph and state design, persistence with Postgres, API and cloud deployment basics, plus evaluation and tracing. Enterprise roles also expect security awareness around tools and data.

How should I prepare for a LangGraph interview?

Build one end-to-end project with state, a tool loop, a checkpointer, an approval interrupt, tests and tracing. Then practise explaining design trade-offs and walking through a production failure using the scenario questions above.

Do I need to learn LangChain before LangGraph?

Learn the LangChain basics first: chat models, messages, tools and create_agent. LangGraph builds on those concepts, and most interviews expect you to know when the higher-level agent API is enough and when to drop down to a custom graph.

Is LangGraph worth learning in 2026?

Yes, if you are aiming at agent or GenAI engineering roles. The concepts it makes explicit, such as state, checkpoints, interrupts and bounded loops, carry over to other agent frameworks and workflow engines, so the learning is not wasted even if a team uses a different tool.

Is LangGraph only available in Python?

No. LangGraph has Python and JavaScript/TypeScript versions with the same core ideas. Python is the more common choice in interviews for AI engineering roles, but the concepts transfer directly.

Do I need LangSmith to use LangGraph?

No. LangGraph is an open-source library that runs anywhere Python runs. LangSmith is optional for tracing, evaluation and managed deployment, and teams also use Langfuse or OpenTelemetry-based tools for observability.

How much LangGraph should a fresher know?

A fresher should be able to build and explain a small graph with state, reducers, a conditional edge, a tool loop and a checkpointer, and describe how human approval works. Deep deployment and multi-agent design are usually expected at mid and senior levels.

What LangGraph projects should I put on my resume?

Pick one or two realistic projects over many demos: for example a policy knowledge assistant with grading and citations, or an IT ticket agent with tool calls and an approval step. Show the graph diagram, tests, eval results and a trace.

How is LangGraph different from CrewAI or AutoGen?

LangGraph is lower level: you define the state and the control flow explicitly and get built-in persistence and interrupts. Role-based multi-agent frameworks give higher-level abstractions with less control over each step. Pick based on how much control and auditability the use case needs.

Ready to turn these answers into working systems? Explore the Cloudsoft APEX program for hands-on AI, ML, cloud and security engineering. If your goal is taking agents like these into customer environments, from discovery and integration to deployment and evaluation, the FDE PRO Forward Deployed Engineer course builds that through five enterprise projects, a simulated customer capstone and placement support until you're placed. Classroom in Ameerpet or live online. Call +91 96660 19191 for a free demo.

New Β· AI Career Guide

Meet Aanya β€” ask anything about courses, fees & placement

Instant answers from verified Cloudsoft info β€” courses, fees, formats, placement support and free demos. Available 24/7, right here on the site.

How Aanya works β†’
Share𝕏infβœ‰
EnrollWhatsAppCall us