New batches starting this week Β· Limited seats

Agentic AI Design Patterns: ReAct, Plan-and-Execute, Reflection and More

Ten framework-agnostic agentic AI design patterns, from tool use and ReAct to orchestrator-workers and memory, with when to use each, failure modes, a comparison table and a decision guide.

Eight agentic AI design patterns: tool use, ReAct, plan-and-execute, reflection, router, supervisor, evaluator loop and human checkpoint
Last updated Β· 14 min read Β· 3,159 words

Agentic AI design patterns are reusable ways of arranging model calls, tools, state and people so that an AI system completes multi-step tasks reliably. The core set is tool use, ReAct, plan-and-execute, reflection, routing, orchestrator-workers, evaluator-optimizer loops, human-in-the-loop checkpoints, deterministic workflows with LLM steps, and memory. Production systems combine two or three, and the classic mistake is choosing the most autonomous pattern when a constrained one would be cheaper, faster and easier to audit.

For the basics, start with what agentic AI is and how the agent loop works. These patterns are framework-agnostic; for how they map onto graph nodes and checkpointers, see LangGraph for enterprise AI.

Think in terms of who controls the next step

The useful question for any pattern is: who decides what happens next, the code or the model? Moving towards model control buys flexibility at the price of predictability, cost and audit effort.

code decides                              model decides
|------------|------------|------------|------------|
workflow    router     plan-and-     ReAct     multi-agent
+ LLM steps            execute                 (open-ended)

Reflection, evaluator loops, human checkpoints and memory are add-ons to whichever control pattern you choose.

Pattern 1: Tool use (function calling), the base layer

How it works. You describe tools to the model as names, descriptions and JSON schemas. The model replies with text or with a structured request to call a tool with specific arguments. Your code validates the arguments, executes the call and returns the result. MCP (Model Context Protocol) standardises how tools are exposed so many agents can share them.

When to use it, and when not to. Use it whenever the model needs fresh data or must take an action. Do not make something a tool when code already knows it must be called: if every request needs the same customer lookup, run it first and put the result in the prompt.

Failure modes. Wrong tool chosen because descriptions overlap; schema-valid but wrong arguments (right field, wrong employee ID); too many tools in one prompt; and tool output carrying injected instructions.

Enterprise considerations. Treat every tool as an API with an owner, an auth model and a blast radius. Separate read tools from write tools, run writes with the user's scoped permissions rather than a shared admin account, make writes idempotent, and log every call. Most of an agent's security lives here, not in the prompt.

Pattern 2: ReAct (reason and act)

How it works. ReAct interleaves reasoning and actions in a loop. The model reasons about what it knows and needs, takes one action (a tool call), observes the result, then reasons again with that observation in context. The key idea is the interleaving: each action is chosen in light of what the previous one returned, not planned upfront.

[goal]
  |
  v
[reason: what do I need next?] <-----+
  |                                   |
  +--> done? --yes--> [final answer]  |
  | no                                |
  v                                   |
[act: call one tool]                  |
  |                                   |
[observe: tool result] ---------------+
        (stop at max steps)

With modern function-calling APIs, the reasoning may be internal or a short visible note, and the action is a structured tool call.

When to use it, and when not to. Use ReAct for exploratory tasks where the next step depends on the last result: investigating an incident or diagnosing a failed deployment. Avoid it when the steps are known in advance, when each model call is expensive, or when an approver must see the full plan before anything runs.

Failure modes. Loops that repeat the same call with near-identical arguments; stopping after one weak observation; drift as tool output crowds out the original goal; and cost and latency that vary widely between runs.

Enterprise considerations. Cap steps and tokens, detect repeated calls, and trace every cycle. Keep a ReAct agent on read-only tools where possible and route writes through a checkpoint.

Pattern 3: Plan-and-execute

How it works. A planner call produces an explicit list of steps. An executor works through them, often with a cheaper model or plain code per step. When a step fails or reveals new information, a re-planner revises the remaining plan or stops.

[goal] --> [planner] --> plan: step1..stepN
                             |
                  +--> [execute next step]
                  |          |
                  |     [check result]
                  |      |         |
                  +--- ok      failed / new info
                         |         |
                     all done   [re-plan]
                         |
                    [final result]

When to use it, and when not to. Use it for longer tasks whose steps vary by request but should be visible: onboarding a vendor or migrating a set of resources. The plan is an artefact you can log and show to a human. Skip it for short tasks and for environments where the plan is stale after step one.

Failure modes. Plausible plans that miss a dependency; rigid execution of a plan early results have invalidated; and re-planning so often that it becomes an expensive ReAct loop.

Enterprise considerations. Validate plans against allowed step types: "delete" in a reporting task should fail validation, not reach a tool. For high-impact steps, put approval on the plan itself.

Pattern 4: Reflection and self-critique

How it works. The model produces a draft; a second call with a critic prompt reviews it against criteria and suggests fixes; the generator revises, for a bounded number of rounds.

When to use it, and when not to. Reflection helps when criteria are explicit: "every claim cites a retrieved document", "the SQL uses only these tables". Without new information, a model asked "are you sure?" about a fact it cannot check often repeats the error or changes a correct answer. Ground the critic in tests, schemas or retrieved sources.

Failure modes. The critic approves its own mistakes, over-corrects good content, and adds cost and latency for marginal gains.

Enterprise considerations. Prefer deterministic checks (validators, tests, policy rules) to self-critique, and prove on a test set that reflection improves outcomes before paying for it. Your AI agent evaluation suite should answer that, not intuition.

Pattern 5: Router

How it works. A classification step, usually a small model call with structured output, sends each request down a specialised path with its own prompt, tools and possibly model. An "unclear" route asks a clarifying question instead of guessing.

[request] --> [router: classify]
               |      |      |       |
            billing  tech  account  unclear
               |      |      |       |
            [path A][path B][path C][ask user]

When to use it, and when not to. Use a router when requests fall into distinct categories needing different tools, policies or models, such as sending simple FAQs to a small model and complex cases to a larger one. Skip it when every request follows the same path or categories overlap heavily.

Failure modes. Misrouting is silent: the wrong path gives a confident, wrong answer. Categories also drift as the business changes.

Enterprise considerations. Log each routing decision, measure routing accuracy on labelled requests, and use rules for categories that must never be misrouted, such as security incidents.

Pattern 6: Orchestrator-workers (supervisor)

How it works. An orchestrator breaks a task into subtasks at runtime, delegates each to a worker with its own tools, and combines the results. In the supervisor variant, it also decides after each worker returns whether another worker is needed.

            [orchestrator]
           /      |       \
    [worker A] [worker B] [worker C]
     logs       runbooks   ticketing
           \      |       /
           [orchestrator: combine]
                  |
              [result]

When to use it, and when not to. Use it when subtasks need different tool sets or permissions and you cannot know upfront which a request needs, such as an IT operations assistant needing logs, runbooks and change tickets. One agent with good tools is usually cheaper and easier to debug, so do not reach for this first. The multi-agent enterprise workflow project walks through a case where the split is justified.

Failure modes. Context lost in handoffs; workers duplicating or contradicting each other; multiplied cost; and errors that could originate in any hop.

Enterprise considerations. Give each worker least-privilege credentials, pass structured handoffs rather than free text, and trace the whole tree under one correlation ID. AI observability becomes mandatory.

Pattern 7: Evaluator-optimizer loops

How it works. A generator produces output; a separate evaluator scores it against an explicit rubric, ideally including non-LLM checks, and returns specific feedback; the generator revises until it passes or hits the iteration limit.

[task] --> [generator] --> draft
               ^              |
               |         [evaluator]
          feedback        (rubric + checks)
               |              |
               +---- fail ----+
                              | pass
                          [output]

When to use it, and when not to. Use it when "good enough" is checkable: SQL that must run and return the expected shape, code that must pass tests. A text-to-SQL data agent is a natural fit because the database itself is a reliable evaluator. Avoid it when latency matters more than marginal quality.

Failure modes. The generator learns to satisfy a weak rubric rather than the user, loops oscillate between two versions, and an uncalibrated LLM evaluator passes poor work.

Enterprise considerations. Calibrate LLM evaluators against human judgements, cap iterations, and store scores alongside outputs for review.

Pattern 8: Human-in-the-loop checkpoints

How it works. Before a high-impact step, the system prepares a proposal, persists its state and waits. A person approves, rejects or edits, and the system resumes from saved state with the decision.

When to use it, and when not to. Use checkpoints for irreversible or regulated actions: payments, access grants, production changes. Do not gate every step. Approval fatigue makes people click "approve" without reading, giving you the cost of a human without the safety.

Failure modes. Rubber-stamping because the request is a raw chat log instead of a clear summary; lost state when the approver replies days later; and actions executed despite rejection because the resume path was never tested.

Enterprise considerations. Persist state durably, key the checkpoint to a business identifier such as the ticket number, authenticate the approver and record who approved what. Pair checkpoints with automated AI guardrails so humans only see cases rules could not settle.

Pattern 9: Deterministic workflow with LLM steps

How it works. Ordinary code defines the sequence, branches and error handling. The model is called only where language understanding is needed: extracting fields, classifying, summarising, drafting. Each LLM step has narrow input, a structured output schema and validation.

[email in] --> [LLM: extract fields] --> [validate]
                                            |
                 [code: look up customer + rules]
                                            |
               [LLM: draft reply] --> [template check]
                                            |
                                     [send / queue]

When to use it, and when not to. This is often the right answer. If you can draw the flowchart, write it as code and use the model only for what code cannot do. You get predictable cost, easy testing and a readable audit trail. Move to a more autonomous pattern only when branches explode or the next step genuinely cannot be known in advance.

Failure modes. Rigidity: unexpected request types fall through to a generic error.

Enterprise considerations. It fits existing change management and audit practice, so it clears security review fastest. A common hybrid is a deterministic outer workflow with one bounded agentic step inside.

Pattern 10: Memory, short-term state vs long-term memory

How it works. Short-term state is everything about the current task: messages, tool results, the plan, approval status. It is checkpointed so the task survives restarts and long waits. Long-term memory persists across tasks: preferences, customer facts, past incident summaries. It sits in a database or vector store (see vector databases explained) and is retrieved when relevant.

AspectShort-term stateLong-term memory
ScopeOne task or threadAcross tasks, users or time
Typical storeCheckpoint table, session storeRelational table, vector index, key-value store
Main riskContext overflow, lost state on restartStale, wrong or poisoned memories; privacy
Key controlTrim or summarise; durable persistenceWrite rules, retention, deletion, access control

When to use it, and when not to. Every multi-step agent needs short-term state. Add long-term memory only for a clear benefit. Many enterprise agents should remember nothing beyond what systems of record hold; the CRM or ticketing system is the memory.

Failure modes and enterprise considerations. Memories can be wrong, stale after a policy change, or deliberately poisoned, and may hold personal data that must be deletable on request. Decide what may be written, who reads it, how long it lives, and log every write.

To build these patterns hands-on, with tool calling, RAG, LangGraph, MCP and evaluation in one curriculum, see Cloudsoft's AI, GenAI and Agentic AI course.

Comparison table

Ratings are relative and qualitative.

PatternComplexityCostLatencyReliabilityAuditability
Workflow + LLM stepsLowLowLowHighHigh
Tool use (single call)LowLowLowHighHigh
RouterLowLowLowMedium-highHigh
Plan-and-executeMediumMediumMediumMediumHigh (plan is visible)
ReActMediumVariableVariableMediumMedium (needs tracing)
ReflectionLow-mediumMedium-highMedium-highDepends on groundingMedium
Evaluator-optimizerMediumMedium-highHighHigh with real checksHigh
Orchestrator-workersHighHighHighMediumLow-medium
Human checkpointMediumLow (compute)Minutes to daysHighHigh
Long-term memoryMedium-highMediumLow-mediumMediumLow unless governed

Decision guide: choosing agentic AI design patterns

  1. Can you draw the flowchart? Build a deterministic workflow with LLM steps, and stop unless a question below forces you on.
  2. Do requests split into distinct types? Add a router, with rules for categories that must never be misrouted.
  3. Does the next step depend on what the last one found? Use a bounded ReAct loop for that part only.
  4. Is the task long, with steps that must be visible upfront? Use plan-and-execute with plan validation.
  5. Is good output checkable? Add an evaluator-optimizer loop grounded in real checks; use plain reflection only if tests show it helps.
  6. Do subtasks need different tools or permissions you cannot predict? Only then consider orchestrator-workers.
  7. Is any action irreversible or regulated? Put a human checkpoint before it, whatever else you chose.
  8. Must the agent recall anything across tasks? Add governed long-term memory; otherwise rely on systems of record.

Then evaluate. A pattern choice is a hypothesis until a scenario suite shows it beats the simpler alternative on task success, cost and safety. And never let the model decide policy: models propose; rules and people decide.

Illustrative example: choosing patterns for an IT access-request agent

Consider a GCC IT team in Hyderabad supporting a global bank. Employees ask in chat for things like "read access to the payments reporting database for the audit". The LangGraph guide shows one implementation as a graph; here the focus is the reasoning behind each pattern choice.

Sub-problemPattern chosenWhy not something more autonomous
Overall flowDeterministic workflow; one LLM step extracts system, role and justificationThe process (identify, classify, check policy, approve, grant, log) is known and auditors expect it fixed
Access vs password reset vs hardwareRouter with an "unclear" routeDistinct categories with different tools and policies
Which entitlement matches "payments reporting"?Bounded ReAct over read-only catalogue and directory toolsThe right lookup depends on what the first search returns; read-only keeps it safe
Policy decisionCode and rules, not the modelThe outcome must be explainable and identical for identical inputs
Approval for sensitive systemsHuman checkpoint with a summarised approval packAccess to financial data is high-impact; the manager must see a clear proposal
Grant and closeTool use with scoped, idempotent write toolsRetries must not create duplicate grants
MemoryShort-term state only; history lives in the ticketing systemAgent memory would create a second, ungoverned record of access decisions

Notice what is absent: no multi-agent supervisor and no self-reflection. Neither helps when decisions are rule-based and the risky step already has a human gate. Taking such a design into a customer's production environment is the work Forward Deployed Engineers do, and the focus of Cloudsoft's FDE PRO program.

FAQ

What are agentic AI design patterns?

They are reusable structures for AI systems that complete multi-step tasks, such as tool use, ReAct, plan-and-execute, reflection, routing, orchestrator-workers, evaluator-optimizer loops, human checkpoints, deterministic workflows and memory.

What is the ReAct pattern?

ReAct interleaves reasoning and actions. The model reasons about what it needs, calls one tool, observes the result and reasons again with that information, repeating until it can finish.

What is the difference between ReAct and plan-and-execute?

ReAct decides one step at a time from the latest observation. Plan-and-execute writes the full list of steps first and re-plans only when needed. The first adapts better; the second is easier to review upfront.

Does the reflection pattern make LLM outputs more accurate?

Sometimes. It works when the critic can check against something external, such as tests, schemas or retrieved sources. Without that grounding it often repeats errors, so measure it first.

When should I use an orchestrator-workers design?

When subtasks need different tools or permissions and you cannot know in advance which a request requires. Otherwise one well-tooled agent is cheaper and easier to debug.

Why is a deterministic workflow often the right answer?

Most enterprise processes have a known sequence. Writing it as code and calling the model only for language tasks gives predictable cost, simple testing and a clear audit trail.

What is the difference between short-term and long-term memory in agents?

Short-term state covers the current task and is persisted so the task can resume. Long-term memory persists across tasks, such as user preferences and needs rules for storage, access and deletion.

Ready to design agents that hold up in production rather than just in a notebook? Explore Cloudsoft's agentic AI training in Hyderabad, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo. Preparing for interviews? Try the agentic AI interview questions.

Share𝕏infβœ‰
EnrollWhatsAppCall us