New batches starting this week Β· Limited seats

Why Agentic AI Is Creating Demand for Forward Deployed Engineers

Once AI can act inside enterprise systems, every deployment needs someone to design tools, scope permissions, set approval rules and evaluate behaviour. This is why agentic AI is creating demand for Forward Deployed Engineers.

Diagram of an AI agent taking an employee request and calling ServiceNow, a knowledge base, Jira and GitHub, with human approval, least privilege and an audit trail
Last updated Β· 15 min read Β· 3,285 words

A chatbot that gives a wrong answer is an inconvenience. An agent that takes a wrong action, such as closing the wrong ticket, emailing the wrong customer or changing the wrong record, is an incident. Agentic AI is creating demand for Forward Deployed Engineers because once AI can act inside enterprise systems, every deployment needs someone to design the tools, scope the permissions, define approval rules and evaluate behaviour against that customer's own systems and policies, and that work cannot be fully packaged into a product. The model is the same across customers. The integration, identity and risk decisions are different every time.

If you want the broader career argument for why integration has become the bottleneck in enterprise AI, read why the future of AI engineering is Forward Deployed Engineering. This article focuses on one narrower question: what specifically about agents raises the bar, and what an agentic AI engineer actually has to build and decide.

Chatbot vs RAG assistant vs agent

The three are often lumped together as "GenAI apps", but they carry very different engineering burdens. The difference is not how clever the model is. It is what the system is allowed to touch.

DimensionChatbotRAG assistantAgent
Reads vs actsReads the prompt, writes textReads enterprise documents, writes textReads data and calls tools that change state in real systems
Blast radiusOne bad answer to one userA wrong or leaked answer, possibly from a document the user should not seeWrong records changed, wrong people notified, actions repeated at machine speed
Permissions neededUsually none beyond the model APIRead access to document stores, ideally filtered per userRead and write access to business systems, scoped per user and per action
Failure costLow; a human reads and discards itMedium; wrong advice may be acted on by a personHigh; the system itself acted, often before anyone looked
Evaluation difficultyJudge answer qualityJudge retrieval and answer faithfulnessJudge multi-step trajectories, tool choices, arguments and side effects

A RAG assistant can be wrong. An agent can be wrong and make the wrong thing happen. That one shift, from reading to acting, turns an AI feature into a piece of operational software that needs the same discipline as a payments integration.

What changes when AI can act

When an agent is given tools, a whole set of engineering concerns that never mattered for a chatbot suddenly become mandatory. None is exotic: it is standard distributed-systems and security practice, applied to a component that makes its own decisions.

Tool design

The model only sees a tool's name, description and input schema. A vague tool like update_ticket(fields) invites the model to change anything. A narrow tool like add_work_note(incident_id, note) or set_priority(incident_id, priority, reason) limits what can go wrong and makes intent auditable. Good agent tools are small, typed, validated on the server side and named for the business action, not the underlying API.

Least-privilege identity, per user

The tempting shortcut is one service account with broad rights that the agent uses for everyone. That means any user can, through the agent, do anything the service account can. The safer pattern is on-behalf-of access: the agent acts with a token derived from the signed-in user's identity, so it can only see and change what that user could. In a Microsoft estate this usually means Entra ID with delegated permissions and narrowly defined scopes; in other systems it means OAuth scopes or per-user roles mapped carefully onto tools.

Human-in-the-loop approvals

Not every action should be autonomous. A sensible design classifies tools by risk: reads run freely, low-risk writes (adding a note) run with logging, and high-risk writes (closing an incident, reassigning a P1, touching production) pause for a human to approve. The approval rule is a business decision, and it differs by customer.

Idempotency and retries

Model calls time out, networks fail and agents sometimes repeat a step. If "create ticket" is retried blindly, you get duplicate tickets. Write tools need idempotency keys or a check-before-create step, and retries need to distinguish "the call failed" from "the call succeeded but the response was lost".

Audit trails

For every action the system should record who the agent acted for, which tool it called, with what arguments, what the model's stated reason was, what came back and whether a human approved it.

Rollback

Some actions can be reversed (reopen a ticket, revert a field) and some cannot (an email already sent). The tool layer should know which is which, keep the prior state for reversible changes and route irreversible ones through approval.

Rate limits and cost per task

An agent in a loop can hammer an API or burn through model tokens. Enterprise systems have their own rate limits, and the customer will ask what each completed task costs. Step limits, per-user quotas, caching of repeated lookups and cost tracking per task are part of the design, not afterthoughts.

MCP: a standard plug, not a finished integration

The Model Context Protocol (MCP) is an open protocol, introduced by Anthropic in late 2024, that standardises how AI applications connect to tools and data. An MCP server exposes a set of tools, resources and prompts in a consistent format; any MCP-compatible client, such as an agent framework or an AI assistant, can discover and call them without custom glue for every pairing.

What MCP does not do is decide anything about a specific customer. For each system, someone still has to:

  • Choose which operations to expose as tools, and which to leave out entirely.
  • Write the server that maps those tools onto the customer's APIs, custom fields and workflows.
  • Wire authentication so the server acts with the right user's permissions, not a shared super-user.
  • Validate inputs, enforce business rules and return errors the model can understand.
  • Host, monitor, version and patch the server like any other production service.

A generic connector for a ticketing tool is a starting point. A server that understands that this customer's "Assignment group" field drives routing, that P1 incidents need a manager's approval to close and that one category is handled by an outsourced vendor is customer-specific engineering. That is where FDE-type work sits.

Orchestration: state, checkpoints and interrupts

Simple agents run a loop: the model picks a tool, the result goes back, repeat. Enterprise agents usually need more structure, and frameworks such as LangGraph model the agent as a graph of steps. Three ideas matter most, and they are easy to explain plainly:

  • State is the shared record of the task so far: the user's request, what has been retrieved, which tools were called and what they returned. Every step reads and updates it, so behaviour is explicit rather than buried in a long prompt.
  • Checkpoints save that state after each step to a store such as PostgreSQL. If the process crashes or a step fails, the run resumes from the last checkpoint instead of starting over and repeating side effects.
  • Interrupts pause the graph at a defined point, typically just before a high-risk tool call, and wait for a human. The approver sees the proposed action, approves, edits or rejects it, and the graph resumes from the checkpoint.

You do not strictly need LangGraph, but you do need these concepts.

An illustrative scenario: an IT service-desk agent at a GCC

Consider a global capability centre in Hyderabad that runs IT support for its parent company's offices worldwide. The service desk handles a high volume of repetitive incidents in ServiceNow: password and access issues, VPN problems, laptop requests, application errors. Leadership wants an agent that triages new incidents, finds the relevant article, adds a suggested resolution and, where safe, updates or resolves the ticket.

A production version looks like this:

Engineer / service-desk analyst
        |
        v
  Chat UI / ServiceNow panel
        |  (user signs in via Entra ID)
        v
  Agent service (FastAPI + LangGraph)
   |  state + checkpoints -> PostgreSQL
   |  traces -> Langfuse / OpenTelemetry
   |
   |-- LLM (Bedrock / Azure OpenAI)
   |
   |-- MCP: ServiceNow server
   |     read_incident, search_incidents
   |     add_work_note, set_priority
   |     resolve_incident  [needs approval]
   |
   |-- MCP: Knowledge base server
   |     search_kb (pgvector, per-user ACL)
   |
   '-- Approval interrupt -> analyst
          approve / edit / reject

The architecture is not the hard part. The decisions are. Here is what the forward deployed engineer on this engagement had to work out with the customer:

  • Scopes. Which ServiceNow tables and fields the agent may read and write, using the analyst's delegated identity. Security incidents and HR-related tickets were excluded entirely, so the tools never return them.
  • Approval rules. Adding work notes and suggesting categories run automatically. Changing priority requires a stated reason. Resolving any incident, and any action on high-priority or VIP-user tickets, pauses for analyst approval.
  • Eval set. A set of historical incidents, anonymised, labelled by senior analysts with the correct KB article, the right next action and the actions that would be wrong. This became the regression suite run on every prompt, tool or model change.
  • Guardrails. Server-side validation on every write tool, a cap on steps per task, refusal to act on tickets outside the user's assignment groups, idempotency keys on updates, and filtering of personal data before it reaches the model or the traces.
  • Monitoring. Traces of every run, dashboards for task completion, approval and rejection rates, tool errors, latency and cost per resolved ticket, and alerts when the rejection rate rises, which is often the first sign of a regression.
  • Rollout. Shadow mode first (suggestions only, no writes), then one assignment group, then wider, with the service-desk lead signing off at each stage.

None of these decisions would carry over unchanged to a different GCC, which might use Jira instead of ServiceNow, run a different identity setup, or have a works council that requires every automated change to be visible to the affected employee. That is the core reason this work needs engineers in the room. For a stage-by-stage view of how such a project moves from pilot to live, see how FDEs take AI from POC to production.

If you want to build this kind of system end to end rather than read about it, the AI Forward Deployed Engineer course at Cloudsoft includes a ServiceNow AI Agent via MCP project and an IT-Ops Multi-Agent Platform project among its five enterprise projects. Classroom in Ameerpet or live online; free demo via +91 96660 19191.

How agents are evaluated

Evaluating a chatbot mostly means judging the final answer. Evaluating an agent means judging the path it took and what it changed along the way. Useful measures, trended over time rather than reduced to one score:

  • Task success. Did the agent achieve the intended outcome for the case? For the service desk: right article found, right note added, ticket in the correct state.
  • Tool-call accuracy. Did it pick the right tool, with correct arguments, in a sensible order, and avoid unnecessary calls?
  • Unsafe-action rate. How often did it attempt an action outside policy, such as touching an excluded ticket type or trying to resolve without approval? Ideally the guardrails block these, but the attempts themselves are a signal.
  • Trajectory review. Humans reading full traces of a sample of runs, especially failures and approval rejections, to find patterns that metrics miss: loops, misread instructions, over-eager actions.
  • Approval outcomes. How often humans approve, edit or reject proposed actions. A rising edit rate tells you where the agent's judgement is drifting.

Teams use tools such as LangSmith or Langfuse to capture traces and run evaluation datasets, and Ragas for the retrieval part of the pipeline. The tooling matters less than the discipline: a labelled eval set built with the customer's experts, run before every change. Many of the reasons AI demos collapse in production trace back to skipping this, which is covered in why AI demos fail in enterprise production.

Why agents create demand for FDE-type engineers

A Forward Deployed Engineer is an engineer who works directly inside a customer's environment to make a technology deliver results there; the complete guide to the FDE role covers the definition in depth. Agents increase the need for this profile for a few structural reasons:

  • Every customer's systems differ. Same ticketing product, different custom fields, workflows, integrations and data quality.
  • Every customer's policies differ. What may be automated, what needs approval, what must be logged and who signs off are governance decisions, not product settings.
  • Identity is customer-specific. Mapping an agent's tools onto a customer's identity provider, roles and scopes requires access to and understanding of that environment.
  • Evaluation needs the customer's experts. Only their analysts know what the right action was for their incidents.
  • Risk sign-off is local. Security, compliance and operations teams want an engineer who can explain exactly what the agent can and cannot do.

Platforms will keep absorbing generic parts: hosted models, standard MCP connectors, orchestration frameworks, tracing. The customer-specific layer on top does not go away; it grows as agents are given more to do. That is why more job descriptions now combine agent frameworks, enterprise integration and security in one role, whether titled FDE, agentic AI engineer or AI solutions engineer.

Skills to build for agentic FDE work

AreaWhat to be able to do
Python and APIsBuild typed, validated services in FastAPI; call and wrap REST APIs of systems like ServiceNow, Jira and GitHub
Agent orchestrationModel agents as graphs with state, checkpoints and interrupts; handle tool errors and step limits
Tool and MCP designDesign narrow tools, build and secure MCP servers, version them safely
Identity and securityImplement on-behalf-of access, scopes and least privilege; threat-model prompt injection via retrieved content
RAGChunking, embeddings, pgvector, permission-aware retrieval
Evaluation and observabilityBuild eval sets, trace runs, review trajectories, track cost and failure modes
Cloud and deliveryContainerise, deploy on AWS, Azure or Google Cloud, use Terraform and CI/CD
Customer engagementRun discovery, turn policy into approval rules, explain risk to non-engineers

Practical next steps:

Common mistakes in agentic AI projects

  • One all-powerful service account. It makes the demo easy and the security review impossible.
  • Exposing raw APIs as tools. Giving the model a generic "call any endpoint" tool hands it the full surface area of the system.
  • Treating approvals as UI polish. If approval is not enforced in the orchestration and the tool server, a prompt change can bypass it.
  • No idempotency. Retries create duplicate tickets, duplicate emails and confused users.
  • Evaluating only final answers. An agent can reach the right answer through an unsafe path; trajectories need review.
  • Trusting retrieved content. Text in a ticket or document can contain instructions; treat it as data, never as commands.
  • Ignoring cost per task. Agent loops multiply model calls; without tracking, the bill surprises everyone.

Frequently asked questions

Is agentic AI just hype?

The marketing around it often is, but the underlying shift is real: models can now reliably call tools, so AI systems can take actions in business software rather than only produce text. The practical value depends on careful engineering of tools, permissions, approvals and evaluation, which is why many agent pilots stall when that work is skipped.

What is an agentic AI engineer?

An agentic AI engineer designs, builds and operates AI systems that use tools to complete multi-step tasks in real systems. Beyond prompting and frameworks, the role covers tool design, identity and least-privilege access, human approval flows, error handling, observability and evaluation of agent behaviour.

Do I need LangGraph to build AI agents?

No specific framework is required, but you need the concepts LangGraph makes explicit: shared state, checkpoints for recovery and interrupts for human approval. LangGraph is a common choice for enterprise agents because it models these directly, so it is worth learning well.

What is MCP?

MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 that standardises how AI applications connect to tools and data. An MCP server exposes tools and resources that any compatible client can discover and call. Someone still has to build, secure and operate the server for each customer system.

Are AI agents safe in enterprises?

They can be, if they are engineered for it. Safety comes from narrow tools, per-user least-privilege access, human approval for high-risk actions, server-side validation, audit trails, rollback where possible and continuous evaluation. An agent with broad credentials and no approvals is not safe regardless of the model.

How do FDEs test AI agents?

They build an evaluation set from real, anonymised cases labelled by the customer's experts, then measure task success, tool-call accuracy and attempts at unsafe actions, and review full trajectories of failures. They run this suite on every change and roll out in shadow mode before allowing the agent to write to live systems.

Agents are where the gap between an AI demo and an enterprise outcome is widest, and where engineers who can close it are most useful. If you want structured, hands-on practice building agents that act safely inside real enterprise tools, look at the Cloudsoft FDE PRO program: 12 weeks, 120+ hours live, 60+ labs, five enterprise projects and the GlobalBank capstone, with placement support until you're placed. Classroom beside Ameerpet Metro or live online; call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us