Most enterprise AI work does not die because the model is weak. It dies in the gap between a promising proof of concept and a system that real users depend on. Forward Deployed Engineers take AI from POC to production by running the work as a sequence of stages, Discovery, Scoped POC, Pilot, Hardening, Production launch and Operate & expand, with an explicit, measurable gate between each one, and a named artefact that proves the gate was passed. Below is the playbook: activities, artefacts, gate criteria and traps for each stage.
If you are new to the role, start with what a Forward Deployed Engineer is. If you want the diagnostic view of why pilots stall, read why AI demos fail in enterprise production. Here we stay with process.
The stage-gate model at a glance
A stage-gate model makes each decision to invest further explicit, taken by the right people on evidence agreed in advance. Each gate has three possible outcomes: proceed, loop back, or stop. Stopping early is a legitimate, cheap result.
[1] Discovery
| GATE A: problem, metric, data access agreed
[2] Scoped POC
| GATE B: eval on fixed test set meets bar
[3] Pilot
| GATE C: real users, real data, value signal
[4] Hardening
| GATE D: security, SLOs, runbook signed off
[5] Production launch
| GATE E: stable in prod, handover accepted
[6] Operate & expand
+-> ROI review -> next use case (back to [1])
The stages map onto the FDE chain from customer problem to business outcome; they decide when each link must be solid enough to bear weight.
Stage 1: Discovery
Goal: agree what problem is worth solving, how success will be measured, and whether the data and access needed actually exist.
What the FDE does:
- Interviews the sponsor, front-line users and source-system owners. The question is not "where can we use AI?" but "which task costs time, money or risk today, and how do you know?"
- Shadows the current workflow to see where people search, wait or rework.
- Maps the data: where it lives, who owns it, its format and quality, what classification it carries and what access path is realistic (API, database read replica, export, nothing).
- Identifies constraints early: data residency, approved model providers and regions, identity provider, change process.
- Records a baseline for the sponsor's metric, even a rough hand-measured sample.
Deliverables: a one-page problem statement with success metrics and a baseline; a data access map; a stakeholder list; a constraints and risks register.
Gate A criteria: the sponsor has signed the problem statement; the metric is measurable (handling time, escalations avoided, accuracy against a reference); the data owner has confirmed access to a representative sample; a named person will judge the POC.
Common traps: starting from a technology ("we want an agent") instead of a task; accepting a vague metric like "productivity"; assuming data access that is later refused; never talking to the people who will use the tool.
Stage 2: Scoped POC
Goal: prove cheaply, on a narrow slice, that the core technical bet works. A POC answers one question: can this approach meet the agreed quality bar on this data?
What the FDE does:
- Builds the thinnest end-to-end path: ingestion, retrieval or tool calls, the model call and a simple interface. Often a RAG pipeline over PostgreSQL with pgvector, or a small agent with one or two tools.
- Builds the evaluation set before tuning: real user questions with reference answers agreed with subject-matter experts, including questions the system should refuse.
- Runs an eval harness (Ragas faithfulness and context metrics plus human grading) and iterates on chunking, retrieval, prompts and model choice.
- Captures cost and latency per request so the business case is not built on guesses.
Deliverables: the POC eval report (test set, metrics, failure categories, examples); a proceed/stop recommendation; an initial cost-per-interaction estimate.
Gate B criteria: the eval pass rate on a fixed test set meets the bar agreed with the sponsor in Discovery; no failure category is a blocker (for example, inventing policy that does not exist); unit cost is viable to the sponsor; the data owner will extend access for a pilot.
Common traps: tuning on the same questions you report on; demoing hand-picked questions to executives; building admin screens before the core bet is proven.
Stage 3: Pilot
Goal: learn whether the system creates value for real users doing real work, under real conditions. A pilot tests adoption and workflow fit, not just accuracy.
What the FDE does:
- Puts the system in front of a small, named user group, usually one team or one region, with single sign-on through the customer's identity provider (often Microsoft Entra ID) and document-level permissions respected.
- Integrates with at least one real system of record or workflow tool (ServiceNow, Jira, a CRM, a document store) rather than a copy.
- Adds tracing from day one so every answer can be inspected: retrieved chunks, tool calls, prompts, latency, tokens. LangSmith, Langfuse or OpenTelemetry-based tracing all work.
- Collects structured feedback in the interface, reviews it with users weekly, and feeds new failure cases into the eval set.
Deliverables: pilot plan (users, duration, success signals, exit criteria); architecture high-level design (HLD); pilot findings report with usage, feedback themes, measured change against the Discovery baseline, and incidents.
Gate C criteria: users return without being chased; the metric moves against the baseline on a method the sponsor accepts; no open high-severity incidents; the sponsor commits budget and IT and security start a formal review.
Common traps: a pilot with no end date that quietly becomes unsupported production; piloting only with enthusiasts; skipping tracing and being unable to explain a bad answer.
Stage 4: Hardening
Goal: turn a working pilot into a system the customer's security, operations and risk teams are willing to own. This is where most of the unglamorous engineering happens.
What the FDE does:
- Writes the low-level design (LLD): data flows, trust boundaries, failure modes, provider fallback.
- Rebuilds infrastructure with Terraform in the customer's approved account and region.
- Sets up CI/CD (GitHub Actions, Argo CD for GitOps on Kubernetes) with the eval suite as a release gate: a change that drops eval scores does not ship.
- Implements guardrails: PII handling, prompt-injection defences for retrieved content and tool outputs, least-privilege credentials for every tool.
- Defines service level objectives for availability, latency and answer quality, with alerting.
- Load-tests and builds the cost model at expected volume.
- Walks the security review: threat model, data handling, log retention, access control evidence.
Deliverables: LLD; security review pack; runbook; SLO document; cost model; rollback and recovery plan; updated eval report.
Gate D criteria: security sign-off with conditions documented; SLOs agreed with the team who will be paged; the runbook exercised (dry-run incident, rollback, provider outage); regression evals pass in CI; the budget owner accepts the cost model.
Common traps: treating security review as an end-of-project form; "click-ops" infrastructure nobody can rebuild; no plan for a provider outage; SLOs that ignore answer quality.
Stage 5: Production launch
Goal: release to the full intended audience safely and hand day-to-day ownership to the customer team.
What the FDE does:
- Plans a phased rollout (by team, region or user percentage) with feature flags and a tested rollback.
- Runs hypercare alongside customer support and operations.
- Trains users and support staff, including how to report a bad answer.
- Pairs with the customer engineers who will own the system until they have shipped a change themselves.
Deliverables: release plan and change record; user guide and support FAQ; handover document (architecture, repositories, environments, credentials ownership, open issues, known limitations, eval process).
Gate E criteria: the system meets its SLOs in production over an agreed observation window; support tickets are triaged by the customer team without FDE escalation for routine issues; the owning team formally accepts the handover.
Common traps: a "big bang" launch with no rollback; the FDE remaining the only person who understands the prompts and evals.
If you would rather practise this whole sequence than read about it, Cloudsoft's AI Forward Deployed Engineer course includes a weekly Customer Engagement Lab alongside the engineering sessions, in Ameerpet or live online.
Stage 6: Operate & expand
Goal: keep the system healthy, prove the business outcome and decide what comes next.
What the FDE does:
- Reviews traces, feedback and eval drift regularly; documents change and model versions are retired.
- Optimises cost and latency: caching, smaller models for simple routes, tighter retrieval.
- Runs the ROI review with the sponsor, comparing the post-launch metric with the Discovery baseline using the method agreed at Gate A.
- Identifies the next use case, often moving from answering questions to taking actions through tools or MCP servers, and restarts at Discovery.
Deliverables: operational review notes; ROI review; expansion proposal.
Gate criteria for expansion: the ROI review shows a result the sponsor is willing to defend to their own leadership; the owning team has capacity; the next use case has its own problem statement, not just "do more".
Common traps: declaring victory at launch and never measuring outcome; letting the eval set go stale.
Deliverables by stage
| Stage | Artefact | Owner | Audience |
|---|---|---|---|
| Discovery | Problem statement and success metrics | FDE with sponsor | Sponsor, leadership |
| Discovery | Data access map | FDE | Data owners, IT, security |
| Scoped POC | POC eval report | FDE | Sponsor, subject-matter experts |
| Pilot | Architecture HLD | FDE | Customer IT, architecture board |
| Pilot | Pilot findings report | FDE with pilot users | Sponsor, product team |
| Hardening | LLD and threat model | FDE | Customer engineering, security |
| Hardening | Security review pack | FDE with customer security | Risk and compliance |
| Hardening | Runbook, SLOs, cost model | FDE with operations | On-call team, budget owner |
| Production launch | Handover document | FDE | Owning customer team |
| Operate & expand | ROI review | Sponsor with FDE | Leadership |
Who does what: a RACI for a typical engagement
R = responsible, A = accountable, C = consulted, I = informed. "Product team" means the vendor's own product engineers behind the FDE. Treat it as a starting default.
| Activity | FDE | Customer sponsor | Customer IT / security | Product team | End users |
|---|---|---|---|---|---|
| Define problem and success metric | R | A | C | I | C |
| Grant data and system access | C | A | R | I | I |
| Build POC and eval set | R/A | I | I | C | C |
| Run pilot and collect feedback | R | A | C | I | R |
| Security review and sign-off | R | I | A | C | I |
| Infrastructure, CI/CD, SLOs | R | I | A | C | I |
| Production go/no-go | C | A | R | I | I |
| Feed product gaps upstream | R | I | I | A | I |
| Day-to-day operation after handover | C | I | A/R | I | I |
| ROI review | R | A | I | I | C |
Note "Feed product gaps upstream": an FDE who patches the same limitation for every customer without telling the product team is doing half the job.
Worked example: a bank's internal policy assistant
This is an illustration, not a real case. Consider a bank, with a GCC operations team in Hyderabad, whose staff need answers from KYC and operations policy manuals and today raise queries to a central policy desk.
Discovery. The real pain is the desk queue: simple lookups crowd out complex cases. The head of operations agrees the metric is the share of routine policy queries resolved without a desk ticket, plus accuracy against the policy text. The manuals sit in a document system with section-level access controls; some are scanned. Gate A decision: proceed, limited to retail operations policies in English; scanned documents and the credit policy for corporate lending are out of scope for now.
Scoped POC. The desk supplies real past queries with correct section references. The FDE builds a RAG pipeline with section-aware chunking and citations and refusal when no policy applies, tested on that fixed set. Early failures trace to outdated policy versions. Gate B decision: proceed, with a condition that only the current approved version of each policy is indexed, and every answer shows the clause it relied on.
Pilot. One region's operations staff get access through Entra ID single sign-on, with document permissions enforced at retrieval time. Langfuse traces show poor answers cluster on questions spanning two policies. Users ask for a one-click ticket when the assistant cannot answer, so the FDE adds a ServiceNow integration pre-filled with the question and retrieved context. Gate C decision: proceed to hardening; pilot-region desk tickets fell against the baseline on the agreed counting method, and usage held.
Hardening. Security asks where prompts are stored and who can read them; the answer goes into the security review pack. The FDE rebuilds the stack in the bank's approved region with Terraform on its Kubernetes platform, adds the eval suite to GitHub Actions and writes a runbook covering provider outages (fall back to search-only mode with citations). Gate D decision: approved with a condition: quarterly re-certification of the eval set by the policy owners.
Production launch. Rollout proceeds region by region. Two bank engineers ship a retrieval fix themselves before handover. Gate E decision: handover accepted once the system held its SLOs through the observation window.
Operate & expand. The ROI review compares desk ticket volume and time-to-answer against the Discovery baseline. The sponsor approves the next use case: scanned documents and the corporate lending manuals, which returns to Discovery because it brings new data owners and risk.
Every gate narrowed or conditioned scope rather than simply saying yes. FDE PRO's GlobalBank capstone simulates exactly this kind of full customer engagement, and its Secure Banking AI Assistant and Enterprise Knowledge Assistant projects cover the build.
How long does POC to production take?
Honestly, it depends; a fixed timeline quoted before Discovery is a guess. Engineering is rarely the longest part. The drivers are:
- Data access. If the data is behind an API you can call this week, the POC moves quickly. If it needs approvals, service accounts and firewall changes, that alone sets the pace.
- Security and risk review cycles. Banks, insurers and hospitals have review boards that meet on a schedule. Starting during the pilot is the biggest lever.
- Number of integrations. Each system of record or workflow tool adds design, access, testing and an owner to coordinate.
- Quality bar and risk of the use case. An internal search assistant tolerates more than an agent that updates customer records.
- Customer platform maturity. An existing Kubernetes platform, CI/CD and observability stack lets you plug in.
- Decision speed. An available sponsor shortens every gate.
Publish an estimate per stage with written assumptions, and re-estimate at each gate.
The FDE's toolkit across stages
| Capability | Typical tools | First needed |
|---|---|---|
| Application and APIs | Python, FastAPI | POC |
| Retrieval and agents | LangChain, LangGraph, PostgreSQL/pgvector, MCP | POC |
| Models | Amazon Bedrock, Azure OpenAI, Gemini | POC |
| Evaluation harness | Ragas, curated test sets, human grading | POC |
| Tracing and observability | LangSmith, Langfuse, OpenTelemetry | Pilot |
| Identity | Microsoft Entra ID, SSO, scoped service credentials | Pilot |
| Infrastructure as code | Terraform | Hardening |
| CI/CD and GitOps | GitHub Actions, Argo CD | Hardening |
| Runtime | Docker, Kubernetes (for example EKS) | Hardening |
The AI layer is needed first, but the platform layer gets you through Gate D. DevOps and SRE engineers already hold that half; to build it, Cloudsoft's Terraform training, Kubernetes training, DevOps course and SRE training cover it. To see how these skills sequence over a career, see the FDE engineer roadmap and the FDE skills employers look for.
Frequently asked questions
What is the difference between a POC and a pilot?
A POC tests whether a technical approach can meet an agreed quality bar on a fixed sample of data, usually without real users. A pilot puts the system in front of a small group of real users doing real work, integrated with real systems, to test whether it creates value and fits the workflow. A POC answers "can it work?"; a pilot answers "does it help?".
What should an AI POC prove?
An AI POC should prove that the core approach meets the success bar agreed with the sponsor on a fixed, representative test set, that its failure modes are understood and not blocking, and that the cost per interaction is viable. It should not try to prove scalability, user adoption or production security; those belong to later stages.
Who signs off production readiness?
Production readiness is normally signed off by the customer, not the FDE. Security and risk approve the security review, operations accepts the SLOs and runbook, and the sponsor makes the go or no-go call. The FDE supplies the evidence.
How do you measure the ROI of an AI assistant?
Agree the metric and the counting method during Discovery, measure a baseline before the system exists, then measure the same thing in the same way after launch. Typical measures are time per task, escalations avoided, error rates and adoption. Set the benefit against the full running cost, including model usage, infrastructure, support and eval upkeep.
What goes in an AI runbook?
An AI runbook covers health checks, what each alert means, steps for common incidents such as provider outages, latency spikes, stale retrieval or a drop in answer quality, how to roll back a prompt, model or index change, how to re-run evals and reindex data, escalation contacts, and where traces live.
Why use a fixed test set for AI evaluation?
A fixed test set lets you compare one version of the system with another fairly. If the questions change between runs, you cannot tell whether a better score means a better system or easier questions. Add new cases deliberately, version the set, and keep it separate from tuning examples.
The gap between a demo and an enterprise outcome is a process gap as much as a technical one. If you want structured, project-based practice running that process end to end, from discovery to handover, explore Cloudsoft FDE PRO: 12 weeks with five live sessions a week, in Ameerpet beside the Metro or live online. Call +91 96660 19191 for a free demo session.



