New batches starting this week Β· Limited seats

10 Projects Every AI FDE Engineer Should Build

Ten AI FDE projects framed as customer deliverables, each with the business problem, what to build, the stack, what it proves and a stretch goal, plus how to present them to hiring managers.

Grid of ten portfolio projects for AI Forward Deployed Engineers, from a knowledge assistant to an executive demo with ROI
Last updated Β· 15 min read Β· 3,299 words

The AI engineer projects that convince FDE hiring managers are not the ones with the cleverest prompt; they are the ones that look like a customer deliverable: deployed, secured, evaluated and documented so someone else could run them. Below are ten AI FDE projects, each framed as a business problem with a build plan, a stack, what it proves and a stretch goal. You do not need all ten. Two or three built properly will carry an interview loop further than a folder of notebooks.

If you are still mapping the role itself, start with what a Forward Deployed Engineer is. For the skills each project is meant to evidence, see the FDE engineer skills employers look for. This page is the build list.

What makes a project convincing to FDE hiring managers

An FDE interviewer reads your GitHub the way a customer reads a vendor's handover pack. They are asking one question: if I put this person in front of a bank or a hospital next month, would they ship something that survives production? Four properties answer it.

  • Deployed. It runs somewhere other than your laptop: a container on a managed cloud service, provisioned with infrastructure as code, behind a real API. A live URL or a recorded walkthrough of the running system both count.
  • Secured. Authentication, authorisation, secrets outside the repo, and a written note on prompt injection and data handling. Enterprise customers ask about these before they ask about accuracy.
  • Evaluated. A fixed test set, named metrics and a results table, including the cases that fail. "It works well" is not evidence; a faithfulness score and three documented failure modes are.
  • Documented like a deliverable. A README a customer's engineer could follow, an architecture diagram and a one-page business summary.

Each project covers part of the FDE chain; the summary table shows how to combine them.

Customer problem -> Discovery -> Data
  -> RAG / Agent -> Tools / MCP -> APIs
  -> Security -> Cloud -> Deployment
  -> Observability -> Evaluation
  -> Business outcome

1. Enterprise knowledge assistant with permission-aware RAG

Business problem. Consider an insurer whose underwriting, claims and HR teams each hold their own policy documents. Staff want one assistant, but HR documents must never surface in an underwriter's answer. Security reviewers check this first.

What you build. An ingestion pipeline that chunks documents and stores embeddings with metadata (department, classification, source, effective date). A retrieval layer that filters by the caller's group membership before similarity search, not after generation. Answers carry citations to the exact source chunk, and the assistant says "I don't know" when retrieval returns nothing relevant.

Stack. Python/FastAPI, PostgreSQL with pgvector, Amazon Bedrock or Azure OpenAI for embeddings and generation, LangChain for the retrieval chain, Microsoft Entra ID for sign-in and group claims, Ragas for evaluation, Docker.

What it proves. You understand that retrieval is an access-control problem as much as a relevance problem, and you can measure answer quality. If RAG fundamentals are new, read what RAG is first.

Stretch goal. Add document versioning so a superseded policy stops appearing in answers the moment a new version is ingested, and add an eval case that proves it.

2. Customer integration service (API, retries, idempotency)

Business problem. A retailer wants AI-generated product enrichment pushed into its order and catalogue systems. Those systems are slow, rate-limited and occasionally down. A naive integration double-writes records on retry or silently drops them.

What you build. A FastAPI service that accepts jobs, calls an LLM to enrich records, and writes results to a downstream API. Include idempotency keys so a retried request never creates a duplicate, exponential backoff with jitter, a dead-letter table for jobs that keep failing, and a status endpoint the customer can poll. Write contract tests against a mock that injects timeouts and server errors.

Stack. Python/FastAPI, PostgreSQL for job state and idempotency records, Docker, GitHub Actions for tests, Terraform to provision on AWS.

What it proves. Plain integration engineering, which is where much of an FDE's week actually goes. Unreliable integration is a common reason pilots stall.

Stretch goal. Add schema validation on LLM output with a structured-output contract, and route records that fail validation to a human review queue instead of the downstream system.

3. ServiceNow or Jira ticket agent via MCP with human approval

Business problem. Consider a GCC IT team in Hyderabad drowning in repetitive service desk tickets. They want an agent that reads a new ticket, finds similar resolved ones, drafts a categorisation and a reply, and updates the ticket, but nobody will let an agent write to the system of record unsupervised.

What you build. An MCP server that exposes a small set of typed tools (search tickets, get ticket, add work note, update category). An agent, built as a LangGraph state machine, plans the action and pauses at an approval node before any write. Read-only tools run freely; every write is logged with its approver.

Stack. Python, MCP (Model Context Protocol), LangGraph, a ServiceNow developer instance or a Jira cloud free tier, Azure OpenAI or Gemini, PostgreSQL for the audit log.

What it proves. You can connect agents to real enterprise tools with least-privilege scopes and a human in the loop. See what MCP is for the protocol basics.

Stretch goal. Swap the ticketing backend from ServiceNow to Jira by writing a second MCP server, with no change to the agent.

4. IT-ops multi-agent incident summariser

Business problem. During an outage, an on-call engineer at a bank flips between logs, metrics, recent deployments and the incident channel. Leadership wants a running summary every few minutes and a draft post-incident review afterwards.

What you build. A supervisor agent that delegates to specialists: a log analyst that queries recent error logs, a change analyst that lists recent deployments from GitHub, and a writer that assembles a timeline and a summary for executives. Agents share state through a typed graph, and each one has a step budget so a confused agent cannot loop forever.

Stack. LangGraph, Python/FastAPI, OpenTelemetry for traces, GitHub API, Docker on Kubernetes (EKS), Langfuse or LangSmith to inspect agent runs.

What it proves. You can decide when multiple agents are justified and when one would do, and you can bound and observe them. What agentic AI is covers the underlying patterns.

Stretch goal. Replay a recorded incident and score the generated timeline against a human-written one, so the summariser has a regression test.

Want to build systems like these with feedback from engineers who deploy them? Cloudsoft's FDE PRO program takes you from customer problem to deployed, evaluated system across five enterprise projects.

5. Secure banking or finance assistant with PII redaction and audit

Business problem. A bank's operations staff want an assistant that answers questions about account-servicing procedures and summarises customer complaints. Complaint text contains account numbers, phone numbers and names, which must not reach the model provider or the logs in clear text.

What you build. A redaction layer that detects and masks PII before any model call, with reversible tokens only where an authorised user needs the original. Input and output guardrails reject prompt injection attempts and off-topic requests. An append-only audit log records who asked what, which documents were retrieved and what was returned. Write a threat model covering the main risks and the control for each.

Stack. Python/FastAPI, Amazon Bedrock with guardrails, PostgreSQL for the audit trail, Microsoft Entra ID, Terraform with private networking on AWS, Docker.

What it proves. You treat security and compliance as design inputs rather than a final checkbox.

Stretch goal. Build a red-team eval set of injection and data-exfiltration prompts and report the block rate alongside answer quality.

6. Document intelligence pipeline (extract, validate, route)

Business problem. Consider a hospital's billing team that receives insurance pre-authorisation forms as scanned PDFs. Staff retype fields into the billing system and chase missing information by email.

What you build. A pipeline that extracts fields into a strict schema, validates them with business rules (dates in range, required fields present, policy number format), assigns a confidence to each field, and routes the document: auto-approve into the system, send to a human for review, or return to sender with a list of what is missing. A small review UI or API lets a human correct fields, and corrections are stored as future eval data.

Stack. Python, a multimodal model on Gemini, Bedrock or Azure OpenAI, Pydantic schemas, PostgreSQL, a queue-backed worker in Docker, GitHub Actions.

What it proves. Not every enterprise AI project is a chatbot, and extraction forces you to measure field-level accuracy.

Stretch goal. Report precision and recall per field on a labelled set, then show how changing the confidence threshold trades human workload against error rate.

7. Eval harness and CI regression gate for an LLM app

Business problem. A team changes a prompt to fix one complaint and quietly breaks five other answers. Nobody notices until users do.

What you build. Take project 1 or 5 and add a versioned test set covering normal questions, edge cases, questions with no answer in the corpus, and adversarial prompts. Run deterministic checks (format, citation present, refusal where expected) and model-graded metrics (faithfulness, answer relevance, context precision). Wire it into GitHub Actions so a pull request that drops a metric below its threshold fails the build, with the diff of changed answers posted as a report.

Stack. Ragas, LangSmith or Langfuse datasets, pytest, GitHub Actions, PostgreSQL to store results over time.

What it proves. You can make AI quality a release criterion rather than an opinion. The LLM evaluation guide explains the metrics and the limits of LLM-as-judge.

Stretch goal. Calibrate your judge by hand-labelling a sample and reporting how often the automated score agrees with you.

8. Observability and cost dashboard for an AI service

Business problem. An AI assistant is live, the cloud bill is rising, and the customer's platform team asks three questions: which feature costs the most, why some requests take far longer than others, and whether quality is drifting.

What you build. Instrument one of your earlier services with OpenTelemetry spans for retrieval, each model call and each tool call. Capture token counts, latency and model name per request, and tag spend by tenant and feature. Build a dashboard showing cost per request, latency percentiles, error rates and a sampled quality score. Add alerts for a cost spike and for a rise in "I don't know" responses.

Stack. OpenTelemetry, Langfuse or LangSmith, Kubernetes (EKS), Terraform, a managed dashboard on AWS, Azure or Google Cloud.

What it proves. Production ownership.

Stretch goal. Add a cheaper model route for simple queries, then use your dashboard and eval harness together to show the cost saving and its quality impact.

9. GitHub PR-review agent with guardrails

Business problem. A product company's platform team wants first-pass review comments on pull requests covering missing tests, risky infrastructure changes and leaked secrets, without an agent that can merge, push or approve anything.

What you build. A GitHub App or Actions workflow that reads the diff, applies deterministic checks first (secret patterns, Terraform plan risk rules), then asks a model for review comments constrained to a schema with file, line, severity and suggestion. Guardrails: read-only repository scope, a hard cap on comments per PR, no comments on files outside the diff, and content from the PR itself treated as untrusted input so an injected instruction in a code comment cannot redirect the agent.

Stack. GitHub Actions, GitHub API, Python, Bedrock, Azure OpenAI or Gemini, a small labelled set of past PRs for evaluation.

What it proves. You can scope an agent's permissions tightly and defend against indirect prompt injection in a workflow developers use every day. It also shows CI/CD fluency (see GitOps and GitHub Actions training).

Stretch goal. Measure precision of the agent's comments on your labelled PRs, and tune the prompt and rules until developers would not mute it.

10. Executive demo and ROI one-pager (customer-engineering artefact)

Business problem. The engineering works, but the sponsor has ten minutes with a CFO to decide whether the pilot gets budget. This project is not new code. It is the artefact that turns one of the projects above into a decision.

What you build. Pick your strongest project and produce three things. A recorded demo of about five minutes that opens with the business problem, shows one realistic user journey, then one failure handled gracefully. A one-page ROI summary that states the baseline process, the measured improvement from your pilot data, the assumptions behind any projection, the running cost from your dashboard and the risks. And a short "what we would do next" plan with stage gates.

Stack. Your earlier project, its eval report and cost data, a screen recorder and a document.

What it proves. Customer engineering: you can translate system behaviour into business language without overclaiming. This is the clearest difference between an FDE and a pure builder, as taking AI from POC to production explains stage by stage.

Stretch goal. Ask someone in finance or operations to watch the demo cold, write down every question they ask, and revise the one-pager to answer them.

Summary: projects, skills and difficulty

ProjectSkills shownDifficulty
1. Permission-aware knowledge assistantRAG, access control, citations, evaluationIntermediate
2. Customer integration serviceAPIs, retries, idempotency, contract testingIntermediate
3. Ticket agent via MCPAgents, MCP, tool scoping, human approvalIntermediate to advanced
4. IT-ops multi-agent summariserMulti-agent design, tracing, bounded autonomyAdvanced
5. Secure banking assistantPII redaction, guardrails, audit, threat modellingAdvanced
6. Document intelligence pipelineStructured extraction, validation, human routingIntermediate
7. Eval harness and CI gateTest sets, RAG metrics, CI/CDIntermediate
8. Observability and cost dashboardOpenTelemetry, cost attribution, alertingIntermediate to advanced
9. PR-review agent with guardrailsLeast privilege, indirect injection defence, GitHub ActionsIntermediate
10. Executive demo and ROI one-pagerDiscovery, communication, business framingBeginner code, hard to do well

A practical combination: projects 1 or 5 as the core system, 7 and 8 layered onto it, 3 as the agent piece, and 10 as the wrapper. Freshers can start with 2 and 6, which need less agent experience; the FDE engineer roadmap shows where each fits in a learning sequence.

How to present your projects

The build is half the work. How you package it decides whether an interviewer spends two minutes or twenty on your repository. The broader portfolio checklist lives in how to become an AI FDE; these are the four artefacts every project above should ship with.

A README a customer could follow

Open with the business problem in two sentences and who the users are. Then prerequisites, setup and configuration in steps that work on a clean machine, the environment variables and what each controls, how to run the evals, and a "known limitations" section.

An architecture diagram

One diagram showing components, data flow, trust boundaries and where each model call happens. Mark where authentication is enforced and where PII is redacted.

An evaluation report

Your test set size and composition, the metrics and why you chose them, a results table, three to five annotated failures and what you changed because of them. Include before-and-after numbers for one improvement.

A demo video

Three to five minutes, problem first. Show the happy path, one failure handled well, and a glimpse of the traces or dashboard. Link it at the top of the README so a busy reviewer sees it first.

Several of these projects map directly onto the FDE PRO curriculum: its five enterprise projects are a Customer Integration Service, an Enterprise Knowledge Assistant, a ServiceNow AI Agent via MCP, an IT-Ops Multi-Agent Platform and a Secure Banking AI Assistant, and the "GlobalBank" capstone is a simulated customer engagement that exercises the discovery, demo and business-summary skills from project 10.

Frequently asked questions

How many projects do I need for an FDE portfolio?

Two or three deep projects are usually enough, provided each is deployed, secured, evaluated and documented. Depth beats count.

Which project should a fresher build first?

Start with the customer integration service or the document intelligence pipeline. Both teach APIs, validation, testing and deployment without needing advanced agent design, and both give you a base to add RAG or agents to later.

Do I need paid enterprise tools like ServiceNow to build these?

No. ServiceNow offers personal developer instances and Jira has a free tier, which is enough for the ticket agent. For identity, a Microsoft Entra ID tenant can be set up for learning, and cloud free tiers or small budgets cover the rest if you shut resources down when you are not using them.

Should I use AWS, Azure or Google Cloud for these projects?

Pick one cloud and deploy everything there properly with Terraform. AWS is a common first choice, with Amazon Bedrock for models; Azure with Azure OpenAI suits teams targeting Microsoft-heavy enterprises. Showing you can map the design to a second cloud in your README is a useful bonus.

Are RAG chatbot projects still worth building?

A basic upload-PDF-and-ask chatbot adds little. A RAG system with permission-aware retrieval, citations, document versioning and an evaluation report is still one of the strongest portfolio pieces, because those are the problems enterprises actually face.

How do I show these projects in an interview?

Lead with the business problem, then walk the architecture diagram, then show the eval results and one failure you fixed. Keep the demo video ready to share. Interviewers will probe trade-offs, so be ready to say what you would change for a real customer.

Ready to build these as a guided engagement rather than alone? Learn Forward Deployed Engineering with FDE PRO: 12 weeks, 60+ labs, five enterprise projects and the GlobalBank capstone, in a classroom in Ameerpet beside Ameerpet Metro or live online, with placement support until you're placed. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us