This is a complete RAG project, walked through the way a Forward Deployed Engineer would deliver it to a customer: from the business problem to a measured return on investment. An enterprise knowledge assistant is only "done" when it answers from the right documents, respects who is allowed to see them, cites its sources, can be evaluated on a fixed test set and runs on monitored cloud infrastructure. The scenario is illustrative; the build plan at the end turns it into a RAG portfolio project you can defend in an interview.
For the fundamentals (embeddings, chunking, vector stores), read what RAG is and how it works; this article focuses on engineering decisions.
Business problem
Illustrative scenario. Consider a mid-size insurer whose operations team (claims processing, policy servicing and customer support back-office) works from a large and growing library of policy wordings, standard operating procedures (SOPs), underwriting guidelines, circulars and process notes. The documents live in SharePoint sites and folders of PDFs. Some are superseded but never deleted; some exist in two regional versions.
The symptoms will be familiar to anyone from a GCC or services delivery team in Hyderabad or Bengaluru:
- New joiners take weeks to become productive because finding "the right SOP" is tribal knowledge.
- Errors happen when someone follows a superseded procedure.
- Subject-matter experts (SMEs) answer the same questions repeatedly.
The customer wants fewer process errors, faster onboarding and less SME interruption, not "a chatbot". That framing, from AI demo to enterprise outcome, drives every decision below.
Requirements
Discovery with the operations head, team leads, an SME and information security produces three groups of requirements.
Functional
- Answer questions about procedures and eligibility rules from approved documents only.
- Every answer cites document, section and effective date, with a link.
- When the documents do not contain the answer, say so plainly and offer to raise a query to the document owner.
- Never answer from a superseded document unless the user asks about history.
Non-functional
- Users see only content their SSO groups already permit in SharePoint.
- Data stays in the customer's cloud account and approved region; no prompts or documents are used to train vendor models.
- An agreed p95 latency target, measured in production.
- Full audit trail: who asked what, which chunks were retrieved, what was answered.
Success metrics
| Metric | How it is measured | Owner |
|---|---|---|
| Answer correctness | Fixed test set scored by SMEs and an LLM judge | Engineering + SME |
| Faithfulness / citation accuracy | Automated eval on every release | Engineering |
| Correct refusal rate | Unanswerable questions in the test set | Engineering |
| Permission leakage | Must be zero on the access-control test suite | Security |
| Adoption and deflection | Weekly active users, SME queries before vs after | Operations |
Agreeing targets for these before writing code is the single most useful thing you can do.
Architecture
An offline indexing pipeline, an online query service, and identity, evaluation and observability across both.
SharePoint / PDFs
| (scheduled + change events)
v
Ingestion worker -> parse -> chunk -> embed
| |
v v
Doc registry (Postgres) pgvector + full-text
^
User (Teams / web) --SSO--> API (FastAPI)|
^ | hybrid retrieve
| | + ACL filter
| v
| Reranker
| v
+---- answer+cites <- LLM (Bedrock /
Azure OpenAI)
|
tool: create_ticket (MCP)
|
Traces -> Langfuse / OpenTelemetry
Key decisions and why:
- PostgreSQL with pgvector as the store. Chunks, metadata, ACLs and vectors live in one database, so permission filters and keyword search join naturally with vector search.
- Stateless API, separate worker, so a large re-index never slows user traffic.
- Model access behind an interface. The application calls a thin
llm_clientso switching between Amazon Bedrock and Azure OpenAI is configuration, not a rewrite.
Data
Most RAG quality problems are data problems.
Ingestion from SharePoint and PDFs
- Use the Microsoft Graph API with an app registration in Microsoft Entra ID, scoped read-only to the agreed sites. Pull files, metadata and permissions.
- Run a full crawl once, then incremental syncs using delta queries and a file hash, so only changed documents are re-chunked and re-embedded.
- Parse PDFs with a layout-aware parser that keeps headings, lists and tables. Flag scanned PDFs: OCR errors silently damage retrieval.
Chunking
SOPs and policy wordings are structured documents, so chunk on structure, not fixed token counts. Split on headings and numbered steps, keep a procedure's steps together where possible, and keep each table with its caption. Prepend a short context header to every chunk, for example Claims SOP > 4. Document verification > 4.2 Missing documents, so a passage like "escalate after the second reminder" still makes sense alone. Use a parent-child approach: retrieve small chunks, send the parent section to the model.
Metadata
| Field | Used for |
|---|---|
| doc_id, title, section_path | Citations and deep links |
| version, effective_date, status | Prefer current; exclude superseded |
| department, product_line, region | Filters and query routing |
| owner (person or team) | Routing "not found" tickets |
| allowed_groups | Permission-aware retrieval |
| source_hash, ingested_at | Incremental sync and audit |
Expect to find duplicates and conflicting versions. Do not resolve them silently in code: produce a report for the document owners. It is often the first value the customer sees.
LLM
Do not pick a model from a leaderboard. Pick it against criteria that matter to this customer, then confirm on your own test set:
- Grounded answering: does it stay within the provided context and refuse when the context is silent?
- Instruction following: does it reliably produce the citation format your code parses?
- Data residency and terms: available in the approved region through the customer's chosen platform (Amazon Bedrock, Azure OpenAI or Gemini on Google Cloud).
- Latency and cost per answer at realistic context sizes.
A practical pattern is two tiers: a capable model for answer generation and a smaller, cheaper model for query rewriting and classification. Choose the embedding model the same way, by retrieval quality on a sample of the customer's own questions. Remember that changing the embedding model means re-embedding the corpus, so decide early.
RAG
The query path has five stages, each with a concrete decision.
- Query understanding. Rewrite follow-up questions into standalone queries using the chat history ("what about for motor?" becomes "What is the missing-documents escalation process for motor claims?"). Extract filters such as product line when the user states them.
- Hybrid retrieval. Run vector similarity search and PostgreSQL full-text search in parallel, both restricted by the user's permitted groups and
status = 'current'. Merge with reciprocal rank fusion. Keyword search matters here because insurance operations are full of exact terms: form codes, clause numbers, product names. - Rerank. Pass the top candidates to a cross-encoder reranker (managed or self-hosted) and keep a small number of the best. It is usually the cheapest large quality gain.
- Prompt assembly. System instructions say: answer only from the numbered sources, cite each claim with its source ID, and reply "I could not find this in the approved documents" when the sources do not answer the question. Include each source's title and effective date.
- Citation validation. In code, check that every cited ID was in the retrieved set and drop or flag any that were not. Render citations as links to the exact SharePoint document and section.
Add a relevance threshold: if the best reranked score is below a tuned cut-off, skip generation and go straight to the "not found" path. This prevents many confident wrong answers.
Agent
Keep the agent light. Every autonomous step adds risk and evaluation work. The one agentic behaviour worth adding is: when the answer is not found, offer to raise a ticket to the document owner.
question -> retrieve -> score ok? --yes--> answer+cites
|
no
v
"Not found. Raise a query?"
|
user confirms
v
create_ticket(owner, question, ctx)
v
ticket ID shown to user
This is a small state machine, easily expressed in LangGraph or plain Python. The model proposes the ticket and drafts the summary; the user confirms; deterministic code calls the tool. The model never decides to create tickets on its own. Clusters of "not found" tickets also show SMEs exactly which documents are missing.
Tools
Define tools narrowly, with typed inputs and server-side validation:
| Tool | Inputs | Guardrails |
|---|---|---|
| search_documents | query, filters | User's ACL applied server-side, never passed by the model |
| get_document_section | doc_id, section_path | Permission re-checked on fetch |
| create_ticket | owner_team, question, summary, source_ids | Requires user confirmation; rate-limited per user |
| get_ticket_status | ticket_id | Only tickets the user raised |
User identity always comes from the authenticated session, never from model output, and every call is logged.
MCP/API
The ticketing integration (ServiceNow or Jira, whichever the customer uses) is exposed through an MCP server. MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data. As an MCP server, the ticket tools can later be reused by other assistants.
- The MCP server exposes only
create_ticketandget_ticket_status, nothing broader. - It uses a least-privilege service account and records the end user's identity in each ticket.
- Calls are idempotent: a client-generated request ID prevents duplicate tickets on retries.
The assistant itself is consumed through a plain REST API (POST /ask, POST /feedback, GET /health) so the web UI or a Teams bot can sit in front of it.
Building this end to end, including the ticket agent via MCP, is what Cloudsoft's AI Forward Deployed Engineer course covers: the Enterprise Knowledge Assistant is one of FDE PRO's five enterprise projects, and the ServiceNow AI Agent via MCP is another.
Security
Reviewers test permissions first. The design is permission-aware retrieval via SSO groups:
- Users sign in with Microsoft Entra ID (OIDC). The API validates the token and reads the user's group memberships.
- At ingestion, each chunk stores the groups allowed to read its source document, mirrored from SharePoint permissions.
- Every retrieval query filters on the intersection of the user's groups and the chunk's allowed groups, inside the database query, before ranking. Filtering after generation is too late.
- Permission changes in SharePoint propagate on the next sync; agree with security how stale that may be, and run a faster sync for revocations.
Beyond access control:
- Prompt injection: treat document text as data. User confirmation and server-side ACLs limit what an injected instruction can do.
- PII: redact personal data in logs and traces; keep raw questions only where the audit policy requires them.
- Secrets and network: a managed secrets store, and private connectivity to model endpoints.
The broader threat model is covered in AI security for enterprises.
Cloud
The right services depend on the customer's existing cloud estate.
- On AWS: the API and worker as containers on ECS or EKS, Amazon RDS for PostgreSQL with pgvector, Amazon Bedrock for generation and embeddings, S3 for raw document copies, Secrets Manager, and VPC endpoints so traffic to Bedrock stays private.
- On Azure: containers on AKS or Container Apps, Azure Database for PostgreSQL with pgvector, Azure OpenAI for models, Blob Storage, Key Vault and private endpoints. An insurer already on Microsoft 365 often prefers this route.
Provision everything with Terraform so environments are identical. Our AWS for AI engineers guide covers the services in depth.
Observability
Every request produces a trace with spans for query rewrite, each retrieval leg, rerank, LLM call and tool calls. Record chunk IDs and scores, tokens, latency, model ID and prompt version. Langfuse or LangSmith give LLM-specific views; OpenTelemetry carries the same spans into the customer's existing monitoring.
Dashboards: questions per day, "not found" rate by topic, thumbs-down rate, p95 latency and cost per answer. Alert on errors, latency and sudden not-found spikes, which often signal a broken sync. More on this in AI observability.
Evaluation
Build a fixed test set before tuning anything. Collect real questions from SMEs, each with the expected answer and supporting source section. Include:
- Questions needing exact terms (form codes, clause numbers).
- Questions whose answer changed between document versions.
- Unanswerable questions, where the correct behaviour is refusal.
- Access-control cases: the same question asked by users in different groups.
Retrieval metrics: did the expected source appear in the top results (recall at k, MRR)? Generation metrics: faithfulness to the retrieved context, answer correctness against the reference, citation accuracy and correct refusals. Tools such as Ragas automate several of these; calibrate any LLM judge against SME scores on a sample before trusting it. The method, and the limits of LLM-as-judge, are in our LLM evaluation guide.
Version the test set and add every production failure to it.
Deployment
CI/CD for an AI application tests behaviour, not just code. A GitHub Actions pipeline runs on every pull request:
- Lint, unit tests and type checks.
- Access-control test suite (any leak fails the build).
- Evaluation on the fixed test set; the build fails if key metrics drop below agreed thresholds versus the main branch.
- Build and scan the container image; push to the registry.
- Deploy to staging with Terraform-managed infrastructure; Argo CD syncs Kubernetes manifests if you run on EKS.
Prompts, chunking settings and model IDs are versioned configuration, so a prompt change goes through the same gate as a code change. Roll out in stages: an internal SME group first, then one team, then the wider operations floor, with a feature flag to fall back to search-only mode. CI/CD for AI applications covers pipeline design in detail.
ROI
ROI is a method agreed with the customer, not a number you announce. Agree inputs, measure a baseline, compare after launch. All figures below are hypothetical placeholders to show the arithmetic, not results.
| Input | Placeholder | How to get the real value |
|---|---|---|
| Users (U) | e.g. 100 staff | Pilot roster |
| Lookups per user per day (L) | e.g. 5 | Time-and-motion sample or survey |
| Minutes saved per lookup (M) | e.g. 4 | Timed tasks, before vs after |
| Working days per year (D) | e.g. 220 | HR calendar |
| Loaded cost per hour (C) | customer's figure | Finance |
| Run cost per year (R) | from billing | Cloud and model usage, support effort |
Annual time value = U Γ L Γ M Γ D Γ· 60 Γ C. Net value = time value β R β build cost amortised. With the placeholder inputs, U Γ L Γ M Γ D Γ· 60 gives roughly 7,300 hours a year to multiply by C. Discount it: not every saved minute becomes productive work. Report softer benefits separately: fewer errors from superseded SOPs, faster onboarding, and documentation gaps closed through not-found tickets.
Build it yourself: milestone plan
Use public or self-written documents, never real confidential files.
| Milestone | Deliverable |
|---|---|
| 1. Corpus and test set | A few dozen documents with versions and two permission groups; a test set of questions with expected sources |
| 2. Ingestion | Parser, structure-aware chunker, metadata, pgvector index, incremental sync by hash |
| 3. Baseline RAG | FastAPI /ask with vector search, citations, first eval scores recorded |
| 4. Quality | Hybrid search, reranker, relevance threshold, citation validation; eval scores compared to baseline |
| 5. Security | OIDC login, group-filtered retrieval, access-control tests |
| 6. Agent and MCP | Not-found flow with user confirmation; MCP server for Jira or a mock ticket API |
| 7. Ops | Tracing, dashboard, Terraform, GitHub Actions with an eval gate, deploy to AWS or Azure |
| 8. Value | ROI one-pager with placeholder method, demo video, README |
Repo structure
knowledge-assistant/
README.md
docs/
architecture.md
eval-report.md
roi-method.md
ingestion/
sharepoint_client.py
parsers.py
chunker.py
sync.py
app/
main.py # FastAPI routes
auth.py # OIDC, group claims
retrieval.py # hybrid + ACL filter
rerank.py
prompts/
llm_client.py
agent.py # not-found flow
mcp_server/
ticket_tools.py
eval/
testset.jsonl
run_eval.py
acl_tests.py
infra/terraform/
.github/workflows/ci.yml
docker-compose.yml
What to put in the README
- The business problem, labelled as an illustrative scenario.
- The architecture diagram and the reasons behind the main choices.
- How to run it locally with one command, and how to deploy it.
- Evaluation results, baseline versus final.
- How permissions are enforced and tested.
- Known limitations and next steps.
- A short demo video showing a refusal and a permission-denied case.
For how this fits alongside other portfolio work, see 10 projects every AI FDE should build.
Frequently asked questions
Is a RAG project still a good portfolio project?
Yes, if it goes beyond a basic chatbot. A RAG project that shows permission-aware retrieval, citations, a fixed evaluation set, tracing and a deployment pipeline demonstrates the engineering that enterprises need, which a notebook demo does not.
Which vector database should I use for an enterprise knowledge assistant?
PostgreSQL with pgvector is a strong default because chunks, metadata, permissions, vectors and full-text search live in one database, which keeps permission filtering simple. Move to a dedicated vector database only when scale or specific features require it.
How do I stop users seeing documents they are not allowed to see?
Store the permitted groups for each source document on every chunk at ingestion, read the user's groups from their SSO token, and filter retrieval inside the database query on those groups before ranking. Test it with an access-control suite that fails the build on any leak.
Do I need an agent in a RAG chatbot for company documents?
Not necessarily. Start with plain RAG. Add a light agent step only where it removes real friction, such as offering to raise a ticket to the document owner when the answer is not found, with the user confirming before any tool runs.
How do I evaluate a RAG project end to end?
Build a fixed test set of real questions with expected answers and source sections, including unanswerable and access-control cases. Measure retrieval recall, faithfulness, answer correctness, citation accuracy and correct refusals, and run the evaluation in CI on every change.
Can I build this project on AWS or Azure?
Either works. On AWS, Amazon Bedrock with RDS for PostgreSQL and containers on ECS or EKS is a common combination. On Azure, Azure OpenAI with Azure Database for PostgreSQL and AKS or Container Apps fits organisations already on Microsoft 365.
If you want to build this assistant with a trainer reviewing your design decisions, Cloudsoft FDE PRO runs it as one of five enterprise projects in a 12-week program, followed by the GlobalBank capstone, in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 for a free demo.



