New batches starting this week Β· Limited seats

Project Walkthrough: Building a Secure Banking AI Assistant

A project walkthrough for an internal banking AI assistant for relationship managers and branch operations, covering entitlement-scoped customer lookups, PII redaction, data residency, audit, model and vendor risk review, red-teaming, evaluation and a private-network deployment.

Secure banking AI assistant: policy RAG, entitled account lookup, PII masking, full audit trail and private deployment
Last updated Β· 14 min read Β· 3,169 words

This walkthrough builds a banking AI assistant the way a Forward Deployed Engineer would deliver it inside a regulated bank, from the business problem to a measured return. A banking AI assistant is ready for production only when it answers policy questions from approved sources with citations, sees customer data strictly on behalf of the signed-in employee and never more, masks personal data everywhere it flows, keeps a complete audit trail, and runs inside the bank's private network after model, vendor and red-team review. The scenario is illustrative. The build plan at the end turns it into a generative AI banking use case you can defend in an interview.

The RAG mechanics match the enterprise RAG knowledge assistant project, so this article covers what banking changes: entitlements, data residency, auditability, vendor risk and the cost of a wrong policy answer.

Business problem

Illustrative scenario. Consider a bank with a large branch network and a relationship management team serving retail, priority and small business customers. Relationship managers (RMs) and branch operations staff field a steady stream of questions: "What documents does a non-resident customer need to open this deposit account?", "What is the current process for a nominee change on a joint account?", "Is this customer's fixed deposit due for renewal, and what did they hold last year?"

The answers are spread across policy notes, operations manuals, circulars that amend earlier circulars, and several core banking and CRM screens. The symptoms:

  • Staff ring the central operations desk for answers buried in documents.
  • Procedures are followed from a superseded circular, which leads to rework, customer complaints and audit observations.
  • Meeting preparation means opening several systems for a simple account summary.

The bank wants faster, correct answers for staff, fewer process errors and better-prepared meetings, with no new regulatory, privacy or reputational risk. That last condition shapes everything below: this AI in banking project is a study in moving from AI demo to enterprise outcome.

Requirements

A secure AI assistant for banks is defined by its controls, so discovery involves branch and RM leadership, the operations policy team, information security, the privacy office, compliance, internal audit, model risk, vendor risk and the owners of core banking, CRM and document systems. Bring the control functions in during week one, not at the end.

Functional

  • Answer product policy and procedure questions from approved, current documents only, citing document, section, circular number and effective date.
  • Show a read-only account summary (products held, masked balances as agreed, maturity dates, open service requests) for customers the user is entitled to see.
  • Draft, never send, customer-facing text such as a meeting summary or a document checklist, for the RM to review and send through existing channels.
  • Very limited writes: log an internal meeting note to the CRM and raise an internal service request, each after explicit user confirmation.
  • Say "not found in approved documents" rather than guess.

Non-functional and regulatory

Banking regulators set expectations on outsourcing, data localisation and auditability, which banks translate into internal policies. Do not interpret regulations yourself; consult the bank's compliance and legal teams and design to their written interpretation. Typical resulting requirements:

  • Customer data and prompts stay within the approved jurisdiction and the bank's controlled environment.
  • No customer data trains vendor models; vendor terms are reviewed and approved.
  • Every query, retrieval, data access, tool call and answer is attributable to a named employee and retained per the bank's record policy.
  • No public internet exposure; corporate network and SSO only.

Success metrics

MetricHow it is measuredOwner
Policy answer correctnessFixed test set scored by policy SMEsOperations policy team
Unsupported claims in answersFaithfulness checks on every release; agreed near-zero thresholdEngineering + model risk
Entitlement violationsMust be zero on the access-control suite and in production monitoringInformation security
PII leakage to logs or modelRedaction tests and log scans; must be zeroPrivacy office
Calls to the central operations deskVolume by topic, before and afterBranch banking

Architecture

Everything runs in a private network segment. Staff reach it through the intranet; it reaches models through private endpoints and bank systems through the existing API gateway.

RM / branch staff (intranet, SSO)
            |
            v
   Assistant API (FastAPI)
   |  PII redaction (in/out)
   |  policy guardrails
   v
 Orchestrator (LangGraph)
   |            |              |
 Policy RAG   Entitlement    Tools via MCP
 (pgvector)   service        (account summary,
   |            |             CRM note, SR)
   |            v              |
   |      API gateway (OBO token)
   |            v              v
   |      Core banking / CRM (read-heavy)
   v
 LLM via private endpoint, in region
   |
 Audit log (immutable) + traces

Key decisions:

  • Two separate paths. Policy questions go to RAG; customer questions go to live, entitled API calls. Customer data is never indexed.
  • The existing API gateway is the only route to core systems, so its controls and logging apply.
  • Redaction and audit are services in the request path, not optional middleware someone can switch off.

For how this fits a wider reference architecture, see enterprise AI architecture.

Data

DataSourceHandling
Product policies, operations manuals, circularsDocument management system, policy portalIndexed for RAG with circular number, version, effective date, supersedes, owner, audience
Customer account dataCore banking, CRM, via API gatewayNever indexed; fetched live, masked, scoped to entitlement
Historic staff queries to the operations deskEmail or ticket exportsRedacted; used for the test set and to prioritise topics

Circulars are the hard part. A new circular often amends one clause of an older one, so "current" is not a document-level flag. Record supersession at section level, and report conflicts to the policy owner instead of resolving them in code.

PII masking and redaction

  • Before the model: replace account numbers, customer IDs, PAN, Aadhaar, phone numbers, emails and addresses with typed placeholders such as [ACCOUNT_1], using pattern rules plus an entity recogniser for names, tested on Indian formats.
  • Tool results: masked identifiers by default. The model reasons over placeholders; the UI re-hydrates only what the user may see.
  • Logs and traces: redacted text only. Raw content the audit policy requires goes to a separate, access-controlled audit store.

LLM

Model choice in a bank starts with risk review. Shortlist models on the approved platform (Amazon Bedrock, Azure OpenAI or Gemini on Google Cloud) in an approved region with acceptable data terms, then compare on your own test set:

  • Grounded answering and refusal when the sources are silent.
  • Reliable tool calling with strict argument formats.
  • Consistent citation formatting that code can validate.
  • Latency and cost at realistic context sizes.

Model and vendor risk review

Expect to produce a model documentation pack: intended and out-of-scope uses, model identifiers, data flows, evaluation results, known limitations, monitoring plan and fallback. Vendor risk separately assesses the provider's data handling, retention, sub-processors, hosting location, incident notification and exit plan. Keep the model behind a thin client in versioned configuration, so a model change triggers a defined re-evaluation. If no hosted model passes review, the same design works with an open-weight model served inside the bank's environment, at more operational cost.

RAG

A banking chatbot with RAG follows the standard RAG query path: rewrite the question to stand alone, run hybrid retrieval (pgvector plus PostgreSQL full-text, because staff search by circular numbers and product codes), rerank, apply a relevance threshold, and generate with numbered sources. Banking-specific adjustments:

  • Effective-date filtering by default; historical versions only when the user asks "what was the rule in…".
  • Audience filtering: some manuals are for operations only, some for specific business lines. Store allowed_groups per chunk and filter inside the query.
  • Quote, then explain. For eligibility rules, document checklists and limits, quote the exact source sentence before summarising.
  • Code-side citation validation: drop any cited source not in the retrieved set, and refuse if nothing valid remains.

Agent

The agent is a small LangGraph state machine. Reads are cheap and frequent; writes are rare and gated.

question -> classify
   |
   +-- policy -> RAG -> cited answer
   |
   +-- customer lookup
   |      |
   |   entitled? --no--> "not in your book"
   |      | yes
   |      v
   |   get_account_summary (masked)
   |
   +-- draft for customer
   |      v
   |   draft -> RM reviews/edits -> RM sends
   |
   +-- write (note / SR)
          v
     preview -> user confirms -> tool

Human-in-the-loop for anything customer-facing

The assistant never contacts a customer. Anything a customer will read is a draft, labelled AI-generated with its sources, that the RM edits and sends through the bank's existing approved channel. The audit log records draft, final text and approver. Product recommendations, credit decisions, fee waivers and complaint responses are out of scope by design.

Tools

ToolTypeGuardrails
search_policiesReadUser's groups and effective date applied server-side
find_customerReadSearches only within the user's entitled customer set; returns masked matches
get_account_summaryReadOn-behalf-of token; masked identifiers; fields limited to an agreed list
get_open_service_requestsReadSame entitlement as the account summary
log_meeting_noteWritePreview and explicit confirmation; internal note only; idempotent
raise_service_requestWriteAllow-listed request types only; confirmation; routed to the normal operations queue

Deliberately missing: anything that moves money, changes customer details or KYC status, or messages a customer.

MCP/API

The banking tools are exposed through MCP servers. MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data. Each server wraps existing gateway APIs with narrow, typed operations.

Customer data access on behalf of the user, never broader

This is the control reviewers will probe hardest. The assistant must not have a powerful service account that can read any customer:

  1. The user signs in with the bank's SSO (for example Microsoft Entra ID). The assistant receives the user's token.
  2. For each tool call, the MCP server exchanges it for a short-lived, narrowly scoped on-behalf-of token for the API gateway (an OAuth token-exchange pattern).
  3. The gateway and core systems apply the bank's existing entitlements: this RM's assigned book, this branch's customers. The assistant inherits those rules rather than reimplementing them.
  4. The customer reference passed to a tool is validated against the user's entitled set server-side. The model can suggest a customer; it cannot widen the scope.

If the gateway cannot accept user-scoped tokens, the fallback is a service account plus a server-side entitlement check on every call, explicitly accepted by security.

Building this end to end is the Secure Banking AI Assistant project, one of five enterprise projects in Cloudsoft's AI Forward Deployed Engineer course. The program then closes with the GlobalBank capstone, a simulated customer engagement.

Security

  • Prompt injection: documents are data, not instructions; tool permissions in code and confirmed writes limit the damage.
  • Insider misuse: per-user rate limits, alerts on unusual lookup patterns, and the audit trail.
  • Output guardrails: block full identifiers, product advice, and rates or charges not drawn from a cited source.

Audit trail

For each interaction, write an append-only record: user and role, timestamp, redacted question, source IDs and versions, masked tool calls and results, model ID, prompt version, answer, drafts with their approved text, and confirmations. Store it with write-once retention, separate from application logs, so internal audit can reconstruct any answer later.

Red-teaming

Before launch and after major changes, red-team with security and experienced RMs: out-of-book lookups, extracting unmasked identifiers, advice on circumventing KYC, injected instructions in documents, and coaxing out rates or fees not in any source. Every successful attack becomes a permanent test case. The broader threat model is in AI security for enterprises, and the control framework that model risk and audit will expect is in enterprise AI governance.

Cloud

The deployment pattern is a private network with no public endpoint:

  • On AWS: EKS in private subnets, linked to the data centre over Direct Connect or VPN; Amazon Bedrock via a VPC interface endpoint; RDS for PostgreSQL with pgvector; Secrets Manager and KMS; an India region where localisation requires it.
  • On Azure: AKS with private endpoints for Azure OpenAI, PostgreSQL and Key Vault, over ExpressRoute.
  • Egress denied by default: outbound traffic allowed only to the model endpoint and the API gateway.

Provision everything, including network rules, with Terraform, so security reviews code rather than screenshots.

Observability

Trace each request with spans for redaction, retrieval, rerank, model call, entitlement check and tool calls, recording source IDs, latency, tokens, model ID and prompt version. Langfuse or LangSmith give LLM views; OpenTelemetry feeds the bank's monitoring and SIEM. Self-host tracing inside the private network: traces carry business context even after redaction.

Dashboards: not-found rate, thumbs-down rate, entitlement denials, guardrail triggers, latency and cost. Alert on any redaction failure, entitlement-denial spikes for one user and not-found spikes after a document sync.

Evaluation

Hallucination risk on policy answers is the core quality risk: a confident wrong checklist turns a customer away or opens an account without required documents. Build a versioned test set with policy SMEs:

  • Questions whose answer changed between circulars, with the current version as reference.
  • Unanswerable questions, where refusal is correct.
  • Entitlement cases: the same customer lookup by an entitled RM, an RM from another branch and an operations user.
  • Redaction and red-team cases.

Measure retrieval recall, faithfulness, answer correctness, citation accuracy and correct refusal. Entitlement, redaction and red-team cases are pass or fail: one failure blocks release. Calibrate any LLM judge against SME scores. See the LLM evaluation guide for method, and RAG evaluation metrics for the retrieval and faithfulness measures. Model risk will ask for these results at every change.

Deployment

GitHub Actions runs unit tests, the entitlement and redaction suites and the evaluation on every change, failing on any safety failure or regression. Images are scanned, signed and pushed to a private registry; Argo CD syncs to EKS. Prompts and model IDs are versioned configuration under the same gate, and model changes follow the model-risk change process.

Roll out in stages:

  1. Policy Q&A only for the central operations desk, who will spot errors.
  2. Pilot branches, policy Q&A plus read-only account summaries.
  3. Drafts and limited writes after internal audit reviews the audit trail.

ROI

ROI is a method agreed with the business and finance, measured against a baseline. All values below are hypothetical placeholders to show the arithmetic, not results or benchmarks.

InputPlaceholderWhere the real value comes from
Staff users (U)e.g. 500Pilot and rollout roster
Lookups per user per day (L)e.g. 4Sampling and assistant usage logs
Minutes saved per lookup (M)e.g. 5Timed tasks, before vs after
Working days per year (D)e.g. 240HR calendar
Loaded cost per hour (C)bank's figureFinance
Run and assurance cost per year (R)from billing and budgetsCloud, model usage, support, periodic reviews and red-teaming

Annual time value = U Γ— L Γ— M Γ— D Γ· 60 Γ— C. Net value = time value βˆ’ R βˆ’ amortised build cost. With the placeholders, U Γ— L Γ— M Γ— D Γ· 60 gives 40,000 hours a year to multiply by C; discount it, since not every saved minute becomes productive work. Report risk benefits separately: fewer procedures followed from superseded circulars, fewer audit observations on process errors.

Build it yourself: milestone plan

Use a fictional bank, self-written policies and synthetic customers. Never use real customer data.

MilestoneDeliverable
1. FixturesFictional circulars with supersession; mock core banking API with RMs, branches, synthetic customers
2. Policy RAGHybrid search, effective-date filter, citation validation
3. RedactionPII detection for Indian identifier formats, placeholders, redacted logs
4. Entitlements and MCPSSO, on-behalf-of tokens, MCP servers for lookups and writes
5. Agent and HITLLangGraph flow with draft review and confirmed writes
6. Audit and red-teamAppend-only audit log; red-team test suite
7. Evaluation and opsEval gate in GitHub Actions, tracing, private-subnet Terraform
8. Assurance packModel documentation, ROI one-pager, demo of a denied out-of-book lookup

Keep auth (SSO and on-behalf-of exchange), redaction, audit and the entitlement, redaction and red-team suites as separate, visible modules: interviewers look for them first. See 10 projects every AI FDE should build for how this fits a wider portfolio.

Frequently asked questions

What makes a banking AI assistant different from a normal RAG chatbot?

The controls around it. A banking AI assistant must enforce existing customer entitlements, mask personal data before it reaches the model or logs, keep an attributable audit trail, run in a private network within approved regions, and pass model, vendor and security reviews before it reaches staff.

How does the assistant access customer data securely?

It calls the bank's API gateway with a short-lived token obtained on behalf of the signed-in employee, so the bank's existing entitlements decide what is visible. Customer data is fetched live, masked and never indexed in the vector store, and the model cannot widen the scope of a lookup.

Should a banking AI assistant talk directly to customers?

Not in this design. Anything customer-facing is a draft that the relationship manager reviews, edits and sends through the bank's existing channels. The audit log records the draft, the final text and the approver.

How do you reduce hallucination risk on policy answers?

Retrieve only current, approved documents, quote the exact source sentence before summarising, validate citations in code, refuse below a relevance threshold, and evaluate on an SME-built test set that includes changed circulars and unanswerable questions on every release.

Which regulations apply to an AI assistant in a bank?

That depends on the jurisdiction and the bank's own interpretation. Banking regulators generally have expectations on outsourcing, data localisation and auditability, so engineers should design to the written requirements of the bank's compliance and legal teams rather than interpreting regulations themselves.

If you want to build a secure banking AI assistant with a trainer reviewing your entitlement, redaction and audit design, explore Cloudsoft FDE PRO: 12 weeks, five enterprise projects and the GlobalBank capstone, in our Ameerpet classroom beside Ameerpet Metro or live online, with placement support until you're placed. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us