This project walkthrough builds a claims triage AI agent for motor and health claims, the way a Forward Deployed Engineer would deliver it inside an insurer. A claims triage agent reads first-notice-of-loss submissions, checks the policy, chases missing documents, raises complexity and fraud signals, and routes each claim to the right adjuster queue with a summary, but it never approves, denies or settles a claim; adjusters decide. The scenario is illustrative. The build plan at the end turns it into a portfolio project you can defend in an interview.
Two companion articles set the context. Our overview of generative AI in insurance maps the use cases and controls across the industry, and the enterprise document intelligence project covers OCR, schema extraction and confidence scoring for claim packs. This article assumes that extraction layer exists and focuses on what happens next: where a claim goes, what is missing and who looks first.
Business problem
Illustrative scenario. Consider an Indian general and health insurer whose claims intake runs from a GCC in Hyderabad. FNOL arrives through the customer app, email, a call-centre form, agents and brokers, garages and network hospitals. A motor claim might be three photos and a message forwarded by an agent; a health reimbursement claim might be a long PDF of bills and a discharge summary with a covering email in Telugu.
Intake staff read each submission, look up the policy, decide which documents are missing, write to the customer, and pick an adjuster queue: motor own damage, motor theft, third-party injury, health reimbursement, health high-value, or the special investigation unit (SIU) referral desk. The pain points:
- Claims wait in a general inbox before anyone triages them, and urgent ones are not obvious.
- Misrouted claims bounce between queues, and every bounce costs days.
- Document requests go out piecemeal: the customer sends the FIR, then is asked for the RC copy, then the driving licence.
The business wants faster, consistent first handling with fewer bounces and fewer document rounds. From AI demo to enterprise outcome: a model that summarises a claim is a demo; a triage step adjusters trust enough to stop re-checking is the outcome.
Requirements
Discovery with claims operations, adjusters, SIU, compliance, the data protection officer and the policy admin team produces these requirements.
| Area | Requirement |
|---|---|
| Intake | Accept forms, emails, photos and PDFs from every channel; create one triage record per claim |
| Extraction | Structured fields per line of business (incident date, location, vehicle, hospital, admission dates, amounts claimed) |
| Policy check | Policy status, cover period, insured vehicle or member, and relevant covers, read-only from the policy admin system |
| Documents | Compare received documents to a checklist per claim type; request missing ones in one consolidated message |
| Flags | Complexity tier and fraud signals with reasons; signals, never accusations |
| Routing | Propose a queue and priority with a cited summary; adjusters can re-route |
| Hard boundary | No approval, denial, settlement, reserve setting or payment, by design |
| Privacy and audit | DPDP-aligned consent and purpose handling, health data minimisation, and a full audit trail per decision |
Success metrics, agreed up front and reported by region, language and channel: routing accuracy, missed-document rate, unnecessary-request rate, time to first adjuster touch and re-route rate.
Architecture
The model reads and writes; deterministic code checks policies, applies checklists and enforces the boundary.
FNOL: app / email / agent / hospital
|
v
[1] Intake + consent/purpose tag
v
[2] Extract fields (doc pipeline)
v
[3] Policy + coverage lookup (read)
v
[4] Document checklist (rules)
| missing? | complete
v |
[5] Request docs, wait |
+---------+---------+
v
[6] Complexity + fraud signals
v
[7] Route proposal + summary
v
[8] Claims system queue
v
Adjuster decides (always)
The orchestration is a state machine, not a free-roaming agent. Each claim moves through explicit states (received, extracted, policy-checked, awaiting documents, triaged, routed), and the waiting state can last days. That makes it a long-running workflow; the patterns for checkpointing and resuming are covered in durable, long-running AI agents.
Data
Four data sources matter, each with an owner:
- FNOL submissions: free text, structured form fields, photos, PDFs. Owned by claims operations.
- Policy admin records: status, cover dates, insured items or members, covers and add-ons. Read-only through an API owned by the policy admin team.
- Document checklists: which documents each claim type needs (for example, an FIR for motor theft, a discharge summary and final bill for health reimbursement). Owned by claims, versioned like code.
- Historical routed claims: past submissions with the queue an adjuster finally worked them in, documents later requested and re-route history. This becomes the evaluation set.
Two traps show up early. Historical routing labels are noisy: a claim may have been routed wrongly and then fixed, so use the final queue after re-routing as the label, and have senior adjusters adjudicate a sample. And history reflects past practice, including any regional unevenness, so it is a benchmark, not ground truth for fairness.
The triage record itself is a typed schema: claim reference, line of business, extracted fields with evidence links, policy check result, checklist result, flags with reason codes, proposed route and confidence.
LLM
The model does three narrow jobs:
- Understand the narrative. Turn "my car was hit from behind near Kukatpally signal, other driver ran away" into incident type, location, third-party involvement and whether injury is mentioned. Use structured output with a schema; the patterns are in our guide to function calling and structured outputs.
- Read images and scanned pages. Photos of a damaged vehicle, a hospital bill or a handwritten form need a vision-capable model. Use it to describe, not to estimate repair cost; multimodal AI in the enterprise covers where image understanding helps and where it misleads.
- Write the adjuster summary. A short, neutral summary that cites the source of every fact: which document, which page, which policy API response.
Language matters here. Submissions arrive in English, Hindi, Telugu, Tamil and other languages, and often in mixed or transliterated text. Test the model per language, and write the summary in the adjuster's working language with the original one click away.
RAG
Retrieval has a supporting role. The agent retrieves three things: the document checklist and triage guidelines for the claim type, the relevant cover description for the product version on the policy, and SIU's published signal definitions. Each is a small, versioned corpus, filtered by product, version and line of business before any semantic search. The model quotes guideline text; it never turns policy wording into a coverage decision.
Agent
The agent is a graph with fixed nodes, built for example in LangGraph, with the model choosing only within each node. This is the part that differs most from a chatbot.
Complexity tiers
Complexity is a rule-led score with model-extracted inputs: injury mentioned, third party involved, multiple vehicles, claimed amount above a threshold set by the business, hospital stay length, pre-existing condition disclosure questions, missing or conflicting dates. Rules assign the tier; the model only supplies the facts, each with its evidence.
Fraud signals, not accusations
This is the most sensitive output. Signals are defined by SIU, not invented by the model: incident shortly after policy inception, photo metadata inconsistent with the stated date, the same vehicle or bill number on another claim, a submitted image that matches an earlier claim, a mismatch between the registration on the policy and in the photos. The rules:
- Every signal has a code, a plain description and the evidence that triggered it.
- Wording is neutral: "registration in photo 2 does not match the policy record" rather than "suspected fraud".
- Signals are visible only to adjusters and SIU, never in any customer message.
- A signal can raise priority for human review; it can never block, delay payment of, or close a claim on its own.
- Referral to SIU is an adjuster's decision.
Human decision points
The agent proposes a route; the claims system places the claim in that queue, where an adjuster accepts or re-routes it. Low-confidence routes go to a triage supervisor queue instead. Every override is captured with a reason code and feeds evaluation. Our human-in-the-loop design guide covers how to stop reviewers rubber-stamping a proposal that is usually right.
Tools
The tool list is the real security boundary: if a tool does not exist, no prompt can make the agent use it.
| Tool | Access | Guardrail |
|---|---|---|
extract_submission | Calls document pipeline | Returns fields with evidence and confidence |
get_policy | Read-only | Masked fields; scoped to the claim's policy |
check_cover_period | Read-only | Deterministic; returns facts, not a coverage decision |
find_related_claims | Read-only | Matches on vehicle, member, bill number |
get_checklist | Read-only | Versioned per claim type |
request_documents | Write (comms) | Approved templates only, one consolidated request, rate-limited, customer's language |
propose_route | Write (triage fields) | Sets queue, priority, summary, flags; idempotent per claim |
What is missing matters as much: no tool to approve, deny, settle, set reserves, issue payments, close claims or change policy data.
MCP/API
Claims system integration is where most project time goes:
- Event in: the claims system emits an event when FNOL is registered; the triage service consumes it from a queue.
- Read-only policy access: a service account that can call only the policy lookup endpoints, with field-level masking so the agent never sees data it does not need.
- Narrow write back: the triage service writes only triage fields (queue, priority, summary, flags, document request status) through a dedicated endpoint, with idempotency keys so a retried event never creates duplicate requests.
- Customer communication goes through the insurer's existing notification service, so templates, opt-outs and delivery logs stay in one place.
MCP (Model Context Protocol, an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data) is useful for exposing get_policy, get_checklist and find_related_claims as a governed server that an adjuster copilot can reuse later. The write paths stay as plain, tightly scoped APIs called by orchestration code.
If you want hands-on practice with this kind of integration, from MCP tool servers to approval gates and evaluation in CI, the AI Forward Deployed Engineer course (FDE PRO) covers it in its ServiceNow AI Agent via MCP and Secure Banking AI Assistant projects, plus the GlobalBank capstone. Claims triage is not one of the five course projects; build it on your own afterwards.
Security
Health data and DPDP consent
Health claims carry diagnoses, procedures and lab results, so treat them at the highest internal sensitivity level. Practical controls, aligned with the approach in our DPDP Act guide for AI applications:
- Consent and purpose as metadata: tag each submission with the notice the claimant saw and the purpose (claim processing). The agent's pipeline accepts only records tagged for that purpose; compliance decides the lawful basis.
- Minimisation: triage needs dates, hospital, amounts and the nature of treatment, not full clinical detail. Send the model only the pages and fields needed; mask identifiers such as Aadhaar numbers on ID copies.
- Dependants and minors: health policies cover children, so confirm with compliance how their data is handled.
- Model access: private endpoints in an approved region, with contractual confirmation that inputs are not used for training.
- Prompt injection: document text is data, and since no tool can change a claim outcome, injected instructions cannot either.
Audit
Every triage produces an immutable audit record: hashes of the input files, extracted fields with evidence, policy API responses as returned, checklist version, rules fired, signal codes, model and prompt versions, the proposed route, and the adjuster's final action. When an auditor asks why a claim waited, the answer is a record, not a reconstruction. Logs and traces carry claim IDs and field names, not health details.
Cloud
On AWS: S3, Amazon Bedrock, SQS, a LangGraph service on EKS, RDS for PostgreSQL, KMS and VPC endpoints. On Azure: Blob Storage, Azure OpenAI, Service Bus, AKS, Azure Database for PostgreSQL and private endpoints. Stay in the region agreed with compliance, and autoscale on queue depth for surges after floods, when motor claims arrive in bulk.
Observability
Trace each claim end to end with OpenTelemetry, and log LLM calls in Langfuse or LangSmith with prompt version, model, tokens and latency. Dashboards claims leadership will actually use:
- Time from FNOL to routed, and to first adjuster touch, by line of business.
- Re-route rate by queue: the earliest sign that routing has drifted.
- Claims in "awaiting documents" by age, and how many needed a second request.
- Signal rates by code, region and language, watched for sudden shifts.
Evaluation
Evaluate against a de-identified set of historical claims routed by adjusters, stratified by line of business, channel, region and language, with deliberate edge cases: multi-vehicle accidents, theft, long hospital stays, mixed-language emails, poor photos.
| Metric | Definition |
|---|---|
| Routing accuracy | Proposed queue matches the adjuster-confirmed final queue |
| Missed-document rate | Claims where an adjuster later requested a document the agent did not flag, over all claims |
| Unnecessary-request rate | Documents requested that the claim did not need, or had already been sent |
| Summary faithfulness | Every fact in the summary traceable to a cited source, checked by sampling |
| Signal precision | Of raised signals, how many SIU reviewers judged worth raising |
Read the confusion matrix, not only accuracy: a high-value health claim in the routine queue is worse than the reverse, so weight errors by cost agreed with claims leadership. Agent-level testing is in our AI agent evaluation guide.
Fairness across regions and languages
A triage agent can be unfair without ever seeing a protected attribute. If extraction is weaker on Telugu emails, those claims miss documents more often and wait longer. If a signal leans on pin code or garage location, it can concentrate on particular districts. So:
- Report every metric per language, region and channel, and set a maximum acceptable gap.
- Run counterfactual tests: the same claim translated, or with only the city changed, should route the same way.
- Review each fraud signal's inputs for proxies, and drop or re-scope any that cannot be justified.
The methods are covered in AI bias and fairness testing.
Deployment
- GitHub Actions runs unit tests, tool contract tests and the evaluation set; a drop in routing accuracy or a rise in missed-document rate beyond agreed limits fails the build.
- Terraform and Argo CD, with checklists and signal definitions shipped as versioned configuration.
- Shadow mode first: the agent triages live claims while humans route as today, and you compare.
- Assist mode next: the proposal and summary appear to intake staff, who accept or change them. Document requests are drafted but sent by a person.
- Routing mode last, one line of business at a time, starting with simpler motor claims, with a kill switch back to manual triage.
ROI
Agree the method with finance before go-live, using measured baselines. Every input below is a hypothetical placeholder.
| Input | Placeholder |
|---|---|
| Monthly FNOL volume | [V] |
| Manual triage minutes per claim, before vs after | [T1], [T2] |
| Loaded cost per intake staff minute | [C] |
| Re-routes per month, before vs after | [R1], [R2] |
| Handling cost per re-route | [CR] |
| Extra document rounds avoided per month | [D] |
| Cost per document round | [CD] |
| Monthly run cost (models, cloud, support) | [RUN] |
Monthly benefit = V Γ (T1 β T2) Γ C + (R1 β R2) Γ CR + D Γ CD. Net = benefit β RUN, with build cost amortised over an agreed period. Report faster first adjuster touch and customer experience separately, not in rupees. The framing is in our enterprise AI ROI guide.
Build it yourself: milestone plan
- Write synthetic FNOL submissions in two languages for motor and health, plus a mock policy admin API in FastAPI.
- Build extraction to a typed triage schema with evidence links.
- Add the checklist engine and a consolidated, templated document request.
- Implement complexity rules and three signal codes with neutral wording.
- Wire the LangGraph workflow with a durable "awaiting documents" state.
- Label a routed evaluation set; report routing accuracy and missed-document rate per language.
- Add the audit record, tracing and a CI evaluation gate, then write a one-page ROI method.
In interviews, lead with the no-decision boundary, the fairness slices and the confusion matrix.
Frequently asked questions
What is claims triage AI?
Claims triage AI reads incoming claim submissions, extracts key facts, checks policy status, identifies missing documents, flags complexity and risk signals, and routes each claim to the right adjuster queue with a summary. Adjusters still make every claim decision.
Can an FNOL AI agent approve or reject insurance claims?
It should not. In this design the agent has no tool that can approve, deny, settle or pay a claim. It proposes a route and a summary, and accountable adjusters decide.
How do you stop fraud signals from becoming accusations?
Use signal definitions written by the investigation team, neutral wording with evidence, visibility limited to adjusters and investigators, and a rule that a signal can only raise priority for human review, never block or close a claim.
How is a claims triage agent evaluated?
Against de-identified historical claims routed by adjusters, measuring routing accuracy, missed-document rate, unnecessary-request rate, summary faithfulness and signal precision, broken down by line of business, channel, region and language.
How does the DPDP Act affect claims intake automation?
It shapes how consent and purpose are recorded, how much personal and health data the model sees, where data is processed, how long it is kept and how rights requests are handled. Compliance decides the lawful basis; engineering enforces it in the pipeline.
Is this a good AI claims processing project for a portfolio?
Yes, if you build it with synthetic data, a mock policy API, a clear no-decision boundary, a labelled evaluation set and fairness slices. Those choices show enterprise judgement, not only model skills.
Ready to build agents that regulated businesses can actually deploy? Explore the Cloudsoft FDE PRO program: 12 weeks, 60+ labs, five enterprise projects and the GlobalBank capstone, in our Ameerpet classroom beside Ameerpet Metro or live online, with placement support until you're placed. Call +91 96660 19191 for a free demo.



