This project walks through how a Forward Deployed Engineer would build an HR AI agent, from the business problem to an ROI method. An HR AI agent is ready for production only when it answers policy questions for the employee's own location, grade and employment type, shows only the employee's own data, asks for confirmation before it takes any action, and passes grievances, harassment, health and performance concerns to a human every time. The scenario is illustrative. The build plan at the end turns it into a portfolio project you can explain in an interview.
The generic retrieval mechanics (chunking, hybrid search, reranking, citation checks) are already covered in our enterprise RAG knowledge assistant project. This article covers what is different about HR: personalisation, effective dates, sensitive routing, privacy and tone.
Business problem
Illustrative scenario. Consider a company with several thousand employees across Indian locations: a head office in Bengaluru, a GCC delivery centre in Hyderabad, sales offices in Mumbai and Delhi NCR, and a plant in Pune. The workforce includes permanent staff at several grades, fixed-term contract employees and interns. An HR shared-services team runs a helpdesk through email, an intranet portal and a ticket queue.
The helpdesk's problems are familiar:
- Most tickets are repeat questions: "How many casual leaves do I get?", "When is the payslip released?"
- Answers depend on who is asking. The leave policy has a different holiday list at each location. Travel limits depend on grade. Contract staff have different benefits.
- Policies change yearly, and employees quote last year's PDF.
- Serious concerns land in the same queue as routine questions.
HR wants faster answers, fewer repeat tickets and a safe channel for serious matters. That shift from AI demo to enterprise outcome shapes every decision below.
Requirements
Discovery covers HR operations, HR business partners, payroll, the Internal Committee that handles POSH complaints, legal, information security and employee representatives.
Functional
- Policy questions: answer from approved policies only, personalised to the employee's location, grade and employment type, and citing the policy, clause and effective date.
- Personal questions: answer leave balance, payslip release date and reporting-manager details from the employee's own HRIS record only.
- Simple requests: apply for leave and raise an HR ticket. Each is an action that the employee confirms before anything is written.
- Sensitive topics: grievances, harassment, health, mental-health concerns, performance disputes and disciplinary matters are never answered by the AI. They go to a named human channel.
- Languages: employees can ask in their preferred Indian language or English.
Non-functional
- An employee sees only their own data. HR staff see more according to their role and scope.
- Employee data stays in the company's cloud account and approved region, and is not used to train vendor models.
- Every answer and action is auditable: question, policy version used, tool called.
Success metrics
| Metric | How it is measured | Owner |
|---|---|---|
| Personalised answer correctness | Test set of questions asked as different employee profiles | Engineering + HR policy owners |
| Sensitive-routing recall | Sensitive test cases sent to a human; target agreed with HR and legal | HR + Engineering |
| Data leakage | Must be zero on the cross-employee access suite | Security |
| Action accuracy | Leave requests and tickets created exactly as confirmed | Engineering |
| Ticket deflection and adoption | Repeat tickets before vs after; weekly active employees | HR operations |
Architecture
The design has a router in front, three paths behind it, and identity, audit and evaluation running across all of them.
Employee (Teams / portal) --SSO--> API
|
load profile from HRIS
(location, grade, type, role)
|
Intent + risk router
/ | \
sensitive policy Q personal / action
| | |
human RAG over HRIS tools
handoff policies (own data only)
(HRBP/IC/ filtered by |
EAP) profile+date confirm step
| | |
+------> answer / ticket <-+
|
audit log + traces (Langfuse)
Key decisions:
- The router runs before retrieval, so a harassment disclosure never reaches the policy-answering prompt.
- The profile comes from the HRIS, not the conversation. Typing "I am a band 7 in Mumbai" changes nothing.
- Policies and personal data stay apart. Policy text sits in PostgreSQL with pgvector; personal data is never embedded, only fetched live through scoped tools.
Data
Policy corpus
Leave, travel, benefits, insurance, code of conduct and holiday lists, as PDFs and intranet pages. Chunk by clause, keep each clause's applicability statement ("Applies to: permanent employees, all locations except plant") in the chunk, and keep grade-wise limit tables whole.
Applicability and versioning metadata
| Field | Used for |
|---|---|
| policy_id, clause_path, title | Citations |
| applies_location, applies_grade, applies_emp_type | Personalised filtering (list or "all") |
| effective_from, effective_to, version, status | Picking the version in force on a given date |
| supersedes | Explaining "this changed from last year" |
| owner | Routing unanswered questions |
Effective dates are where HR assistants most often go wrong. A question about "my leave next month" must use the policy in force on that date, possibly a published version not yet in force. Filter on the relevant date, not today by default, and tag changed clauses so the agent can say "This changed from 1 April; the earlier rule wasβ¦".
Applicability tags are rarely written cleanly in the source. Extract them with an LLM-assisted pass, have the policy owner review every tag, and report contradictions (two holiday lists for one location) back to HR rather than fixing them in code.
Employee data (HRIS)
Location, grade, employment type, manager, leave balances and payroll calendar come from the HRIS API, and only what one answer needs enters the prompt.
LLM
Choose the model by testing it on HR questions. Accept it only if it meets these criteria:
- It stays strictly within the retrieved clauses and refuses when they are silent, because a guessed leave rule creates an HR dispute.
- It writes well in the Indian languages your employees use, judged by native speakers.
- It follows tone instructions consistently.
- It is available in the approved region on the company's platform: Amazon Bedrock, Azure OpenAI or Gemini on Google Cloud.
Use a small, fast model for intent and risk classification and a more capable one for answers. The classifier is safety-critical and gets its own evaluation set.
RAG
The retrieval pipeline matches the knowledge assistant walkthrough, with two HR-specific additions:
- Profile filtering inside the query. Each retrieval filters on
applies_location,applies_gradeandapplies_emp_typematching the employee's HRIS profile, or "all". It also filters on the effective window covering the date in question. The filter runs in the SQL before ranking, so the model never sees the travel limits for another grade. - Applicability in the answer. The prompt asks the model to say which rule applies and why: "As a permanent employee at the Hyderabad office, you getβ¦". This makes personalisation visible and lets the employee spot a wrong HRIS record.
For HR staff answering on behalf of others, the profile filter becomes an explicit parameter, limited to their scope and logged.
Multilingual questions
Employees may write in Hindi, Telugu, Tamil, Kannada, Marathi or a mix with English. If policies exist only in English, translate the query for retrieval, answer in the employee's language, keep the cited clause title in English and state that the English policy is authoritative. Add languages only after native-speaker reviewers approve quality.
Agent
The agent is a small, explicit state machine, which fits well in LangGraph. The important design choice is what the agent will not do.
message -> risk check --sensitive--> handoff
|
routine
v
intent: policy | personal | action
|
draft request (dates,
type, balance check)
v
"Apply 2 days CL,
12-13 Nov? Confirm"
|
user confirms
v
HRIS / ticket tool call
v
reference shown to user
Sensitive routing by design
The risk check looks for grievances, harassment, discrimination, health and mental-health concerns, performance or disciplinary disputes, and signs of distress. When it triggers, the agent gives no policy summary, advice or probing questions. It acknowledges the message, explains who will handle it and how confidentially, and offers the right channel: the HR business partner, the Internal Committee process, or the employee assistance programme. With consent it creates a restricted case; urgent safety signals show emergency contacts at once.
The router is tuned to over-route: a routine question sent to a human costs minutes, while a harassment disclosure answered with a policy summary destroys trust. When unsure, route.
Actions with confirmation
The model drafts leave applications and tickets; deterministic code executes them only after the employee confirms a draft showing type, dates and balance before and after. The HRIS approval workflow stays in place: the agent submits requests, never approves them.
Tools
| Tool | Scope | Guardrails |
|---|---|---|
| search_policies | Profile and date filters | Filters taken from the HRIS profile, never from model output |
| get_my_leave_balance | Caller only | Employee ID from the session token |
| get_payroll_calendar | Caller's pay group | Dates only; no salary figures shown in chat |
| apply_leave | Caller only | Confirmation required; idempotency key; balance re-checked |
| create_hr_ticket | Caller only | Confirmation required; category from a fixed list |
| create_sensitive_case | Restricted queue | Minimal content; visible only to the assigned HRBP or the Internal Committee |
No tool accepts an employee_id from the model; identity comes from the session, closing off "show me my colleague's balance" attacks.
MCP/API
The HRIS and the HR ticketing system (ServiceNow HR Service Delivery, Jira Service Management or an in-house tool) are exposed as MCP servers. MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data. Each server exposes only the narrow tools above, uses a least-privilege service account (or on-behalf-of access where the HRIS supports it) and records the end user's identity for audit. See what MCP is for the protocol, and the ServiceNow AI agent project for a fuller MCP integration walkthrough.
Cloudsoft's AI Forward Deployed Engineer course does not include an HR agent project. Its Enterprise Knowledge Assistant and ServiceNow AI Agent via MCP projects teach the two building blocks this design combines: permission-aware RAG and confirmed tool actions through MCP.
Security
- Own data only. Employees sign in with Microsoft Entra ID; every personal tool resolves the employee from the token.
- Role-based HR access. HR roles (business partner, payroll, operations) come from directory groups, scoped to a business unit or location, with queries logged for review.
- Minimal logging. Redact personal details from traces. Sensitive-case conversations keep only what the case needs, under restricted access and a retention period agreed with legal, in line with India's data protection law.
- Prompt injection. Treat policy text and pasted content as data. Confirmation steps and identity-bound tools limit what an injected instruction can do.
The wider threat model is covered in AI security for enterprises. Approval, accountability and policy ownership are covered in enterprise AI governance.
Bias and fairness
The agent never ranks, recommends or judges employees, so fairness work focuses on consistency:
- Equivalent profiles get the same answer, whatever the name, language or gender-coded wording.
- Answer quality and routing rates are comparable across languages, locations and groups.
Test with paired prompts differing in one attribute and review results with HR.
Tone and empathy guidelines
Write a short tone guide with HR: warm and plain, no legal jargon unless quoting a clause, never cheerful about bad news, never minimising a concern, and offering a human when unsure. It goes into the system prompt, and graders score answers against it.
Cloud
Follow the company's existing estate: on AWS, ECS or EKS with RDS for PostgreSQL, Amazon Bedrock and VPC endpoints; on Azure, common where Entra ID and Teams are in use, AKS or Container Apps with Azure OpenAI and private endpoints. Keep HRIS connectivity private and provision with Terraform. The Microsoft Entra ID course covers the identity groundwork that HR role scoping depends on.
Observability
Trace every request through router decision, retrieval filters, policy version used, tool calls and confirmation outcome, using Langfuse or LangSmith with OpenTelemetry. Redact personal details at the span level. Useful dashboards: sensitive-routing rate and time to human pickup, "not found" by policy area (where policies are unclear), confirmation abandonment (confusing drafts) and thumbs-down rate by language.
Alert on HRIS tool failures and on a sudden drop in sensitive routing, which may mean the classifier has drifted.
Evaluation
Build the test set with HR before tuning anything. Each case has an employee profile, a question, a date context, and the expected behaviour (answer, route or act).
- Personalisation cases: the same question asked by different profiles, such as a Pune plant contractor and a Hyderabad permanent employee, with different correct answers.
- Effective-date cases: questions about dates before and after a policy change.
- Sensitive-routing cases: direct, indirect and mixed messages ("I need leave because my manager keeps humiliating me"), written by HR and Internal Committee members. The mixed ones matter most. The correct behaviour is to route, not to process the leave request first.
- Privacy cases: attempts to read another employee's data, and HR users querying outside their scope.
- Action and tone cases: edits, cancellations, insufficient balance; tone scored by native speakers.
Measure routing as a classifier with sensitive-case recall as the headline number, then faithfulness and correctness for answers and exact-match for actions. Calibrate any LLM judge against HR reviewers. Our LLM evaluation guide covers the method, and how to evaluate AI agents covers multi-step and tool-use scoring.
Deployment
GitHub Actions runs unit tests, the privacy suite and the full evaluation; any leak, or routing recall below the agreed threshold, fails the build. Prompts, tone guide, classifier thresholds and applicability tags are versioned configuration behind the same gate.
Roll out in stages (HR team, one location, all locations) with a flag to fall back to "policy answers only". Adoption depends on trust:
- Launch with a note from HR on what the agent does, what it never does, and how sensitive matters are handled.
- Put it in Teams or the HR portal, always with a "talk to a person" option.
- Review feedback weekly with policy owners and add every failure to the test set.
ROI
Agree the method with HR and finance, measure a baseline, then compare. All inputs below are hypothetical placeholders to show the arithmetic, not results.
| Input | Placeholder | Real source |
|---|---|---|
| Routine HR tickets per month (T) | e.g. 2,000 | Ticket system baseline |
| Share resolved by the agent (S) | e.g. 0.4 | Pilot measurement |
| HR handling minutes per ticket (H) | e.g. 10 | Time sample |
| Loaded HR cost per hour (C) | company figure | Finance |
| Annual run and support cost (R) | from billing | Cloud, model usage, maintenance |
Annual HR time value = T Γ S Γ H Γ 12 Γ· 60 Γ C. With the placeholders, T Γ S Γ H Γ 12 Γ· 60 gives 1,600 hours a year to multiply by C. Net value = time value β R β build cost amortised. Report faster pickup of sensitive cases separately, never as a cost saving.
Build it yourself: milestone plan
Use self-written policies and synthetic employees. Never use real HR data.
| Milestone | Deliverable |
|---|---|
| 1. Corpus and profiles | Leave and travel policies for three locations and two versions; a mock HRIS with a few dozen synthetic employees |
| 2. Test set | Personalisation, effective-date, sensitive, privacy and action cases |
| 3. Personalised RAG | Applicability and date filters in SQL; cited answers |
| 4. Router | Risk classifier with handoff flow; routing recall reported |
| 5. Actions | apply_leave and create_hr_ticket via an MCP server with confirmation |
| 6. Security | OIDC login, own-data tools, HR role scopes, privacy suite |
| 7. Ops | Tracing with redaction, CI eval gate, Terraform deploy to AWS or Azure |
| 8. Value | ROI method one-pager, tone guide, demo of a routed sensitive case |
In your README, show the routing evaluation as prominently as answer quality: interviewers remember a candidate who can explain why the agent declined to answer.
Frequently asked questions
What is an HR AI agent?
An HR AI agent is an assistant that answers employee policy questions from approved HR documents, answers personal questions from the employee's own HRIS data, and completes simple requests such as applying for leave after the employee confirms. It routes sensitive matters to human HR staff.
How is an HR chatbot with RAG different from a general knowledge assistant?
HR answers depend on who is asking and when. An HR chatbot with RAG filters policies by the employee's location, grade and employment type from the HRIS, and by the policy version in force on the relevant date, before it retrieves anything. A general knowledge assistant usually filters only by document permissions.
Should an HR AI agent handle harassment or grievance complaints?
No. By design it should recognise grievances, harassment, health and performance concerns, acknowledge them, explain the confidential human channel, and hand them to the HR business partner, the Internal Committee process or the employee assistance programme, without giving advice or a policy summary.
How do you stop employees seeing each other's HR data?
Personal tools take the employee identity only from the authenticated SSO session and never accept an employee ID from the model. HR staff get scoped access through directory roles. A cross-employee privacy test suite runs in CI and fails the build on any leak.
How do you evaluate an HR AI agent?
Use a test set built with HR in which each case has an employee profile, a question, a date context and the expected behaviour. Measure sensitive-routing recall as the headline metric, then personalised answer correctness, faithfulness, action accuracy, privacy and tone across languages.
Can AI for HR introduce bias?
It can, so keep the agent out of decisions about individuals and test for consistency. Paired prompts that differ in one attribute, such as name or language, should get equivalent answers. Routing and answer quality should be compared across groups and languages and reviewed with HR.
Want to build systems like this with a trainer reviewing your design decisions? Cloudsoft FDE PRO is a 12-week program with five enterprise projects and the GlobalBank capstone, taught in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 for a free demo.



