New batches starting this week Β· Limited seats

Project Walkthrough: Building an HR AI Agent for Employee Questions and Requests

A project walkthrough for an HR AI agent at an illustrative multi-location Indian company, covering personalised policy answers, effective dates, confirmed leave requests, routing sensitive topics to humans, privacy, fairness, evaluation and adoption.

HR AI agent flow: employee question, policy by location and grade, own data only, simple requests, sensitive topics routed to a human
Last updated Β· 15 min read Β· 3,299 words

This project walks through how a Forward Deployed Engineer would build an HR AI agent, from the business problem to an ROI method. An HR AI agent is ready for production only when it answers policy questions for the employee's own location, grade and employment type, shows only the employee's own data, asks for confirmation before it takes any action, and passes grievances, harassment, health and performance concerns to a human every time. The scenario is illustrative. The build plan at the end turns it into a portfolio project you can explain in an interview.

The generic retrieval mechanics (chunking, hybrid search, reranking, citation checks) are already covered in our enterprise RAG knowledge assistant project. This article covers what is different about HR: personalisation, effective dates, sensitive routing, privacy and tone.

Business problem

Illustrative scenario. Consider a company with several thousand employees across Indian locations: a head office in Bengaluru, a GCC delivery centre in Hyderabad, sales offices in Mumbai and Delhi NCR, and a plant in Pune. The workforce includes permanent staff at several grades, fixed-term contract employees and interns. An HR shared-services team runs a helpdesk through email, an intranet portal and a ticket queue.

The helpdesk's problems are familiar:

  • Most tickets are repeat questions: "How many casual leaves do I get?", "When is the payslip released?"
  • Answers depend on who is asking. The leave policy has a different holiday list at each location. Travel limits depend on grade. Contract staff have different benefits.
  • Policies change yearly, and employees quote last year's PDF.
  • Serious concerns land in the same queue as routine questions.

HR wants faster answers, fewer repeat tickets and a safe channel for serious matters. That shift from AI demo to enterprise outcome shapes every decision below.

Requirements

Discovery covers HR operations, HR business partners, payroll, the Internal Committee that handles POSH complaints, legal, information security and employee representatives.

Functional

  • Policy questions: answer from approved policies only, personalised to the employee's location, grade and employment type, and citing the policy, clause and effective date.
  • Personal questions: answer leave balance, payslip release date and reporting-manager details from the employee's own HRIS record only.
  • Simple requests: apply for leave and raise an HR ticket. Each is an action that the employee confirms before anything is written.
  • Sensitive topics: grievances, harassment, health, mental-health concerns, performance disputes and disciplinary matters are never answered by the AI. They go to a named human channel.
  • Languages: employees can ask in their preferred Indian language or English.

Non-functional

  • An employee sees only their own data. HR staff see more according to their role and scope.
  • Employee data stays in the company's cloud account and approved region, and is not used to train vendor models.
  • Every answer and action is auditable: question, policy version used, tool called.

Success metrics

MetricHow it is measuredOwner
Personalised answer correctnessTest set of questions asked as different employee profilesEngineering + HR policy owners
Sensitive-routing recallSensitive test cases sent to a human; target agreed with HR and legalHR + Engineering
Data leakageMust be zero on the cross-employee access suiteSecurity
Action accuracyLeave requests and tickets created exactly as confirmedEngineering
Ticket deflection and adoptionRepeat tickets before vs after; weekly active employeesHR operations

Architecture

The design has a router in front, three paths behind it, and identity, audit and evaluation running across all of them.

Employee (Teams / portal) --SSO--> API
                |
       load profile from HRIS
     (location, grade, type, role)
                |
          Intent + risk router
     /          |            \
 sensitive   policy Q     personal / action
    |           |              |
 human       RAG over      HRIS tools
 handoff     policies      (own data only)
 (HRBP/IC/   filtered by       |
  EAP)       profile+date   confirm step
    |           |              |
    +------> answer / ticket <-+
                |
      audit log + traces (Langfuse)

Key decisions:

  • The router runs before retrieval, so a harassment disclosure never reaches the policy-answering prompt.
  • The profile comes from the HRIS, not the conversation. Typing "I am a band 7 in Mumbai" changes nothing.
  • Policies and personal data stay apart. Policy text sits in PostgreSQL with pgvector; personal data is never embedded, only fetched live through scoped tools.

Data

Policy corpus

Leave, travel, benefits, insurance, code of conduct and holiday lists, as PDFs and intranet pages. Chunk by clause, keep each clause's applicability statement ("Applies to: permanent employees, all locations except plant") in the chunk, and keep grade-wise limit tables whole.

Applicability and versioning metadata

FieldUsed for
policy_id, clause_path, titleCitations
applies_location, applies_grade, applies_emp_typePersonalised filtering (list or "all")
effective_from, effective_to, version, statusPicking the version in force on a given date
supersedesExplaining "this changed from last year"
ownerRouting unanswered questions

Effective dates are where HR assistants most often go wrong. A question about "my leave next month" must use the policy in force on that date, possibly a published version not yet in force. Filter on the relevant date, not today by default, and tag changed clauses so the agent can say "This changed from 1 April; the earlier rule was…".

Applicability tags are rarely written cleanly in the source. Extract them with an LLM-assisted pass, have the policy owner review every tag, and report contradictions (two holiday lists for one location) back to HR rather than fixing them in code.

Employee data (HRIS)

Location, grade, employment type, manager, leave balances and payroll calendar come from the HRIS API, and only what one answer needs enters the prompt.

LLM

Choose the model by testing it on HR questions. Accept it only if it meets these criteria:

  • It stays strictly within the retrieved clauses and refuses when they are silent, because a guessed leave rule creates an HR dispute.
  • It writes well in the Indian languages your employees use, judged by native speakers.
  • It follows tone instructions consistently.
  • It is available in the approved region on the company's platform: Amazon Bedrock, Azure OpenAI or Gemini on Google Cloud.

Use a small, fast model for intent and risk classification and a more capable one for answers. The classifier is safety-critical and gets its own evaluation set.

RAG

The retrieval pipeline matches the knowledge assistant walkthrough, with two HR-specific additions:

  1. Profile filtering inside the query. Each retrieval filters on applies_location, applies_grade and applies_emp_type matching the employee's HRIS profile, or "all". It also filters on the effective window covering the date in question. The filter runs in the SQL before ranking, so the model never sees the travel limits for another grade.
  2. Applicability in the answer. The prompt asks the model to say which rule applies and why: "As a permanent employee at the Hyderabad office, you get…". This makes personalisation visible and lets the employee spot a wrong HRIS record.

For HR staff answering on behalf of others, the profile filter becomes an explicit parameter, limited to their scope and logged.

Multilingual questions

Employees may write in Hindi, Telugu, Tamil, Kannada, Marathi or a mix with English. If policies exist only in English, translate the query for retrieval, answer in the employee's language, keep the cited clause title in English and state that the English policy is authoritative. Add languages only after native-speaker reviewers approve quality.

Agent

The agent is a small, explicit state machine, which fits well in LangGraph. The important design choice is what the agent will not do.

message -> risk check --sensitive--> handoff
              |
           routine
              v
      intent: policy | personal | action
                                  |
                       draft request (dates,
                       type, balance check)
                                  v
                       "Apply 2 days CL,
                        12-13 Nov? Confirm"
                                  |
                           user confirms
                                  v
                     HRIS / ticket tool call
                                  v
                      reference shown to user

Sensitive routing by design

The risk check looks for grievances, harassment, discrimination, health and mental-health concerns, performance or disciplinary disputes, and signs of distress. When it triggers, the agent gives no policy summary, advice or probing questions. It acknowledges the message, explains who will handle it and how confidentially, and offers the right channel: the HR business partner, the Internal Committee process, or the employee assistance programme. With consent it creates a restricted case; urgent safety signals show emergency contacts at once.

The router is tuned to over-route: a routine question sent to a human costs minutes, while a harassment disclosure answered with a policy summary destroys trust. When unsure, route.

Actions with confirmation

The model drafts leave applications and tickets; deterministic code executes them only after the employee confirms a draft showing type, dates and balance before and after. The HRIS approval workflow stays in place: the agent submits requests, never approves them.

Tools

ToolScopeGuardrails
search_policiesProfile and date filtersFilters taken from the HRIS profile, never from model output
get_my_leave_balanceCaller onlyEmployee ID from the session token
get_payroll_calendarCaller's pay groupDates only; no salary figures shown in chat
apply_leaveCaller onlyConfirmation required; idempotency key; balance re-checked
create_hr_ticketCaller onlyConfirmation required; category from a fixed list
create_sensitive_caseRestricted queueMinimal content; visible only to the assigned HRBP or the Internal Committee

No tool accepts an employee_id from the model; identity comes from the session, closing off "show me my colleague's balance" attacks.

MCP/API

The HRIS and the HR ticketing system (ServiceNow HR Service Delivery, Jira Service Management or an in-house tool) are exposed as MCP servers. MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data. Each server exposes only the narrow tools above, uses a least-privilege service account (or on-behalf-of access where the HRIS supports it) and records the end user's identity for audit. See what MCP is for the protocol, and the ServiceNow AI agent project for a fuller MCP integration walkthrough.

Cloudsoft's AI Forward Deployed Engineer course does not include an HR agent project. Its Enterprise Knowledge Assistant and ServiceNow AI Agent via MCP projects teach the two building blocks this design combines: permission-aware RAG and confirmed tool actions through MCP.

Security

  • Own data only. Employees sign in with Microsoft Entra ID; every personal tool resolves the employee from the token.
  • Role-based HR access. HR roles (business partner, payroll, operations) come from directory groups, scoped to a business unit or location, with queries logged for review.
  • Minimal logging. Redact personal details from traces. Sensitive-case conversations keep only what the case needs, under restricted access and a retention period agreed with legal, in line with India's data protection law.
  • Prompt injection. Treat policy text and pasted content as data. Confirmation steps and identity-bound tools limit what an injected instruction can do.

The wider threat model is covered in AI security for enterprises. Approval, accountability and policy ownership are covered in enterprise AI governance.

Bias and fairness

The agent never ranks, recommends or judges employees, so fairness work focuses on consistency:

  • Equivalent profiles get the same answer, whatever the name, language or gender-coded wording.
  • Answer quality and routing rates are comparable across languages, locations and groups.

Test with paired prompts differing in one attribute and review results with HR.

Tone and empathy guidelines

Write a short tone guide with HR: warm and plain, no legal jargon unless quoting a clause, never cheerful about bad news, never minimising a concern, and offering a human when unsure. It goes into the system prompt, and graders score answers against it.

Cloud

Follow the company's existing estate: on AWS, ECS or EKS with RDS for PostgreSQL, Amazon Bedrock and VPC endpoints; on Azure, common where Entra ID and Teams are in use, AKS or Container Apps with Azure OpenAI and private endpoints. Keep HRIS connectivity private and provision with Terraform. The Microsoft Entra ID course covers the identity groundwork that HR role scoping depends on.

Observability

Trace every request through router decision, retrieval filters, policy version used, tool calls and confirmation outcome, using Langfuse or LangSmith with OpenTelemetry. Redact personal details at the span level. Useful dashboards: sensitive-routing rate and time to human pickup, "not found" by policy area (where policies are unclear), confirmation abandonment (confusing drafts) and thumbs-down rate by language.

Alert on HRIS tool failures and on a sudden drop in sensitive routing, which may mean the classifier has drifted.

Evaluation

Build the test set with HR before tuning anything. Each case has an employee profile, a question, a date context, and the expected behaviour (answer, route or act).

  • Personalisation cases: the same question asked by different profiles, such as a Pune plant contractor and a Hyderabad permanent employee, with different correct answers.
  • Effective-date cases: questions about dates before and after a policy change.
  • Sensitive-routing cases: direct, indirect and mixed messages ("I need leave because my manager keeps humiliating me"), written by HR and Internal Committee members. The mixed ones matter most. The correct behaviour is to route, not to process the leave request first.
  • Privacy cases: attempts to read another employee's data, and HR users querying outside their scope.
  • Action and tone cases: edits, cancellations, insufficient balance; tone scored by native speakers.

Measure routing as a classifier with sensitive-case recall as the headline number, then faithfulness and correctness for answers and exact-match for actions. Calibrate any LLM judge against HR reviewers. Our LLM evaluation guide covers the method, and how to evaluate AI agents covers multi-step and tool-use scoring.

Deployment

GitHub Actions runs unit tests, the privacy suite and the full evaluation; any leak, or routing recall below the agreed threshold, fails the build. Prompts, tone guide, classifier thresholds and applicability tags are versioned configuration behind the same gate.

Roll out in stages (HR team, one location, all locations) with a flag to fall back to "policy answers only". Adoption depends on trust:

  • Launch with a note from HR on what the agent does, what it never does, and how sensitive matters are handled.
  • Put it in Teams or the HR portal, always with a "talk to a person" option.
  • Review feedback weekly with policy owners and add every failure to the test set.

ROI

Agree the method with HR and finance, measure a baseline, then compare. All inputs below are hypothetical placeholders to show the arithmetic, not results.

InputPlaceholderReal source
Routine HR tickets per month (T)e.g. 2,000Ticket system baseline
Share resolved by the agent (S)e.g. 0.4Pilot measurement
HR handling minutes per ticket (H)e.g. 10Time sample
Loaded HR cost per hour (C)company figureFinance
Annual run and support cost (R)from billingCloud, model usage, maintenance

Annual HR time value = T Γ— S Γ— H Γ— 12 Γ· 60 Γ— C. With the placeholders, T Γ— S Γ— H Γ— 12 Γ· 60 gives 1,600 hours a year to multiply by C. Net value = time value βˆ’ R βˆ’ build cost amortised. Report faster pickup of sensitive cases separately, never as a cost saving.

Build it yourself: milestone plan

Use self-written policies and synthetic employees. Never use real HR data.

MilestoneDeliverable
1. Corpus and profilesLeave and travel policies for three locations and two versions; a mock HRIS with a few dozen synthetic employees
2. Test setPersonalisation, effective-date, sensitive, privacy and action cases
3. Personalised RAGApplicability and date filters in SQL; cited answers
4. RouterRisk classifier with handoff flow; routing recall reported
5. Actionsapply_leave and create_hr_ticket via an MCP server with confirmation
6. SecurityOIDC login, own-data tools, HR role scopes, privacy suite
7. OpsTracing with redaction, CI eval gate, Terraform deploy to AWS or Azure
8. ValueROI method one-pager, tone guide, demo of a routed sensitive case

In your README, show the routing evaluation as prominently as answer quality: interviewers remember a candidate who can explain why the agent declined to answer.

Frequently asked questions

What is an HR AI agent?

An HR AI agent is an assistant that answers employee policy questions from approved HR documents, answers personal questions from the employee's own HRIS data, and completes simple requests such as applying for leave after the employee confirms. It routes sensitive matters to human HR staff.

How is an HR chatbot with RAG different from a general knowledge assistant?

HR answers depend on who is asking and when. An HR chatbot with RAG filters policies by the employee's location, grade and employment type from the HRIS, and by the policy version in force on the relevant date, before it retrieves anything. A general knowledge assistant usually filters only by document permissions.

Should an HR AI agent handle harassment or grievance complaints?

No. By design it should recognise grievances, harassment, health and performance concerns, acknowledge them, explain the confidential human channel, and hand them to the HR business partner, the Internal Committee process or the employee assistance programme, without giving advice or a policy summary.

How do you stop employees seeing each other's HR data?

Personal tools take the employee identity only from the authenticated SSO session and never accept an employee ID from the model. HR staff get scoped access through directory roles. A cross-employee privacy test suite runs in CI and fails the build on any leak.

How do you evaluate an HR AI agent?

Use a test set built with HR in which each case has an employee profile, a question, a date context and the expected behaviour. Measure sensitive-routing recall as the headline metric, then personalised answer correctness, faithfulness, action accuracy, privacy and tone across languages.

Can AI for HR introduce bias?

It can, so keep the agent out of decisions about individuals and test for consistency. Paired prompts that differ in one attribute, such as name or language, should get equivalent answers. Routing and answer quality should be compared across groups and languages and reviewed with HR.

Want to build systems like this with a trainer reviewing your design decisions? Cloudsoft FDE PRO is a 12-week program with five enterprise projects and the GlobalBank capstone, taught in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us