This walkthrough builds a WhatsApp AI chatbot for an Indian business from business problem to measured return, the way a Forward Deployed Engineer would deliver it. A production WhatsApp AI chatbot is mostly channel and trust engineering: it must respect the 24-hour customer service window and template rules, cope with Telugu, Hindi, English and code-mixed text and voice notes, verify who is behind a phone number before sharing personal data, and hand over to staff without making the customer repeat themselves. The scenario is illustrative; the build plan at the end turns it into a portfolio project.
Generic customer-facing patterns (policy engine, honest deflection metrics) are in the enterprise AI customer support agent project; this one focuses on the WhatsApp Business Platform and Indian customers.
Business problem
Illustrative scenario. Consider a two-wheeler service chain with workshops across Hyderabad and nearby towns. Most customers already message the nearest workshop's phone number on WhatsApp: "bike service ki slot undha Saturday?", "meri scooty ready hai kya?", or a Telugu voice note describing an engine noise. Advisors answer between jobs from personal phones, with no record in the dealer management system (DMS).
The symptoms:
- Saturday and pre-festival mornings bring a burst of booking messages that nobody answers for hours.
- "Is my vehicle ready?" calls interrupt advisors all afternoon, though the status sits in the DMS.
- Additional-work approvals ("brake pads are worn, shall we replace?") stall because the customer missed a call.
The owner wants one official WhatsApp number for the chain that books slots, answers FAQs, gives service status, collects additional-work approvals and routes anything else to the right workshop's staff.
Requirements
Discovery covers the owner, workshop managers, advisors and the DMS vendor. Tag a sample of exported chats (with permission, redacted) by intent and language before writing code.
Functional
- Book, reschedule and cancel service slots by workshop, vehicle model and service type.
- Answer FAQs (service intervals, pick-up areas, timings) from approved content.
- Give job status, estimates and invoices only to a verified owner.
- Present additional-work estimates and record the customer's approval or refusal.
- Understand text and voice notes in Telugu, Hindi and English, in native script, Roman script and code-mixed form.
- Hand off to the workshop's advisor with context.
Platform constraints from the WhatsApp Business Platform
From Meta's developer documentation and policies at the time of writing; Meta revises them, so recheck before you build.
- Access. Use the Cloud API, which Meta hosts. You can integrate directly as a business building for its own use, or go through a Meta partner (historically called a Business Solution Provider, or BSP; Meta now uses terms such as Solution Partner and Tech Provider) who handles onboarding, billing and often an agent inbox.
- Customer service window. When a customer messages the business, a 24-hour window opens, refreshed by each new customer message. Inside it the bot can send free-form messages; outside it, only pre-approved template messages.
- Templates. Each template is submitted with a category (marketing, utility or authentication) and reviewed by Meta before use; templates can later be paused or disabled based on quality. "Your vehicle is ready for pick-up" is a utility template; a festive offer is marketing.
- Pricing. Pricing is per delivered template message and depends on category and the recipient's country; service replies inside the window are not charged, and utility templates delivered inside an open window are also free. Take prices from Meta's current rate card, never hard-code them.
- Opt-in. The business may message a person only after opt-in that names the business and says what they will receive, and it must honour opt-outs.
- Human escalation. Meta's Business Messaging Policy requires automated experiences to offer prompt, clear escalation paths to a human, and forbids asking people to share full payment card numbers, financial account numbers or personal ID numbers in chat.
- AI-provider restriction. Since January 2026 the WhatsApp Business Solution Terms bar AI providers from using the platform to distribute general-purpose assistants as the primary service. A service chain's own booking and support bot, where AI is incidental, is a different case; keep it scoped to the business and have legal read the current terms.
Success metrics
| Metric | Definition | Owner |
|---|---|---|
| Self-served bookings | Slots booked and actually attended without staff involvement | Operations |
| Status calls avoided | Inbound "is it ready" calls per job, before and after | Workshop managers |
| Approval turnaround | Time from estimate sent to customer decision | Service advisors |
| Data exposure incidents | Must be zero: status or invoice sent to the wrong person | Data owner |
Architecture
Meta delivers every incoming message to your webhook, which only verifies, stores and queues; workers do the slow work.
WhatsApp user
|
Meta Cloud API --webhook--> Ingest API (FastAPI)
verify signature,
dedupe, 200 OK
|
queue
|
worker: media fetch -> speech-to-text
|
orchestrator (LangGraph state machine)
| | |
FAQ RAG identity + tools via MCP
consent (DMS, slots,
estimates)
|
outbound sender: window check,
template vs free-form, rate limiter
|
handoff -> advisor inbox
Key decisions:
- Webhook and processing are decoupled. Meta retries undelivered webhooks for up to several days, and retries can produce duplicates, so the ingest API deduplicates on the WhatsApp message ID and returns 200 before any LLM call.
- One outbound sender owns channel rules. Nothing else calls the send API. It checks whether the 24-hour window is open, picks a free-form message or an approved template, applies rate limits and records delivery status webhooks.
- Per-customer ordering. Customers send three short messages in a row: process each sender in order and debounce briefly so the bot answers the combined thought.
Data
| Data | Source | Handling |
|---|---|---|
| FAQs, service menu, workshop timings, pick-up areas | Owner-approved content, in English with Telugu and Hindi versions | Indexed for RAG; versioned; owner signs off |
| Customers, vehicles, job cards, estimates, invoices | DMS | Never indexed; fetched live via tools after verification |
| Slots and workshop capacity | Booking system or DMS calendar | Live tool calls only |
| Chat history, voice notes, transcripts | The bot itself | Personal data: retention limits, redaction in logs |
| Consent and opt-in records | The bot and DMS | Timestamped, purpose-tagged, auditable |
The usual surprise: the DMS "registered mobile" is often a family member's number or blank, which caps automatic verification. Measure it in week one.
LLM
Select on your own redacted WhatsApp messages, scored by native Telugu and Hindi speakers, for:
- Understanding Roman-script Telugu and Hindi with English words mixed in ("service ayyaka call cheyyandi", "kal subah drop kar sakte ho?").
- Reliable tool calling and extraction of dates and vehicle numbers.
- Short replies: two or three lines plus buttons.
- Availability on the client's cloud (Amazon Bedrock, Azure OpenAI or Gemini).
Use a small model for language, intent and guardrail checks and a stronger one for tool planning. Non-Latin scripts often use more tokens per word, so send a rolling summary plus recent turns, not the whole thread; tokens and context windows explains the budget.
RAG
RAG answers only general questions over a small corpus: service menu, intervals, warranty basics, timings. Hybrid search over pgvector and PostgreSQL full-text with a relevance threshold is enough. Two WhatsApp-specific points:
- Store approved answers in all three languages so warranty terms never drift in translation. Retrieve cross-lingually, reply in the customer's language.
- Below the threshold, the bot does not guess; it offers buttons: "Talk to the workshop" or "Book a slot".
Agent
A LangGraph state machine: the model chooses within a state; code decides which tools each state exposes.
message (text / voice / image)
|
voice? -> speech-to-text -> transcript
|
detect language + intent
|
+-- FAQ ---------------> RAG answer + buttons
+-- booking -----------> slot tools -> confirm
+-- status / invoice
| |
| verified? --no--> verify owner
| | yes
| v
| job status / invoice / estimate
+-- approve extra work -> show estimate
| -> explicit button
+-- complaint / "human" -> handoff
Voice notes
The webhook carries a media ID, not the audio; the worker exchanges it for a download URL that expires within minutes. Voice notes arrive as Opus audio in OGG. Transcribe with a speech-to-text model tested on Telugu, Hindi and code-mixed speech, keep the confidence score, and treat the transcript as untrusted input. Low confidence, or any action from voice, gets a text read-back: "You want a general service at Kukatpally on Saturday at 10 am. Confirm?" with buttons. More on speech in voice AI agents.
Identity verification
A message from the DMS-registered mobile is a useful signal, not proof: Indian families share phones and bikes.
- Sender matches the registered mobile: allow low-risk status ("Your job is in progress, expected by 5 pm").
- Invoices, estimates with amounts, address or other personal details: step up with a one-time code sent through an approved authentication template, or SMS, to the registered number; verification expires.
- Sender does not match: share nothing personal; offer to message the registered number, or hand off.
- Vehicle and job card numbers are never proof; they are printed on the vehicle and paper.
Consent under DPDP
Meta's opt-in and India's Digital Personal Data Protection Act consent are different things, and you need both. On first contact, send a short notice in the customer's language: what data the bot processes (messages, voice notes, service records), why, for how long, and how to withdraw. Record consent with purpose tags: service communication is one purpose, marketing templates another, and declining marketing must not block booking. Withdrawal ("STOP", "aapandi", "band karo") must work as easily as consent. Most DPDP Rules obligations apply from May 2027; build them now. Mapping consent, retention and processors (Meta, partner, LLM and speech providers) is covered in DPDP Act for AI applications. Not legal advice; involve counsel.
Staff handoff
Hand off when the customer asks for a person in any language, on complaints, warranty or insurance disputes, repeated low confidence or payment problems. The advisor gets a structured summary (verified status, vehicle, intent, language, voice transcripts, what was tried) in the agent inbox, and the bot goes silent on that thread until the advisor closes it. Staff replies after the 24-hour window must use a template. More in human-in-the-loop AI.
Tools
| Tool | Inputs | Guardrails |
|---|---|---|
| find_slots | workshop, service_type, date range | Read-only; live capacity |
| book_slot | slot_id, vehicle, pickup flag | Explicit button confirmation; idempotent; per-number booking limit |
| get_job_status | none (vehicle from session) | Registered-mobile match; returns stage and ETA only |
| get_invoice | job_card_id | Step-up verified; sends PDF as WhatsApp document |
| record_estimate_decision | estimate_id, decision | Step-up verified; only from an interactive button reply, never from free text; logged with message ID |
| handoff_to_staff | reason, summary, workshop | Always available; never rate-limited |
Arguments are schema-validated, and "next Saturday" or "repu sayantram" is resolved by code to an absolute IST timestamp. See function calling and structured outputs. Deliberately missing: discounts, refunds and payments (a gateway payment link instead).
MCP/API
The DMS, booking calendar and estimates are exposed through MCP servers with narrow, typed operations. MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data. The orchestrator injects the verified customer and vehicle IDs; the model cannot set them.
The WhatsApp side is plain HTTPS: the Cloud API's send endpoint, the media endpoints, template management, and webhooks for messages and delivery statuses. Use reply buttons and list messages for choices; a tap is unambiguous where "ok" is not. Wrapping dated DMS APIs safely is integration work of the kind Cloudsoft's AI Forward Deployed Engineer course practises in its Customer Integration Service and ServiceNow AI Agent via MCP projects.
Security
- Webhook authenticity: verify the
X-Hub-Signature-256HMAC with the app secret on every request, and answer the subscription challenge only with your verify token. Reject anything else. - Wrong-person disclosure: the top risk. Enforce verification in code, and test shared-phone and mismatched-number cases on every build.
- Prompt injection via text, transcripts or captions: tool permissions live in code, so a jailbreak cannot approve work or reveal an invoice.
- Sensitive identifiers: never ask for Aadhaar, card or bank numbers; detect and mask them if customers send them anyway, including in images.
- Secrets: the Cloud API access token lives in a secrets manager, rotated.
Cloud
On AWS: the ingest API behind a load balancer, SQS FIFO grouped by sender for ordering, ECS or EKS workers, RDS for PostgreSQL with pgvector, Amazon Bedrock and Secrets Manager, in the Mumbai or Hyderabad region where residency matters. On Azure: Container Apps, Service Bus sessions, Azure OpenAI and Key Vault. Terraform provisions both.
Rate limits, in four layers
- Meta throughput: each business number has a messages-per-second ceiling; pace festive reminder campaigns.
- Per-recipient pacing: Meta also limits how quickly you can message one user. Combine replies into one message with buttons instead of firing several.
- Messaging limits: business-initiated templates are capped by unique recipients in a rolling 24 hours, in tiers that rise with business verification and quality.
- Your own limits: LLM and speech provider quotas, plus per-sender caps against abuse. Route model calls through an LLM gateway for quotas, retries and fallback.
Observability
Trace each message with OpenTelemetry and Langfuse: dedupe result, transcription confidence, language, intent, tool calls, outbound type, delivery status and handoff reason. Owner dashboards show bookings by workshop and language, handoff reasons, verification failures, template delivery, quality ratings and cost per conversation. Alert on webhook failures, queue lag and template pauses.
Evaluation
Grade whole conversations from redacted chats and scripted personas (method in AI agent evaluation):
- Bookings in Telugu script, Roman Telugu, Hinglish and English, with relative dates.
- Voice notes from varied speakers with traffic noise.
- Identity cases: family member's phone, unregistered number, job card number offered as proof.
- Window cases: follow-up due 30 hours later must use a template; opt-out respected.
- Duplicate webhooks, DMS timeouts, injection via captions and voice.
Native speakers review language results each release. Wrong-person disclosure and approvals from free text are pass/fail: one failure blocks release.
Deployment
GitHub Actions runs unit tests, the safety suite and the conversation evals on every change; prompts, templates (as code), model IDs and FAQ content are versioned. Use a separate staging number, never production. Roll out in stages: one workshop, FAQs and bookings first; then status; then estimate approvals once verification holds. Keep a kill switch that routes the number to staff instantly.
If you want to build this kind of system under review rather than alone, FDE PRO runs 12 weeks with five enterprise projects, the GlobalBank capstone and a weekly Customer Engagement Lab, in Ameerpet or live online.
ROI
Agree the method with the owner and measure against a baseline. Every value below is a hypothetical placeholder, not a result.
| Input | Placeholder | Source of the real value |
|---|---|---|
| Bookings self-served per month (B) | e.g. 1,000 | Bot logs, attended only |
| Staff minutes saved per booking (m1) | e.g. 4 | Time study at a pilot workshop |
| Status calls avoided per month (C) | e.g. 2,000 | Call logs before vs after |
| Minutes per status call (m2) | e.g. 2 | Time study |
| Cost per staff minute (w) | client's figure | Payroll |
| Extra approved work from faster approvals (E) | measured, not assumed | DMS, against a holdout workshop |
| Running cost (R) | from billing | Meta template charges, partner fees, model, speech, cloud |
Monthly value = ((B Γ m1) + (C Γ m2)) Γ w + margin on E β R. With the placeholders, staff time freed is 8,000 minutes a month before costs. Credit E only with holdout evidence; report after-hours bookings separately.
Build it yourself
Use a fictional service chain, a Meta test number and seeded data.
| Milestone | Deliverable |
|---|---|
| 1. Channel | Webhook with signature check, dedupe, queue, outbound sender with window logic |
| 2. Mocks | Fake DMS, slots and estimates APIs behind MCP servers |
| 3. FAQ RAG | Trilingual approved answers, threshold, buttons |
| 4. Agent | LangGraph flow with verification, consent and handoff states |
| 5. Voice and evals | Speech-to-text with read-back; multilingual and safety suites in CI |
| 6. Ops and value | Tracing, dashboard, Terraform, ROI one-pager |
In the README, show a refused wrong-person invoice request, a Telugu voice-note booking with read-back, and a template follow-up outside the window.
Frequently asked questions
Do I need a BSP to build a WhatsApp AI chatbot?
No. A business can integrate directly with Meta's hosted Cloud API for its own use. A Meta partner (formerly called a Business Solution Provider) can simplify onboarding, billing and a shared agent inbox for staff handoff.
What is the WhatsApp 24-hour customer service window?
When a customer messages the business, a 24-hour window opens, refreshed by each new customer message. Inside it the bot can send free-form replies; outside it, only pre-approved template messages can be sent.
How should a WhatsApp chatbot verify identity?
Treat the sender's number matching the registered mobile as a signal for low-risk updates only. For invoices, amounts and personal details, step up with a one-time code to the registered number, and never accept vehicle or job card numbers as proof.
How do you handle voice notes in Telugu or Hindi?
Download the audio using the media ID from the webhook, transcribe it with a speech-to-text model tested on Indian languages and code-mixed speech, and confirm any action with a text read-back and buttons before acting.
Is WhatsApp opt-in the same as DPDP consent?
No. Meta's opt-in permits the business to message the person; DPDP consent covers processing their personal data for specific purposes. Collect both, record purposes separately, and make withdrawal as easy as giving consent.
Can a business use an LLM on WhatsApp under Meta's rules?
Meta's terms bar AI providers from distributing general-purpose assistants as their main service on the platform; a business using AI for its own service and bookings is a different case. Keep the bot scoped and have legal review the current terms.
Ready to take channel-integrated AI from demo to a system a business can rely on? Explore the Cloudsoft FDE PRO program, in our Ameerpet classroom beside Ameerpet Metro or live online, with placement support until you're placed. Call +91 96660 19191 for a free demo.



