New batches starting this week Β· Limited seats

Voice AI Agents Explained: Building Real-Time Conversational AI

An engineering guide to voice AI agents: the cascade and speech-to-speech architectures, turn-taking and latency, telephony integration, tool use with confirmation, human handoff, compliance, evaluation and cost per minute.

Voice AI agent pipeline: speech in, speech-to-text, LLM with tools, text-to-speech, natural reply
Last updated Β· 15 min read Β· 3,215 words

A voice AI agent holds a spoken conversation in real time: it listens, works out what the caller wants, takes actions such as looking up a record or booking a slot, and answers in a natural voice. Voice AI agents are built either as a cascade (speech-to-text, then an LLM, then text-to-speech) or on a speech-to-speech realtime model, and in both cases the hard engineering is not the model but latency, turn-taking, telephony, entity accuracy and safe handoff to humans. This guide explains both architectures and what it takes to run an AI phone agent that callers will tolerate.

If you have built a chat agent, most of the agent logic carries over; the enterprise AI customer support agent project covers identity verification, policy limits and resolution metrics for chat. Voice adds its own problems: no screen, no scrollback, no time to think, and every mistake is heard immediately.

What a voice AI agent is (and is not)

A voice AI agent is not an IVR menu with better speech recognition. An IVR says "press 1 for billing" and follows a fixed tree. Conversational voice AI lets callers speak in their own words, asks clarifying questions, keeps context across turns, and calls back-end tools mid-conversation. It is an LLM agent whose input and output are audio, running under a strict real-time constraint.

Typical uses are inbound service lines, outbound reminders, IT helpdesks and front-door triage. For audio alongside images and documents, see multimodal AI for enterprise.

Voice AI agent architecture: cascade vs speech-to-speech

There are two mainstream ways to build the conversational core. Both sit inside the same outer system of telephony, tools, logging and handoff.

The cascade: STT, LLM, TTS

Streaming speech-to-text (STT, or ASR) turns the caller's audio into text as they speak. An LLM reads the transcript and history, decides what to say or which tool to call, and streams a text reply. Text-to-speech (TTS) converts the reply to audio, ideally starting before the full sentence is generated.

  • Control and inspectability. Every turn has text you can log, filter and evaluate.
  • Component choice. Pick STT that handles Indian accents well, an LLM strong at tool calling, and a TTS voice that suits the brand, and swap each independently.
  • Guardrails slot in naturally. Checks for PII, policy and prohibited commitments run between the LLM and TTS, as described in AI guardrails.
  • Mature tool use. See function calling and structured outputs.

The weaknesses: each hop adds delay, and tone, hesitation and emphasis are mostly lost when audio becomes text.

Speech to speech AI (realtime models)

A speech-to-speech model takes audio in and produces audio out within one model, usually over a persistent streaming connection. Several model providers now offer realtime APIs of this kind, often with tool calling and a parallel text transcript.

Its strengths are more natural responsiveness with fewer hand-offs, awareness of how something was said, and a simpler pipeline for prototypes. Its weaknesses are the mirror image of the cascade's: inserting a text check between "thinking" and "speaking" gives back some of the latency advantage; the audio the caller heard is the ground truth and may not match the transcript exactly; you are coupled to one provider's voices, languages and pricing; and quality on Indian languages and code-mixed speech must be tested, not assumed.

Choosing between them

FactorCascade (STT β†’ LLM β†’ TTS)Speech-to-speech realtime
ResponsivenessGood with streaming; more hops to tuneTypically more natural, fewer hops
Control and guardrailsText checks between stagesHarder to intercept before audio plays
ObservabilityText at every stepAudio-first; transcript secondary
FlexibilitySwap STT, LLM, TTS independentlyMostly tied to one provider
Indian languagesSpecialised STT/TTS per languageDepends on model coverage
Good fitRegulated, tool-heavy, multilingual flowsConversational feel matters most, simpler tools

Hybrids are common, for example a realtime model whose tool calls are validated in a text pipeline.

Latency budget, turn-taking and barge-in

In chat, a slow reply is an annoyance. On a call, silence is interpreted: the caller wonders whether the line dropped, repeats themselves, or talks over the agent. A reply that comes too fast cuts them off mid-thought. Callers judge a voice agent against their instinct for human turn-taking.

Treat the delay a caller experiences as a latency budget spent across stages: telephony transport both ways, detecting the end of the caller's turn, finalising the transcript (or the model ingesting audio), the LLM's first useful tokens plus any tool call, and TTS producing the first audio chunk.

The practical rules: stream everything; measure time-to-first-audio, not total response time; keep model and voice services in a region near the telephony edge; keep prompts and history compact; and never leave the caller in silence for a slow tool. Have the agent say "Let me check that" while the lookup runs. Measure each stage on your own real calls.

End-of-turn detection

Simple voice activity detection (VAD) waits for silence, which fails both ways: people pause while recalling a date of birth, and a long threshold makes every reply feel slow. Better systems combine VAD with semantic signals: does the partial transcript sound complete, and is the caller still reading out a number? Tune per step; collecting a phone number needs more patience than a yes/no question.

Barge-in and interruptions

Callers interrupt: "no, no, the other branch". Handling barge-in means detecting caller speech while the agent talks (with echo cancellation so it does not hear itself), stopping playback quickly and discarding queued audio, and updating state to reflect what the caller actually heard rather than the full reply generated. Otherwise the agent believes it confirmed something the caller never heard. It should also ignore backchannels like "hmm", "okay" or "haan".

Telephony integration: SIP, PSTN and WebRTC

  • PSTN is the public telephone network ordinary calls travel over, reached through a telephony provider or carrier that gives you numbers.
  • SIP (Session Initiation Protocol) is the signalling standard for setting up and transferring calls over IP. Enterprises often connect a SIP trunk from their carrier or contact-centre platform to the agent's media server.
  • WebRTC is the browser and mobile standard for real-time audio, used for "talk to us" buttons in apps and websites, usually with better audio quality than phone lines.
  • Media streaming: many telephony platforms fork call audio to your service over a WebSocket, which is how the agent sends and receives audio in real time.

Phone audio is narrowband and compressed, so test STT on real call recordings. Plan for DTMF keypad input as a fallback for sensitive numbers and SIP transfer to human queues. Automated and outbound calling in India has its own regulatory requirements; your telephony partner and legal team should confirm what applies.

Reference architecture

 Caller (phone / app)
        |
  PSTN / SIP trunk  or  WebRTC
        |
 Media gateway  (audio stream)
        |
 Voice orchestrator
  |- VAD + end-of-turn
  |- barge-in control
  |- STT -> LLM -> TTS  (cascade)
  |   or realtime speech model
        |
 Agent logic + guardrails
  |- tools: lookup, book, cancel
  |- confirmation before actions
  |- handoff -> human queue (SIP)
        |
 Back-end APIs  (CRM, scheduling)
        |
 Logs: transcripts, tool calls,
       sampled audio, latency

The orchestrator owns call state, turn detection, barge-in and per-stage timing; the agent logic is a chat-style tool-calling agent constrained for voice.

To build the foundations this stack sits on, tool-calling agents, RAG, guardrails and evaluation, Cloudsoft's AI, GenAI and Agentic AI course works through them in hands-on labs.

Conversational voice AI in the Indian context

  • Code-mixed speech. Callers switch between English and Hindi, Telugu, Tamil or Kannada within one sentence: "Mera appointment Thursday ko reschedule karna hai". STT must not force a single language, and replies should match the caller's register.
  • Accents and audio. Regional accents, traffic noise, speakerphones and weak signal raise error rates. Build test sets from real calls across regions.
  • Numbers. Phone numbers, amounts and dates are spoken in mixed languages and groupings ("double five", digits in Hindi). Normalise carefully and read them back.
  • Names and addresses. Indian names, localities, landmarks ("opposite the metro station") and PIN codes are where ASR errors cluster. Use custom vocabulary where supported, constrain recognition with known lists such as branch and doctor names, and confirm by offering choices.

Detect language early, let the caller switch, and never ask them to confirm details in a language they did not choose.

Tool use during calls, with confirmation

Voice agents become useful when they can act: check availability, look up an order, book or cancel. The tool layer matches a chat agent's, with voice-specific rules:

  • Confirm before every write. Read back the key details ("Dr. Rao, Kukatpally branch, Thursday at four in the evening. Shall I book it?") and act only on a clear yes. A misheard entity is the most common way a voice action goes wrong.
  • Validate in code. The model proposes a structured call; your code checks the slot exists, the patient matches and the action is permitted before anything is written.
  • Fill the gap. Slow APIs are dead air; cache reference data and speak a short holding line.
  • Limit scope. No tool for anything you would not let it do unsupervised. Payments, refunds and medical advice usually stay with humans or secure non-voice channels.

Handoff to humans

A voice agent that traps callers does more damage than no agent. Transfer whenever the caller asks for a person, in any phrasing or language, and on repeated misunderstanding, out-of-scope requests, distress, complaints and anything safety-related. Make it a warm transfer: pass a structured summary (who, what they wanted, what was verified, what was attempted) to the human's screen so the caller does not repeat everything. If no one is available, offer a callback rather than looping.

This is not legal advice; rules differ by country, sector and call type and change over time, so check them with your legal and compliance teams. Common themes:

  • Disclose that the caller is speaking to an AI, plainly, at the start of the call.
  • Recording and transcription consent: tell callers the call is recorded or transcribed and why.
  • Sensitive data: health, financial and identity data in audio and transcripts need retention limits, access control, masking and data-residency decisions aligned with data protection law such as India's DPDP Act where applicable.
  • Outbound rules: automated and promotional calls are regulated in many countries, including India; confirm consent and registration requirements first.
  • Voices: use licensed voices and never imitate a real person without consent.

Evaluating voice AI agents

Voice evaluation extends AI agent evaluation with extra dimensions:

  • Task completion, verified against the back-end record (was the right appointment booked?), not the transcript.
  • Entity accuracy. Measure ASR errors on names, phone numbers, dates, times and branch names specifically. Overall word error rate hides the one wrong digit that breaks a booking.
  • Interruption handling. Script calls where the caller barges in, corrects themselves or goes silent, and check the agent's state matches what was said and heard.
  • Latency distribution of time-to-first-audio, not just averages; callers remember the slow turns.
  • Handoff correctness and summary quality, and results per language reviewed by native speakers.

Keep a regression set of recorded or synthetic calls with varied accents and noise, run it on every prompt, model or voice change, and combine LLM-as-judge scoring of transcripts with human review, as in LLM evaluation.

Observability: transcripts, audio and traces

Each call should produce a replayable trace: per-turn transcripts, LLM inputs and outputs, tool calls and results, guardrail decisions, barge-in events and per-stage latency. General practice is in AI observability; voice adds sampled audio (with consent and retention limits), because transcripts are wrong in exactly the places you need to debug; turn-level timing that shows whether delay came from STT, model, tool or TTS; outcome tags (completed, transferred, abandoned, failed); and redaction of sensitive entities in stored logs.

Cost per minute, conceptually

Voice AI is usually budgeted per call minute. A cascade minute is roughly telephony (carrier, numbers, SIP or streaming fees) plus STT by audio duration, LLM tokens that grow as history is re-sent each turn, and TTS by characters or duration. A speech-to-speech model replaces the middle three with realtime audio charges. Orchestration and storage come on top.

Levers: short spoken replies, summarised histories, cached reference data and smaller models for simple calls. Compare cost per successful task: a cheaper pipeline that fails more bookings costs more overall.

Illustrative example: appointment booking for a clinic chain

Illustrative scenario. Consider a clinic chain with branches across Hyderabad and other cities whose front-desk lines are busy every morning with booking, rescheduling and cancellation calls in English, Telugu, Hindi and mixes of all three. The voice agent handles out-patient bookings only: no medical advice, reports or billing disputes.

  1. The agent says it is an AI assistant, notes that the call is recorded, and asks for a preferred language.
  2. It identifies the patient by registered mobile number (from caller ID where available) and confirms with a second detail such as date of birth.
  3. The caller says, "Thursday evening, Dr. Lakshmi at the Kukatpally branch, for my mother." The agent matches doctor and branch against the clinic's lists, checks availability with a read-only tool, and offers two slots.
  4. The caller picks one; the agent reads back patient, doctor, branch, date and time and waits for a clear yes before calling the booking tool.
  5. Code validates the request, writes it to the scheduling system and sends an SMS confirmation.
  6. If the caller mentions chest pain or another urgent symptom, the agent stops, gives the clinic's approved emergency guidance, and transfers to staff immediately.

A cascade fits here: the clinic needs text-level guardrails against medical advice, STT tuned for Telugu-English mixing, and auditable transcripts. Evaluation tracks correct bookings in the scheduling system, entity accuracy on doctor names and dates, transfers on urgent symptoms, and recovery when callers interrupt the read-back.

Taking a system like this into a client's real telephony, scheduling systems and compliance review is the work Forward Deployed Engineers do, and several companies hiring for Forward Deployed Engineer roles build voice AI. Cloudsoft's FDE PRO program covers that delivery side.

Frequently asked questions

What is the difference between a cascade and a speech-to-speech voice agent?

A cascade converts speech to text, sends it to an LLM, and turns the reply into speech with TTS. A speech-to-speech model handles audio in and audio out within one model. Cascades give more control and text-level guardrails; speech-to-speech usually feels more natural but is harder to inspect and constrain.

Why does latency matter so much for voice AI agents?

Callers read silence as a dropped line or confusion, so they repeat themselves or talk over the agent. Stream every stage, measure time-to-first-audio per turn, and speak a short holding line when a tool call will take a moment.

How does a voice agent handle interruptions?

It detects caller speech while talking, uses echo cancellation, stops playback quickly, and updates its state to reflect only what the caller actually heard. It ignores short backchannels such as okay or hmm.

How do I connect an AI phone agent to real phone numbers?

Use a telephony provider that connects the PSTN to your system through a SIP trunk or a media streaming API over WebSocket. For apps and websites, WebRTC carries audio directly. SIP transfer hands calls to human queues.

Can voice AI agents handle Indian languages and code-mixed speech?

Yes, with deliberate engineering: STT and TTS that support your languages, testing on real call audio across accents, custom vocabulary for names and places, and read-back of numbers, dates and addresses. Review results per language with native speakers.

Do I have to tell callers they are speaking to an AI?

Disclosing it plainly at the start is sound practice everywhere, and some jurisdictions and sectors require it, along with recording consent. Rules vary, so check what applies with your legal and compliance teams.

How should voice AI agents be evaluated?

Measure task completion against back-end records, ASR accuracy on key entities, interruption recovery, latency distribution, handoff correctness and results per language, and rerun a regression set of test calls on every change.

How is the cost of a voice AI agent calculated?

Per call minute: telephony, STT, LLM tokens and TTS for a cascade, or realtime audio charges for speech-to-speech, plus orchestration and storage. Compare cost per successfully completed task, since failed calls still cost money.

Ready to build voice and chat agents that call real tools, pass evaluation and hold up in production? Cloudsoft's GenAI and Agentic AI training in Hyderabad teaches agents, RAG, guardrails and evaluation through hands-on labs, in our Ameerpet classroom or live online. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us