Voice AI interview questions in 2026 test whether you can make a spoken agent feel responsive and stay safe on a real phone line: streaming speech recognition, endpointing and barge-in, a latency budget measured to the first byte of audio, identity checks that never put an OTP in a prompt, and a clean handoff to a human. This guide collects 50 high-value questions with model answers, from the ASR β LLM β TTS cascade versus speech-to-speech models, through WER and code-mixed Telugu-English, SSML and voice cloning consent, SIP, PSTN and WebRTC, to evaluation with simulated callers, AI disclosure and recording consent, cost per minute and ten production scenarios.
How to use this guide
These questions come up in voice agent interviews for conversational AI engineers, AI engineers on contact-centre platforms, speech engineers and Forward Deployed Engineers who put voice bots into banks, hospitals, insurers and retailers. The architecture background is in our explainer on voice AI agents; this page turns it into interview answers and goes deeper on the parts interviewers probe. What is typically tested at each level:
- Freshers and developers new to speech: what ASR, TTS, VAD and WER are, why a voice reply must be shorter than a chat reply, and why a phone line is not the same as a laptop microphone.
- Mid-level engineers: streaming pipelines, endpointing trade-offs, barge-in mechanics, custom vocabulary, SSML, DTMF fallback and how tools are called mid-call with confirmation.
- Senior engineers and architects: cascade versus speech-to-speech decisions, per-stage latency percentiles, evaluation with simulated callers, identity verification design, compliance and cost per successful call.
Speech recognition interview rounds often overlap with NLP and multimodal rounds, so where a concept belongs there we link rather than repeat. Answer each question aloud before reading the model answer. For a voice role, saying it out loud is the practice.
- Fundamentals (Q1βQ7)
- Architecture: cascade vs speech-to-speech (Q8βQ12)
- Latency, endpointing and barge-in (Q13βQ18)
- Speech recognition (ASR) (Q19βQ24)
- Text-to-speech and voices (Q25βQ27)
- Telephony integration (Q28βQ30)
- Dialogue design, tools and handoff (Q31βQ35)
- Evaluation, compliance and cost (Q36βQ40)
- Real-world scenarios (Q41βQ50)
- Key takeaways
- Interview preparation checklist
- FAQ
Fundamentals
1. What is a voice AI agent, and how is it different from a traditional IVR?
Answer: A voice AI agent holds a real-time spoken conversation: it listens, works out what the caller wants in their own words, calls back-end tools, and answers in a synthesised voice. A traditional IVR plays a fixed menu ("press 1 for billing") and follows a decision tree, maybe with keyword speech recognition. The agent keeps context across turns, asks clarifying questions, handles corrections ("no, the other account") and can complete a task end to end. Underneath, it is an LLM agent whose input and output are audio and which must answer within a conversational time limit. It still needs the boring parts of an IVR: authentication, routing rules, business hours and a way to reach a human.
Interview tip: Say that the IVR is not thrown away. Most production voice bots keep DTMF menus as a fallback and reuse the IVR's routing and queue logic.
2. Walk me through the components of a voice agent pipeline.
Answer: Audio arrives from a phone call or a browser. Voice activity detection (VAD) decides when the caller is speaking. Streaming automatic speech recognition (ASR, or speech-to-text) produces partial and final transcripts. An endpointing step decides the caller has finished their turn. The LLM, with system prompt, history and tool definitions, decides what to say or which tool to call. Text-to-speech (TTS) turns the reply into audio, streamed back sentence by sentence. Around that sit an orchestrator that manages turn state and interruptions, tool integrations, guardrails, logging and tracing, and a transfer path to human agents.
Caller audio -> VAD -> streaming ASR -> endpointing -> LLM (+ tools, guardrails) -> streaming TTS -> caller (barge-in: caller speech stops TTS)
3. What do ASR, TTS, VAD and diarisation each mean?
Answer: ASR converts speech to text. TTS converts text to speech. VAD classifies short audio frames as speech or non-speech, which drives turn detection and barge-in. Diarisation answers "who spoke when", splitting a recording into speaker segments. In a live two-party call, diarisation is often unnecessary because each side arrives on its own channel; it matters for mono recordings, conference calls, and post-call analytics such as QA scoring.
4. Why can't you take a chat agent and just add speech on both ends?
Answer: Because voice removes everything chat relies on. There is no screen, so a bulleted list or a table becomes a monologue the caller cannot scroll back through. There is a hard time limit: silence of a couple of seconds feels broken, while a chat user happily waits. Input is noisier: ASR errors on names, numbers and dates arrive as confident-looking text. Callers interrupt, mumble, talk to someone else in the room and say "haan" mid-sentence. And mistakes are heard immediately, by a customer who may be recording the call. The agent logic carries over; prompts, response length, confirmation style, error recovery and latency engineering all have to be redesigned.
5. What does "time to first audio" mean, and why is it the key latency metric?
Answer: Time to first audio (sometimes called voice-to-voice latency) is the time from when the caller stops speaking to when they hear the first sound of the reply. It is what the caller experiences as "the bot is slow". Total generation time matters less, because a streamed reply that starts quickly feels responsive even if it continues for several seconds. Measure it end to end at the caller's side of the media path, not just inside your server, and break it down per stage so you know whether endpointing, ASR, the LLM, a tool call or TTS is responsible.
6. What is barge-in?
Answer: Barge-in is the caller interrupting while the bot is speaking. A good agent detects the caller's speech, stops its own audio quickly, discards queued audio, and listens. Without it, callers have to wait out every long sentence, and they will press zero or hang up. Barge-in is harder than it sounds on a phone line because the bot's own voice can echo back into the caller channel and trigger the VAD, and because short acknowledgements ("okay", "hmm") should not always stop the bot. Q17 covers the mechanics.
7. Where do voice agents typically deliver value in an enterprise?
Answer: High-volume, well-defined call types where the back-end action is clear: appointment booking and rescheduling, order and delivery status, card blocking, balance and due-date queries, payment reminders, IT helpdesk password resets, and front-door triage that collects the reason for the call before routing to the right human queue. They deliver less value on complaints, disputes, bereavement, complex advice or anything where the caller needs judgement or empathy over time. A strong answer names one call type, the system of record it touches and the outcome metric, for example "bookings completed without transfer, verified against the scheduling system".
Real-world example: Consider a hospital chain whose reception lines are swamped every morning. Scoping the agent to out-patient bookings only, with no clinical questions, makes it safe to launch and easy to measure.
Architecture: cascade vs speech-to-speech
8. Compare the ASR β LLM β TTS cascade with a speech-to-speech (realtime) model.
Answer: The cascade chains three components with text in between. A speech-to-speech model takes audio in and produces audio out within one model, usually over a persistent streaming connection, often with tool calling and a parallel transcript. Several model providers now offer realtime APIs of this kind.
| Factor | Cascade (ASR β LLM β TTS) | Speech-to-speech |
|---|---|---|
| Latency | Each hop adds delay; needs careful streaming | Fewer hops, typically more natural turn-taking |
| Control and guardrails | Text between every stage to check, log and redact | Inserting a text check gives back part of the latency gain |
| Paralinguistics | Tone, hesitation and emphasis mostly lost | Can respond to how something was said |
| Component choice | Choose the ASR, LLM and voice that suit your languages | Coupled to one provider's voices, languages and pricing |
| Debugging | Per-stage transcripts and timings | Transcript may not match the audio the caller heard |
| Indian languages | Can pick an Indic-strong ASR and TTS | Must be tested per language and code-mix |
I default to the cascade for regulated, tool-heavy flows (banking, insurance, healthcare) and evaluate speech-to-speech for conversational, lower-risk flows where naturalness matters most. Either way the outer system is the same: telephony, tools, verification, logging and handoff.
Interview tip: Avoid declaring one architecture "better". Interviewers want the decision criteria and a statement that you would measure both on your own call recordings before choosing.
9. How do you build guardrails into a speech-to-speech pipeline?
Answer: Layer them where you can without blocking speech. Before the model: authentication state, tool permissions and a tightly scoped system prompt, so the model cannot do much harm even if it says something wrong. During the call: tool calls go through your own server, which enforces policy (amount limits, allowed record types) regardless of what the model asked for. On output: run the parallel text transcript through fast checks and, if a check fails, interrupt the audio and replace it, accepting that a fragment may already have been heard. After the call: review audio-based transcripts, not just the model's own transcript. For flows where any wrong sentence is unacceptable, such as reading a policy clause, use the cascade or a fixed pre-approved prompt for that step. The general patterns are in AI guardrails.
10. What does a reference architecture for a production voice agent look like?
Answer: Telephony provider or SIP trunk at the edge, a media gateway that converts call audio into a stream your service can consume, a voice orchestrator (stateful per call) that runs VAD, ASR, turn-taking, the LLM loop and TTS, a tool layer behind an API gateway with its own authorisation, a session store for call state, and an observability pipeline for traces, redacted transcripts and sampled audio. A warm-transfer path connects back to the contact-centre platform with a summary.
PSTN/SIP or WebRTC
-> media gateway (audio stream)
-> voice orchestrator (per-call state)
ASR | turn-taking | LLM | TTS
-> tool API gateway -> CRM, core, booking
-> transfer -> contact centre + summary
-> traces, redacted transcripts, audio
The orchestrator should run in a cloud region close to callers and to the telephony provider's media servers, because every network round trip sits inside the latency budget.
11. Why is a voice orchestrator stateful, and how do you scale it?
Answer: Each call holds a long-lived bidirectional audio stream, ASR and TTS connections, turn state, the conversation history and a pending interruption flag. You cannot load-balance individual audio frames across instances. So you scale by concurrent calls: sticky routing of a call to one instance for its lifetime, autoscaling on active sessions rather than CPU alone, graceful draining during deploys (stop accepting new calls, let current ones finish), and checkpointing call state externally so that if an instance dies you can at least transfer the caller with context instead of dropping them. Capacity planning starts from peak concurrent calls, not requests per second.
12. When would you use a smaller or self-hosted model in a voice agent?
Answer: When latency, cost or data residency dominate. A smaller model often has a faster time to first token and is good enough for routing, slot filling and short confirmations, while a larger model handles complex reasoning turns. Self-hosting near your media servers removes a network hop and keeps audio and transcripts in your environment, but you take on GPU capacity, batching and peak-hour scaling. A common split is a fast model for the conversational turn and a larger model or a deterministic workflow for the hard decisions. Serving trade-offs are covered in the LLM inference and serving interview questions.
Latency, endpointing and barge-in
13. Build a latency budget for one turn of a cascade voice agent.
Answer: Break the gap between end of caller speech and first audio into stages and give each a target, then measure the real distribution of each.
| Stage | What drives it | Main levers |
|---|---|---|
| Endpointing | Silence threshold or turn-detection model | Adaptive thresholds, semantic turn detection |
| ASR final | Time to finalise the last words | Streaming ASR, regional deployment |
| LLM time to first token | Model size, prompt length, queueing | Smaller model, prompt caching, short history |
| First speakable chunk | Tokens until a sentence or phrase boundary | Prompt for short openings, chunk on clauses |
| TTS first audio | Voice model, streaming support | Streaming TTS, warm connections |
| Network and telephony | Region, carrier, jitter buffers | Co-locate, keep connections open |
The targets themselves depend on the use case; the method is to set a voice-to-voice target with the business (for example, "most turns well under a second, and no turn without audio feedback for more than about two seconds"), then hold each stage to its share. Tool calls are budgeted separately because they can blow the whole budget on one turn. General LLM techniques are in LLM latency optimisation.
14. What is endpointing, and what is the trade-off in tuning it?
Answer: Endpointing decides that the caller has finished their turn so the agent can respond. The simplest version waits for a fixed silence after speech. A short threshold makes the bot fast but cuts people off when they pause to think or read out a number; a long threshold avoids interruptions but adds that silence to every turn. Better approaches adapt: a longer threshold when the bot has just asked for an account number or address, a shorter one after a yes/no question, and a turn-detection model that uses the partial transcript ("my number is nine eight four..." is clearly unfinished) and sometimes prosody to predict end of turn. Measure both false endpoints (bot interrupts) and late endpoints (dead air).
Interview tip: Mention digit strings specifically. Indian callers often read a mobile number in groups with pauses, which a fixed short threshold will chop into fragments.
15. How does streaming ASR reduce latency, and what are partial versus final transcripts?
Answer: Batch ASR waits for the whole utterance; streaming ASR transcribes while the caller speaks, so at the endpoint only the last fraction of a second needs finalising. Streaming engines emit partial hypotheses that may change ("I want to cancel" may become "I want to cancel my Thursday slot") and then a final, more stable transcript for each segment. You use partials for turn detection, barge-in decisions and, carefully, for speculative work such as pre-fetching the caller's record or starting an LLM call that you abandon if the final transcript differs materially. Never execute a tool action on a partial.
16. How do you start speaking before the LLM has finished its answer?
Answer: Stream LLM tokens into a chunker that releases text to TTS at natural boundaries: a sentence end, or a clause boundary after a minimum length, so the voice does not sound choppy. Prompt the model to lead with a short, complete opening sentence ("Sure, let me check that.") and put detail after it. Keep TTS connections warm, and pipeline so the next chunk is synthesised while the first plays. Two cautions: anything already spoken cannot be unsaid, so output guardrails must run on each chunk before it reaches TTS; and if a tool call is needed, emit a brief holding phrase, then the result, rather than letting the model speculate about the answer.
17. Explain how you implement barge-in correctly.
Answer: Five parts. First, run VAD and ASR on the caller channel continuously, even while TTS is playing. Second, remove the bot's own voice from that channel with acoustic echo cancellation, or by relying on separate channels where the telephony stack provides them, so the bot does not interrupt itself. Third, classify the interruption: a backchannel ("okay", "hmm", "haan") may only pause or be ignored, while real speech or a short "stop" or "wait" halts playback. Fourth, on a real barge-in, stop audio within a short time, flush queued TTS audio, and cancel in-flight LLM generation. Fifth, and most often missed, truncate the assistant's message in the conversation history to what the caller actually heard, using playback position. Otherwise the model believes it said things the caller never heard and the dialogue drifts.
18. What fillers and feedback do you use when a tool call is slow?
Answer: Silence on a phone line reads as a dropped call. If a tool call may exceed the latency budget, the agent says a short, honest holding phrase ("Let me check your booking, one moment") before calling, and if it runs long, a second update. Some teams play a soft background tone. Vary the phrases so they do not sound robotic, and never fill time with invented content. On the engineering side: set per-tool timeouts, pre-fetch likely data once the caller is identified, run independent lookups in parallel, and if a tool fails, say so and offer a callback or a human rather than retrying silently.
Speech recognition (ASR)
19. What is word error rate, and how is it calculated?
Answer: WER is the number of substitutions, deletions and insertions needed to turn the ASR output into the reference transcript, divided by the number of words in the reference: (S + D + I) / N. Lower is better, and because insertions count, WER can exceed one. It is computed after text normalisation (case, punctuation, number formats), and the normalisation rules change the score, so two vendors' WER figures are only comparable on the same test set with the same normaliser. For scripts and languages where word boundaries are ambiguous, character error rate (CER) is often reported alongside.
20. What are the limits of WER as a quality metric for a voice agent?
Answer: WER weights every word equally, but the agent does not. Mishearing "the" is harmless; mishearing one digit in a policy number or "cancel" as "can't sell" breaks the task. A system can have an acceptable overall WER and still fail on exactly the entities that matter. So I add entity-level metrics: accuracy on names, numbers, dates, amounts and domain terms; intent accuracy (did the downstream system understand the request); and task success. WER also hides distribution: an average across languages can conceal that one accent group or noisy-line segment is much worse. Report WER per language, accent, channel (landline, mobile, app) and noise condition, on your own recordings.
Interview tip: Saying "WER is necessary but not sufficient; I measure entity accuracy and task success" is one of the clearest signals that you have shipped a voice system.
21. How do you handle code-mixed Indian languages such as Telugu-English or Hindi-English?
Answer: Code-mixing means switching languages within a sentence: "Naa card block cheyyali, last transaction dispute undi." Practical steps: choose an ASR that explicitly supports the code-mixed pair, not just each language separately, and test it on your own call audio. Decide the output script convention (Telugu words in Telugu script, English words in Latin script, or everything romanised), because the LLM and your entity parsers must expect it, and because a script mismatch inflates WER without any real error. Use language identification per segment where the engine supports it. Bias recognition towards your domain vocabulary. Make sure the LLM is prompted, and tested, to understand mixed input and to reply in the caller's preferred language mix. Measure per language pair with native-speaker reviewers.
22. What is custom vocabulary or contextual biasing, and when do you need it?
Answer: Most ASR services let you supply phrase lists or hints that raise the probability of specific words: product names, branch and locality names, doctor names, plan codes. You need it whenever your entities are rare in general speech, which in Indian enterprises is constant: locality names, Indian personal names and brand-specific terms. Make the list dynamic where possible, for example the doctor names of the branch the caller has just chosen, because an oversized static list can cause false matches. Where biasing is not enough, options include adapting or fine-tuning a model on in-domain transcribed audio (with consent for that use), and post-ASR correction by fuzzy-matching against the known entity list before a tool call.
23. How do accents, noise and narrowband phone audio affect ASR, and what do you do about it?
Answer: Traditional phone calls carry narrowband audio (telephony codecs such as G.711 sample at 8 kHz), which drops high frequencies that help distinguish consonants. Mobile calls add compression, packet loss and background noise: traffic, TV, a crowded shop. Accent variation across regions adds another layer. Mitigations: pick models trained or tuned on telephony audio rather than studio speech; evaluate on real calls from your caller population; avoid aggressive noise suppression that damages speech; design the dialogue to confirm critical entities; and offer DTMF for digits when recognition confidence is low. Do not resample narrowband audio upward and expect wideband accuracy.
24. When does diarisation matter in voice AI, and what makes it hard?
Answer: In a live bot call it rarely matters, because the bot knows its own audio and the caller arrives on a separate stream. It matters in post-call analytics on mono recordings (who said the disclosure, the agent or the customer), in conference calls, and when a caller hands the phone to a relative mid-call. It is hard with overlapping speech, short turns, similar voices and code-switching. If you control the recording, record stereo (one channel per party) so you do not need to diarise at all. Our call centre QA AI project shows why speaker attribution is critical when scores are about named employees.
Text-to-speech and voices
25. What makes a TTS voice suitable for a production voice agent?
Answer: Intelligibility on a phone line first, naturalness second. Check: streaming support and low time to first audio; correct pronunciation of names, numbers, dates, currency ("βΉ2,499" read naturally) and abbreviations in your languages; natural prosody (stress and intonation) so questions sound like questions; consistent voice across turns; support for each language and code-mixed text you need; licensing terms that allow your commercial use; and stable output so the same sentence does not sound different every time. Evaluate with listeners from your caller population, over the actual telephony codec, not on studio headphones.
26. What is SSML, and how would you use it in a voice agent?
Answer: Speech Synthesis Markup Language is a W3C XML standard that many TTS engines support, with varying coverage. It controls how text is spoken: <break> for pauses, <say-as> to read content as digits, a date or a telephone number, <prosody> for rate and pitch, <emphasis>, <sub> for aliases, and <phoneme> for exact pronunciation of a brand or place name.
<speak>Your slot is on
<say-as interpret-as="date">2026-10-08</say-as>
<break time="300ms"/> at the
<sub alias="Kukatpally Housing Board">KPHB</sub>
branch.
</speak>
In an LLM pipeline, do not let the model write free-form SSML; it can produce invalid markup. Have the model output plain text or structured fields, then a deterministic formatter applies SSML for dates, amounts and confirmations. Check which tags your chosen engine actually supports.
27. What are the consent and security risks of voice cloning, and how do you handle them?
Answer: Voice cloning creates a synthetic voice from recordings of a real person. Risks: using someone's voice without informed, written consent; a voice actor's voice being used beyond the agreed scope or after the contract ends; impersonation of executives or relatives in fraud; and defeating voice-biometric authentication. Controls: written consent that names the purpose, duration and revocation; use licensed stock voices for most agents; never clone a real person to make callers think they are speaking to that person; restrict access to voice models like any other credential; and, where the provider supports it, watermark or label synthetic audio. On the receiving side, treat voice biometrics as one weak factor, not sole proof of identity. Provenance and labelling of synthetic media are covered in AI content provenance and watermarking.
Telephony integration
28. Explain SIP, PSTN and WebRTC in the context of a voice agent.
Answer: The PSTN is the public telephone network that ordinary phone numbers live on. SIP (Session Initiation Protocol) is the signalling protocol most business telephony uses to set up, transfer and end calls; the audio itself usually travels as RTP media. A SIP trunk connects your platform to a carrier so calls from the PSTN reach you. WebRTC is the browser and mobile-app standard for real-time audio and video, usually with wideband codecs such as Opus, and handles NAT traversal and encryption. Integration options, described generally: a cloud telephony or contact-centre provider that forwards call audio to your service over a streaming connection; a SIP trunk into your own media server; or WebRTC for in-app voice. The choice affects audio quality, latency, transfer options and who holds the recording.
Interview tip: You do not need to recite SIP headers. Interviewers want to hear that PSTN audio is narrowband, that transfer and hang-up are signalling events your orchestrator must handle, and that in-app WebRTC gives better audio than a phone call.
29. What is DTMF, and when do you fall back to it?
Answer: DTMF (dual-tone multi-frequency) is the tones sent when a caller presses keypad buttons; on IP telephony it is commonly carried as RTP events or SIP messages rather than audio. Fall back to it when speech is unreliable or inappropriate: entering an account number or OTP after repeated misrecognition, noisy environments, callers who prefer it, and sensitive data you do not want spoken aloud or passed through ASR. Design the dialogue so the switch is natural ("You can also type the number on your keypad"), configure the platform to deliver DTMF digits directly to your tool layer, and suppress the tones from recordings where payment data is involved.
30. How do you handle call transfer technically?
Answer: Two forms. A cold (blind) transfer hands the call to a queue or number with nothing else. A warm transfer passes context: the orchestrator writes a summary, the verified identity status, the intent and collected details to the contact-centre platform (as call-attached data or a CRM record keyed to the call), then transfers via SIP signalling or the provider's API. The human agent sees the summary before they speak. Handle edge cases: queue full or out of hours (offer a callback), transfer failure (do not drop the caller; tell them and retry or schedule a callback), and the caller hanging up mid-transfer (log it as abandoned, not resolved).
Dialogue design, tools and handoff
31. How should prompts and replies differ for voice compared with chat?
Answer: Short turns: one or two sentences, one question at a time. No markdown, lists, URLs or tables, because they will be read out literally. Put the most important information first. Offer at most two or three options in a turn, since callers cannot hold more in memory. Spell out numbers in the way people say them, and read back critical entities in chunks ("nine eight four, eight six..."). Use natural spoken acknowledgements, and match the caller's language. Tell the model explicitly in the system prompt that it is on a phone call, that the caller cannot see anything, and how long replies may be, then enforce length in the formatter as well.
32. How do you design confirmations so the bot is safe but not tedious?
Answer: Confirm in proportion to risk and confidence. For irreversible or costly actions (cancelling a policy, blocking a card, a payment) use explicit confirmation: read back the key details and ask for a clear yes. For low-risk steps use implicit confirmation by folding the understood value into the next question ("Thursday evening at Kukatpally. Which doctor?"), which lets the caller correct it without an extra turn. Use ASR confidence and entity validation: a date in the past or a branch that does not exist triggers a clarifying question. After two failed attempts on the same entity, change strategy: spell it, use DTMF, offer options, or transfer.
33. How do tools and actions work during a call, and how do you keep them safe?
Answer: The LLM calls functions such as find_patient, get_slots or block_card with structured arguments, and your server executes them. Safety comes from the server, not the prompt: each tool checks the session's verified identity and permissions, validates arguments against schemas and business rules, enforces limits, and is idempotent so a retried call after a timeout does not book twice. Read tools can run freely after verification; write tools require the confirmation step from Q32 and are logged with the call ID. Return concise, speakable results to the model and keep sensitive fields out of its context entirely. See function calling and structured outputs for the mechanics.
34. How do you verify a caller's identity in a voice agent, including OTPs?
Answer: Combine factors appropriate to the risk of the action. Caller ID matched to a registered number is a hint, not proof, because it can be spoofed. Knowledge factors (date of birth, last transaction) are weak alone. A one-time password sent to the registered mobile is stronger. Design rules for OTPs: the caller enters the OTP by DTMF where possible; the digits go straight from the telephony layer to a verification service; the LLM receives only "verified: true/false", never the OTP; the OTP never appears in transcripts, logs or recordings; attempts are rate-limited and expire. Never let an outbound bot ask a customer to read out an OTP, because that is exactly the pattern fraudsters use and customers are told to refuse. Step up verification for sensitive actions instead of verifying everything at the start.
caller asks to block card -> tool: send_otp(registered mobile) -> caller types OTP (DTMF) -> telephony -> verify service -> LLM sees: verified=true -> confirm -> block_card()
35. When and how should a voice agent hand off to a human?
Answer: On the caller's request, every time, without arguing or making them repeat "agent" three times. Also on triggers: repeated misunderstanding, low confidence on a critical entity, anger or distress, topics outside scope (complaints, medical or financial advice, legal threats), failed verification, vulnerable-customer signals, and any action the policy reserves for humans. Hand off warm, with a summary and verified status so the caller does not start again. Measure handoff correctness both ways: calls that should have been transferred but were not, and calls transferred unnecessarily. The wider design of escalation and approvals is in human-in-the-loop AI.
Evaluation, compliance and cost
36. How do you evaluate a voice agent before and after launch?
Answer: On several layers. Task success, verified against the system of record (was the right slot actually booked?), not judged from the transcript. Entity accuracy on names, numbers and dates. Latency as percentiles: p50, p90 and p95 of time to first audio per turn, plus the worst turn per call, because callers remember the slowest moment. Interruption handling, handoff correctness, containment (calls completed without transfer) balanced against complaints and repeat calls, and safety checks for disclosures and prohibited statements. All of it per language and channel. Pre-launch, use a regression suite of recorded and simulated calls; post-launch, sample real calls for human review and track outcome tags. The general agent methods are in AI agent evaluation.
Interview tip: Containment alone is a trap metric. A bot that never transfers has high containment and angry customers. Pair it with task success and repeat-call rate.
37. How do you build simulated callers for voice agent testing?
Answer: An LLM plays the caller from a persona and goal ("elderly caller, Telugu-English, wants to move her mother's appointment, gives the wrong date first"), its text goes through TTS with varied voices, accents, speaking rates and added noise or codec degradation, and the audio is streamed into your agent over the same path as real calls. The harness scripts behaviours: barge-in at random points, long pauses mid-number, self-correction, silence, background talk. Assertions check the back-end end state, tool calls, latency percentiles and policy. Simulated callers give coverage and repeatability, but they speak more cleanly than real people, so calibrate against real recorded calls and native-speaker review. More on generating test data safely is in synthetic data for AI testing.
38. What compliance points apply to call recording and transcripts in India?
Answer: This is not legal advice; involve legal and compliance. Engineering themes: tell callers at the start that the call is recorded or transcribed and why; under India's DPDP Act, recordings and transcripts are personal data, so you need a lawful basis such as consent with a clear notice, purpose limitation, retention limits, erasure, security safeguards and breach response that reach every copy (audio store, transcripts, LLM logs, traces, eval datasets). The DPDP Rules were notified in November 2025 with most business obligations applying from May 2027. Sector regulators (banking, insurance, health) may add retention and audit rules, and outbound and promotional calls fall under telecom regulator rules on commercial communication. Engineering detail is in DPDP Act for AI applications.
39. Do you have to tell callers they are talking to an AI? What does the EU AI Act say?
Answer: Disclose it plainly at the start of every call; it is good practice everywhere and trust collapses when callers discover it later. Under the EU AI Act, Article 50 sets transparency duties: providers must design AI systems that interact directly with people so that those people are informed they are dealing with an AI, unless that is obvious from context, and there are related duties to mark synthetic audio and to disclose deep fakes. Article 50 duties apply from August 2026. This matters to Indian teams that build or run voice agents serving people in the EU. If the bot is asked "Am I talking to a real person?", it must answer truthfully. See the EU AI Act for Indian IT teams for roles and timelines, and check current guidance.
40. How do you think about cost per minute for a voice agent?
Answer: Build it bottom-up for one call minute, then convert to cost per successful task. For a cascade: telephony (carrier minutes, numbers, SIP or media streaming fees), ASR billed by audio duration, LLM tokens (which grow each turn because history is re-sent), TTS billed by characters or audio, plus orchestration compute, storage and observability. A speech-to-speech model replaces the middle three with realtime audio pricing.
Real-world example (illustrative, placeholder rates): Suppose telephony costs βΉT per minute, ASR βΉA per audio minute, TTS βΉS per minute of speech, and the LLM βΉL per turn with about N turns per minute. A cascade minute is roughly T + A + (S Γ share of time the bot speaks) + (L Γ N) + overhead. If a cheaper configuration completes fewer bookings and transfers more calls, its cost per successful booking (including the human minutes of transferred calls) can be higher. Plug in your vendors' current rates; levers are shorter replies, summarised history, prompt caching, a smaller model for simple turns and early routing of out-of-scope calls.
If you want to practise this end to end, from streaming pipelines and tool calling to evaluation and cloud deployment, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers the engineering foundations behind production AI systems, in Ameerpet or live online.
Real-world scenarios
41. Callers keep talking over the bot, and transcripts show the bot answering questions nobody asked. What do you do?
Answer: This is a barge-in and turn-state problem. The likely causes are that interruptions are not stopping playback fast enough, the bot's own audio is leaking into the caller channel, the history still contains the full reply the caller never heard, or partial transcripts from overlapping speech are being treated as complete turns.
What I would check:
- Traces for barge-in events: time from caller speech onset to TTS stop, and whether queued audio and in-flight LLM generation were cancelled.
- Whether echo from the bot's own voice triggers VAD (self-interruptions with no caller speech in the audio).
- Whether assistant messages are truncated in history to the played portion.
- How backchannels like "haan" and "okay" are classified.
- Reply length: long monologues invite interruptions.
Production consideration: Fix the mechanics first, then shorten replies and ask one question per turn. Add simulated callers who interrupt at random points to the regression suite so this does not return.
42. WER on Telugu-English calls is far worse than on English calls, and bookings fail. How do you approach it?
Answer: First establish whether the problem is real recognition error, a measurement artefact, or entity failure. Code-mixed WER is often inflated by script and normalisation mismatches (the reference has Telugu words in Telugu script, the output romanises them), so fix the evaluation before the model.
What I would check:
- A labelled sample of real Telugu-English calls with a consistent script convention, reviewed by native speakers.
- WER and CER after consistent normalisation, plus entity accuracy on names, localities, dates and numbers.
- Whether the ASR is configured for the code-mixed pair or forced to one language.
- Custom vocabulary for branch, locality and doctor names, loaded dynamically per context.
- Whether the LLM understands the transcript even when ASR is imperfect, and whether failures cluster in specific entities.
Production consideration: Compare candidate ASR engines on this test set, add fuzzy matching to known entities before tool calls, use explicit confirmation and DTMF fallback for numbers, and report quality per language on the dashboard so the gap stays visible. The text side of code-mixed language handling is covered in the NLP interview questions guide.
43. During testing, the bot reads a customer's full card number back to them. What is wrong, and how do you fix it?
Answer: Sensitive data has reached the model's context and nothing stops it reaching TTS. Speaking a full card number exposes it to anyone near the caller and puts it in recordings, transcripts and logs, creating a payment-card compliance problem (PCI DSS) and a privacy one.
What I would check:
- Where the number entered the context: spoken by the caller and transcribed, or returned by a tool.
- Whether tools return full card numbers when the last four digits would do.
- Whether redaction runs on transcripts, prompts, logs, traces and recordings.
- Whether an output check blocks long digit sequences before TTS.
Production consideration: Tools should return masked values only; card entry should use DTMF routed to a payment service with tones suppressed and recording paused or masked; ASR output should be redacted before the LLM; and a deterministic output filter should block card-like numbers from being spoken at all. Confirm with "the card ending in four three two one". Treat this as a release blocker and add a test for it.
44. Callers in rural areas experience more than two seconds of latency before the bot responds. How do you investigate?
Answer: Split the delay into what you control and what the network adds. Rural mobile networks bring jitter, packet loss and fluctuating bandwidth, which enlarge jitter buffers and can also degrade ASR, causing re-prompts that feel like latency.
What I would check:
- Per-stage latency percentiles for affected calls versus urban calls: endpointing, ASR final, LLM first token, TTS first audio, tool calls.
- Telephony and media metrics: jitter, packet loss, round-trip time, codec in use.
- Whether endpointing waits longer because noise keeps VAD from detecting silence.
- Whether the orchestrator, ASR, LLM and TTS run in the same Indian cloud region as the carrier's media servers.
- Re-prompt rate due to recognition failures on those calls.
Production consideration: Co-locate all components in-region, keep connections warm, tune VAD for noisy audio, use shorter replies and a faster model, and play an immediate acknowledgement so the caller hears something quickly. For in-app calls, WebRTC with a resilient codec helps. Offer DTMF and a callback option when the line is poor.
45. A caller asks, "Am I talking to a real person?" and the bot says yes. What do you change?
Answer: This is a disclosure failure with legal and trust consequences. The model was probably given a persona with a human name and no explicit instruction, or it is mirroring the caller.
What I would check:
- The system prompt and persona text for anything implying a human.
- Whether the opening disclosure was played on this call.
- Eval coverage for "are you human" questions in each language.
Production consideration: Play a fixed, pre-recorded disclosure at call start, independent of the model. Add an explicit instruction and a deterministic check for questions about being human, answered with an approved truthful line and an offer to transfer. Add these probes in every supported language to the regression suite.
46. Transferred callers complain they have to repeat everything to the human agent. How do you fix the handoff?
Answer: The transfer is cold in practice even if it is warm in design: the summary is missing, late or not displayed.
What I would check:
- Whether the summary and verified identity status are attached to the call before the transfer, keyed to the right call ID.
- Whether the contact-centre desktop displays that data, and how agents are trained to use it.
- Summary quality: intent, collected details, what was already tried, why it transferred.
- Whether verification is honoured or policy forces re-verification (sometimes a legitimate rule, which should be explained to callers).
Production consideration: Generate the summary as structured fields, not prose, write it before initiating the transfer, and have the bot tell the caller what the human will already know. Track the repeat-question rate on transferred calls.
47. A model upgrade improves average latency, but complaints rise. What happened?
Answer: Averages hide tails and quality. Possible causes: worse tail latency, more verbose replies, weaker tool calling, different handling of code-mixed input, or a voice or speaking style callers like less.
What I would check:
- p90 and p95 time to first audio and worst-turn-per-call, not the mean.
- Task success and handoff rates per language before and after.
- Reply length and interruption rate.
- Tool-call argument errors and confirmation failures.
Production consideration: Gate model changes on the full voice regression suite, roll out to a small share of traffic with a fast rollback, and compare outcome metrics, not just speed.
48. An outbound payment-reminder bot is being planned. What do you raise before building it?
Answer: Outbound changes the risk profile: the bot calls people who did not ask to speak to it, about money.
What I would check:
- Consent and telecom rules for automated and commercial calls, calling hours and do-not-call handling, with the compliance team.
- How the bot proves it is from the lender without asking the customer for secrets; it must never ask for an OTP, PIN or card details.
- Right-party verification before any account detail is disclosed, since a family member may answer.
- Scripts for hardship, disputes and vulnerable customers, with transfer to a human.
- Rate limits on call attempts per customer.
Production consideration: Disclose AI and recording at the start, keep a full audit trail, and measure complaints and opt-outs alongside promise-to-pay outcomes.
49. Calls from a retail store's customers have constant TV and crowd noise, and the bot keeps responding to the background. What do you do?
Answer: Background speech is being treated as the caller. VAD detects speech, not the caller's speech.
What I would check:
- Recordings of false triggers to see whether it is TV dialogue, nearby people or the bot's echo.
- VAD and barge-in thresholds, and whether noise suppression is on and appropriate.
- Whether transcripts of background speech are low-confidence and could be filtered.
Production consideration: Tune VAD sensitivity, require a minimum speech duration or confidence before barge-in, ignore low-confidence fragments that do not fit the dialogue state, and offer DTMF. Do not over-suppress, because quiet callers will then be ignored. Add noisy simulated calls to the suite.
50. A bank wants its existing WhatsApp assistant to also answer phone calls. How would you plan it?
Answer: Reuse the back end, redesign the front end. Tools, policies, knowledge sources and verification services can be shared; conversation design, response formatting, latency engineering, telephony and evaluation are new work.
What I would check:
- Which intents work on voice: drop anything that depends on reading long text, links or images.
- How verification changes without links and in-app flows: OTP by DTMF, step-up for sensitive actions.
- Whether tool latency is acceptable for real-time speech, and where pre-fetching is needed.
- ASR and TTS quality for the bank's languages over phone audio.
- Disclosure, recording consent and retention for audio, which chat did not have.
Production consideration: Share one tool layer and one policy engine across channels so a rule changed once applies everywhere, but keep separate prompts and eval suites per channel. The chat side is covered in our WhatsApp AI chatbot project.
Key takeaways
- Voice AI is an LLM agent under a real-time constraint; the hard parts are latency, turn-taking, entity accuracy and safe handoff, not the model.
- Choose between the cascade and speech-to-speech on control, guardrails, languages and latency, and measure both on your own calls.
- Budget time to first audio per stage and track percentiles; the slowest turn is what callers remember.
- WER is necessary but not sufficient: measure entity accuracy and task success per language, including code-mixed speech.
- Keep OTPs, card numbers and other secrets out of the LLM, transcripts and speech; use DTMF and server-side verification.
- Disclose the AI, state recording, honour every request for a human, and transfer with context.
- Evaluate with simulated callers calibrated against real recordings, and judge cost per successful task, not per minute.
Interview preparation checklist
- Draw the voice pipeline from memory, with streaming, endpointing and barge-in marked.
- Explain cascade versus speech-to-speech with three trade-offs and when you would choose each.
- Prepare a latency budget table and explain how you would measure each stage.
- Calculate WER by hand on a short example and explain its limits.
- Be ready to discuss code-mixed Indian language recognition, script conventions and custom vocabulary.
- Know what SIP, PSTN, WebRTC and DTMF are and how call transfer works.
- Describe an OTP verification flow where the model never sees the OTP.
- Explain how you would build simulated callers and which metrics you would report.
- Know the compliance themes: AI disclosure, recording consent, DPDP Act and Article 50 of the EU AI Act.
- Build a small voice agent (browser or phone) with streaming ASR, an LLM tool call and TTS, and bring latency numbers you measured yourself.
FAQ
What skills are required for a voice AI engineer role?
You need solid Python or another backend language, LLM tool calling, streaming and asynchronous programming, basics of audio and telephony, and evaluation skills. Cloud deployment, observability and an understanding of privacy rules for recordings matter as much as knowledge of any specific speech model.
How should I prepare for a voice agent interview?
Build a small voice agent that streams speech recognition, calls one tool and speaks the reply, then measure its latency per stage. Practise explaining the pipeline, barge-in, endpointing and verification aloud, and prepare one story about a bug you found and fixed.
Do I need a speech processing background to work on voice AI?
Not for most application roles. Most teams use managed speech recognition and text-to-speech services or open models, so the work is integration, latency, dialogue design and evaluation. A speech research background helps for roles that train or adapt acoustic models.
Is voice AI a good career path in 2026?
It is a useful specialisation because contact centres, healthcare, banking and IT helpdesks all handle large call volumes and need engineers who can make voice agents reliable. The skills also transfer to general AI agent engineering, so the risk of specialising is low.
What is the difference between speech recognition and conversational AI?
Speech recognition converts audio into text. Conversational AI covers the whole interaction: understanding intent, managing dialogue state, calling tools, generating replies and speaking them. A conversational AI interview in 2026 usually expects you to know both and how they connect.
Which speech to text interview topics come up most often?
Streaming versus batch recognition, word error rate and its limits, custom vocabulary, accents and noisy phone audio, diarisation, code-mixed languages and how recognition errors on names and numbers affect downstream tasks.
Should I learn telephony protocols in detail?
Learn the concepts rather than the specifications: what SIP, PSTN, WebRTC and DTMF do, why phone audio is narrowband, and how transfers and hang-ups reach your application. Most teams integrate through a telephony or contact-centre provider.
How do Indian languages change voice AI work?
Callers often mix Telugu, Hindi, Tamil or other languages with English in one sentence, and names and localities are rare in general training data. You need to test recognition and voices on real local audio, agree on script conventions and confirm critical details carefully.
Can I build a voice AI portfolio project without telephony costs?
Yes. Start with a browser-based agent over WebRTC or a local microphone, using streaming speech services or open models. Add a phone number later through a provider's trial if needed. Document latency, evaluation results and how you handled sensitive data.
Voice agents sit at the point where AI meets real customers, so the engineering has to hold up on a noisy phone line, not just in a demo. To build that depth across LLMs, agents, cloud and security, explore the APEX program at Cloudsoft. If you want to deliver systems like this inside customer organisations, the AI Forward Deployed Engineer course (FDE PRO) covers integration, deployment, observability and evaluation, with placement support until you're placed. Both run in Ameerpet beside the metro or live online; call +91 96660 19191 for a free demo.



