New batches starting this week Β· Limited seats

Project Walkthrough: Building a Call-Centre Quality Assurance AI

A project walkthrough for post-call quality assurance AI at an illustrative BFSI contact centre in Hyderabad, covering diarised Telugu-Hindi-English transcription, redaction before the LLM, evidence-cited scoring, human review and disputes, fairness across accents, evaluation against QA-scored calls and an ROI method.

Call-centre QA AI flow: call recording, transcription and redaction, scoring against the scorecard with cited evidence, QA reviewer decides
Last updated Β· 15 min read Β· 3,197 words

This walkthrough builds a call center QA AI system the way a Forward Deployed Engineer would deliver it to a bank's contact centre: from the business problem to a measured return. A call-centre quality assurance AI is production-ready only when every score it gives points to the exact transcript lines behind it, card numbers and OTPs are redacted before any LLM sees the call, compliance misses go to a human reviewer instead of being decided by the model, and agents can see and dispute what was said about them. This is post-call analytics on recorded calls, not a live voice bot. The scenario is illustrative, and the build plan at the end turns it into a portfolio project.

For the real-time side, where an AI talks to callers, read voice AI agents. For a customer-facing chat agent, see the AI customer support agent project.

Business problem

Illustrative scenario. Consider a BFSI contact centre in Hyderabad serving a bank's credit card, personal loan and savings customers. Calls run in Telugu, Hindi and English, often all three in one call: "Sir, mee card block chesanu, aapka new card seven working days mein aayega." Every call is recorded; QA manually scores a small monthly sample per agent.

The symptoms:

  • Coverage is thin. Most calls are never reviewed, so a missed disclosure on a loan call is found only if that call happens to be sampled, or when a complaint arrives.
  • Scoring is inconsistent. Two reviewers score the same call differently, especially on judgement items such as empathy.
  • Feedback is late and vague. Agents hear "work on empathy" weeks after the call, with no example.

The customer does not want to replace QA reviewers. It wants every call screened and reviewers' time spent where it matters. That is the gap from AI demo to enterprise outcome: summarising one call with an LLM takes an afternoon; a system that compliance, QA, the agents themselves and the privacy office will all accept takes engineering.

Requirements

Discovery covers QA, operations, compliance, the privacy office, security, HR (these are scores about employees) and the platform owner. Collect the scorecard, its guidance, QA-scored calls and the disclosure scripts per product.

Functional

  • Transcribe with speaker separation, handling code-mixed Telugu, Hindi and English.
  • Redact card numbers, OTPs, CVVs, account numbers and identity numbers before any LLM processing.
  • Score each call against the QA scorecard with a cited transcript quote for every item.
  • Flag compliance risks into a human review queue.
  • Draft coaching notes; show trends by product, queue and language.
  • Let agents view their scored calls and raise a dispute.

Non-functional

  • The AI score is a draft. Only a reviewer-confirmed score enters performance records.
  • No unredacted payment or authentication data in transcripts, prompts, logs or traces.
  • Batch results within hours, not seconds.
  • Comparable quality across languages and accents, measured, not assumed.

The scorecard as a contract

ItemTypeWhat counts as evidence
GreetingRule-likeAgent opens with the approved greeting and name
Verification doneComplianceRequired identity checks completed before account details are discussed
Disclosure readComplianceProduct-specific disclosure delivered in substance, in the caller's language
ResolutionJudgementIssue resolved, or a clear next step and timeline given
EmpathyJudgementAcknowledges the customer's situation, no dismissive language
Compliance phrasesComplianceMandatory statements present; prohibited promises absent

Each item gets one of four verdicts: met, not_met, not_applicable or insufficient_evidence. The fourth matters most: on poor audio the system says so instead of failing the agent.

Architecture

Contact-centre platform
  (recording + call metadata)
        |  call-ended event
        v
Ingest service -> object storage (audio)
        |
Speech-to-text + diarisation
        |
PII redaction (text + audio)
        |
Scoring pipeline (LLM, per item)
   |-- scorecard + scripts (RAG)
   |-- evidence validator
        |
Results store (PostgreSQL)
   |            |            |
Review queue  Coaching    Trends
 (QA team)    drafts     dashboard
   |
Agent view + disputes

Key decisions:

  • Batch, queue-driven pipeline. Each stage reads from a queue and writes its output, so a slow speech-to-text batch never blocks scoring, and any stage can be re-run.
  • Redaction sits before the LLM boundary. Nothing downstream of redaction ever receives raw transcript text.
  • Deterministic checks first. Timing rules (was verification before the balance was read out?) are code over timestamps; the LLM handles meaning.

Data

DataSourceHandling
Call audioPlatform recording store or exportEncrypted object storage; short retention for working copies
Call metadataPlatform APIs (agent ID, queue, product, disposition, timestamps)Joined to each call; drives which disclosures apply
TranscriptsGeneratedRedacted before storage; utterance IDs and timestamps kept
Scorecard and scriptsQA and compliance teamsVersioned; every score records the version used
Human-scored callsQA team's historyGold set for evaluation, never shown to the model as answers

Two practical points decide transcript quality. First, ask for dual-channel (stereo) recordings, agent on one channel and customer on the other, if the platform supports it. Speaker separation then becomes reliable almost for free. With mono audio you depend on a diarisation model, and its errors turn "agent asked for OTP" into "customer asked for OTP". Second, decide the transcript script early. Code-mixed Telugu and Hindi may come back in native script, Roman script or a mix, and the approved disclosure scripts must be compared in the same form.

LLM

The model works on redacted text only. Choose it on your own scored calls, judged by native speakers, with these criteria:

  • Code-mixed comprehension: understanding a Telugu disclosure delivered in substance rather than word for word.
  • Structured output: reliable JSON with verdicts and utterance IDs (schema-validated, with a retry on malformed output).
  • Data terms and region: available on the bank's approved platform, such as Amazon Bedrock or Azure OpenAI, in a region compliance accepts.

Speech-to-text is the other model decision, and often the harder one. Test engines on your own narrowband call audio and measure accuracy on the words that drive scores: product names, disclosure phrases, numbers and "yes" or "no" answers. A general word error rate hides exactly these.

RAG

Retrieval is small and precise: per call, fetch the scorecard guidance, the disclosure script for that product and language, and a few QA calibration examples of met and not_met. Filter on metadata first; semantic search is secondary. The domain-specific rule: the model must judge against the script version in force on the call date, not today's.

Agent

Be honest with the customer: scoring is a fixed workflow, not an agent, so the same call gives the same result. Agentic behaviour helps in two narrow places:

  • Coaching notes. From an agent's reviewer-confirmed scores, the model finds a pattern ("verification skipped on calls transferred from the IVR") and drafts one note with cited examples. A team leader edits and sends it.
  • Trend questions. A QA lead asks "which queue has the most disclosure misses on loan calls this month?" and an assistant answers through read-only tools.

Scoring flow per item

redacted transcript + metadata
        |
item applies? --no--> not_applicable
        | yes
low STT confidence? --yes--> insufficient_evidence
        | no
LLM: verdict + utterance IDs + quote
        |
validator: quote exists at that ID?
        |-- no --> retry once, then review
        v
compliance item not_met? --yes--> review queue
        | no
draft score

The evidence validator is the most important small component. It checks each quote appears verbatim in the cited utterance from the right speaker. A score without valid evidence is never shown.

Tools

ToolPurposeGuardrails
get_disclosure_scriptApproved script by product, language and dateRead-only; versioned
get_confirmed_scoresAgent's reviewer-confirmed history for coachingScoped to the requesting team leader's team
create_review_itemSend a flag to the QA queueIdempotent per call and item
draft_coaching_noteSave a draft for a team leaderDraft only; a person sends it
query_trendsAggregates for dashboards and questionsAggregates only; minimum group size so individuals are not exposed

Deliberately missing: anything that writes to HR or performance systems, or that contacts customers.

MCP/API

Contact-centre platforms differ, but the shape is similar: a call-ended webhook or event signals a recording; a recordings API or secure scheduled export provides audio; reporting APIs, sometimes joined with the CRM, provide metadata. Results go back into the platform's quality module through its API, or into a separate QA app.

Expose the results store and trend queries through an MCP server, so a supervisor assistant or later agent-assist tool can reuse them without new integration code. MCP, the Model Context Protocol, is an open protocol introduced by Anthropic in late 2024 for connecting AI applications to tools and data. Keep each operation narrow and read-only, with the caller's identity and team scope enforced on the server side.

Wiring post-call pipelines into real platforms, CRMs and identity systems is the delivery work Cloudsoft's AI Forward Deployed Engineer course practises, notably in its Customer Integration Service and Secure Banking AI Assistant projects.

Security and privacy

Customers hear a recording notice at the start of the call; check with counsel that the notice and lawful basis cover quality monitoring and AI-assisted analysis, not only "recorded for quality and training purposes" read narrowly. Agents are data principals too: their voices and scores are personal data, so their notice must explain AI scoring, human review and disputes. Training models on calls is a separate purpose. The engineering side of consent, purpose tagging and erasure is in our guide to the DPDP Act for AI applications. This is not legal advice; sector regulators may add their own expectations.

PII redaction before the LLM

  • Prevent first. Where the platform supports it, take card details via keypad input or pause recording during payment, so the numbers are never recorded.
  • Redact spoken numbers in every language. Customers say "four four double seven", or digits in Telugu or Hindi. The redactor needs number-word handling per language, plus patterns and context ("OTP cheppandi", "CVV").
  • Use typed placeholders such as [OTP] or [CARD_NUMBER]. The scorer can still see that the agent asked for an OTP and the customer gave one, which is what the verification item needs.
  • Bleep the same spans in stored audio copies.
  • Test redaction recall as a release gate. A missed card number is an incident; an over-redacted word is just a slightly worse transcript.

Retention and access

Set retention separately for raw audio, redacted transcripts, scores and disputes, in line with the bank's records policy. Agents see their own calls only; team leaders their team; QA and compliance all calls. Log every access.

Cloud and cost per call

On AWS: S3, SQS between stages, ECS or EKS, RDS for PostgreSQL, Amazon Bedrock and KMS, in an India region where residency requires it. Azure has direct equivalents. Provision with Terraform.

Finance will ask for cost per call:

cost per call =
    STT rate x audio minutes
  + LLM tokens (in + out) x rates
  + redaction and pipeline compute
  + storage for the retention period
  + reviewer time on flagged calls

Levers: score all items in one structured call rather than one call per item where accuracy holds; cache the scorecard and script portion of the prompt where your provider supports prompt caching; use a smaller model for rule-like items; use batch or off-peak processing, since nobody needs a score within seconds; and strip hold music before speech-to-text.

Observability

Trace every call through every stage: STT confidence, redaction counts, scorecard and model versions, verdicts, validator failures and cost. Use Langfuse or LangSmith for the LLM steps and OpenTelemetry for the pipeline.

Dashboards: backlog, failures and cost per call for operations; overturn rate per item, disputes and flags per queue, split by language, for QA. A spike in "disclosure not read" on one queue may be coaching, or a changed IVR prompt confusing diarisation; traces tell you which.

Evaluation

The ground truth is the QA team's own judgement. Build a gold set of calls scored independently by two reviewers, covering every product, language and a spread of accents, including known compliance misses.

MetricWhat it tells you
Agreement with QA, per itemHow often the AI verdict matches the reviewers' agreed verdict; report a chance-corrected measure such as Cohen's kappa alongside raw agreement
Human-human agreementThe realistic ceiling; judgement items like empathy will be lower for people too
Missed-compliance-flag rateCalls with a real compliance miss that the AI did not flag. The primary safety metric and a release gate
False-flag rateFlags reviewers reject; drives reviewer workload and trust
Evidence validityShare of verdicts whose citations pass the validator and that reviewers find relevant

Tune the flagging threshold towards catching misses: a false flag costs a reviewer a few minutes, while a missed disclosure on a loan call is a regulatory problem. Method details, including LLM-judge limits, are in our LLM evaluation guide.

Fairness across accents and languages

The likeliest bias is not in the LLM's opinions but in speech-to-text. If transcription is weaker for one regional accent or for Telugu-heavy calls, the disclosure item fails more often for those agents, and the AI quietly turns an STT weakness into lower scores for real people. Stratify every metric by call language and agent accent region, compare overturn and dispute rates across groups in production, and treat a gap as a defect to fix (better STT, custom vocabulary, routing to insufficient_evidence) before rollout. The methods are in AI bias and fairness testing.

Deployment

GitHub Actions runs the redaction recall suite and gold-set evaluation on every prompt, scorecard, model or STT change, failing on any regression in missed flags or fairness gaps. Argo CD deploys to EKS.

Roll out in stages:

  1. Shadow mode. The AI scores the calls QA already samples; reviewers score blind, then compare. Nothing reaches agents.
  2. Assisted review. Reviewers confirm, edit or reject AI drafts on their sample.
  3. Full screening. Every call is screened; flags and low-confidence calls go to the queue, and a random sample is still human-scored.

Human review and disputes

Reviewers make the final decision on every score that counts. Agents see the item, verdict, quote and a link to the audio, and can dispute within an agreed window; a different reviewer decides, and every overturned score goes into the evaluation set. Watch for rubber-stamping; patterns are in human-in-the-loop AI. The wider control set banks expect around LLMs is in generative AI in banking.

ROI

ROI is a method agreed with operations, QA and finance against a measured baseline. Every value below is a hypothetical placeholder to show the arithmetic, not a result or benchmark.

InputPlaceholderWhere the real value comes from
Calls per month (V)e.g. 100,000Platform reports
Calls manually reviewed today (M)customer's figureQA team records
Reviewer minutes per manual review (T1)e.g. 20Time study of current process
Reviewer minutes per AI-assisted review (T2)measured in pilotReview queue timestamps
Calls flagged for review per month (Q)measured in pilotReview queue
AI cost per call (A)from billingSTT, LLM, compute and storage
Fixed run cost per month (F)from billingPlatform, support, scorecard upkeep

Reviewer hours freed = (M Γ— T1 βˆ’ Q Γ— T2) Γ· 60, valued at the loaded cost per reviewer hour. Monthly net = that value βˆ’ (V Γ— A) βˆ’ F. A negative result can still be right: coverage went from a sample to every call. Report that gain, and misses the sample would not have caught, separately.

Build it yourself: milestone plan

Use role-played or synthetic calls, never real customer recordings.

MilestoneDeliverable
1. FixturesRole-played calls, mock metadata API, scorecard, disclosure scripts
2. TranscriptionSTT with speaker separation, utterance IDs, confidence
3. RedactionCard number and OTP redaction including spoken number words; recall test suite
4. ScoringStructured per-item verdicts with quotes; evidence validator
5. Review appFastAPI and a simple UI: review queue, confirm or override, agent dispute flow
6. EvaluationHand-scored gold set; agreement, missed-flag and false-flag rates split by language
7. OpsQueues, tracing, cost per call, Terraform, CI gate
8. ValueROI one-pager; demo of a cited missed-disclosure flag

The README should report redaction recall, per-language results and limitations. See projects every AI FDE should build for portfolio fit.

Frequently asked questions

What is call center QA AI?

It is a post-call system that transcribes recorded calls, scores them against the quality scorecard with transcript evidence, flags compliance risks for human reviewers, drafts coaching notes and shows trends. It screens every call, while QA reviewers keep the final decision on every score that counts.

How do you stop card numbers and OTPs reaching the LLM?

Prevent capture where possible with keypad input or paused recording, then redact transcripts before the LLM boundary, including numbers spoken as words in Telugu, Hindi or English. Replace them with typed placeholders so verification can still be scored, bleep stored audio, and test redaction recall on every release.

How do you evaluate an AI QA scorer?

Compare it with calls scored independently by two QA reviewers. Track agreement per scorecard item against the human-human ceiling, the missed-compliance-flag rate as the main safety metric, the false-flag rate, and evidence validity, all split by language and accent.

Can agents challenge an AI-generated score?

They should be able to. Agents see each item with its quote and audio link, raise a dispute within an agreed window, and a different human reviewer decides. Overturned scores feed back into the evaluation set.

Does code-mixed Telugu, Hindi and English speech work?

It can, but test it on your own call audio. Speech-to-text quality varies by language and accent, so measure accuracy on the phrases that drive scores, use dual-channel recordings, and mark low-confidence segments as insufficient evidence instead of failing the agent.

Want to build AI systems that a bank's compliance team will actually sign off? The Cloudsoft FDE PRO program runs 12 weeks with five enterprise projects and the GlobalBank capstone, in our Ameerpet classroom beside Ameerpet Metro or live online, with placement support until you're placed. Call +91 96660 19191 for a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us