All Blogs

    Mis-Selling Controls for Voice Sales Calls in India 2026: What RBI and IRDAI Expect, and How to Audit 100% of Calls

    17 Mins ReadSep 2, 2026
    Mis-Selling Controls for Voice Sales Calls in India 2026: What RBI and IRDAI Expect, and How to Audit 100% of Calls

    The compliance head at a mid-sized NBFC got the complaint on a Thursday. A borrower in Nagpur said he had been told the insurance bundled with his personal loan was mandatory. It was not. He wanted it unwound and he had escalated past the branch.

    She pulled the call. It took two days, because the recording was in a system her team could search by phone number and date but not by content. When she finally heard it, the executive had not said "mandatory". He had said something worse and harder to defend: "sir, without this the file will not move." Technically not a false statement about the product. Functionally a statement that the loan depended on buying the insurance.

    Then came the question that actually mattered, from her CEO: how many other calls said something similar?

    She had no answer. Her team sampled roughly 2% of sales calls for quality, chosen largely at random, scored against a rubric that asked whether the executive was polite and whether he did the closing. Nothing in that rubric would have caught this. The honest answer was that she did not know, could not know, and would not know until the next complaint arrived.

    Mis-selling in Indian financial services is not primarily a training problem or a bad-apple problem. It is a detection problem. The conduct happens on voice calls, the evidence sits in recordings nobody listens to, and the sampling rate is so low that the practice has to be near-universal before random sampling finds it. This post covers what the regulators actually require on sales conduct, the specific call patterns that create liability, why sampling fails structurally, and how automated auditing of 100% of calls changes the economics of the problem.

    A note on scope: regulation here moves, and some of it is in consultation. Everything below describes settled requirements or well-established regulatory direction. Verify current text against the relevant regulator before you build controls on it.

    Why this got urgent

    Three forces converged.

    Regulatory attention shifted from disclosure to outcomes. For years the compliance answer to mis-selling was a longer disclosure document. Regulators across RBI, IRDAI and SEBI have moved steadily toward asking whether the customer actually understood and whether the product was suitable, which is a question a signed form cannot answer and a call recording can.

    Bundled products became the flashpoint. Credit-linked insurance, add-ons at loan origination, and cross-sell at renewal are where the complaints cluster. The pattern is almost always the same: a product that is genuinely optional is presented in a way that makes it sound conditional.

    Voice AI raised the stakes both ways. When an AI agent makes the sales call, every word is deterministic, logged and reproducible, which is a compliance gift. It also means a badly configured agent mis-sells at perfect consistency across a hundred thousand calls, which is a compliance catastrophe with a clean audit trail pointing at you. The same technology that audits the problem can industrialise it.

    What counts as mis-selling on a call

    Mis-selling is not a single offence. Across the Indian regimes it decomposes into five behaviours, and the distinction matters because the controls differ.

    TypeWhat it sounds like on a callPrimary regime
    MisrepresentationStating a return, benefit or feature the product does not haveRBI, IRDAI, SEBI
    ConditionalityImplying an optional product is required to get the main oneRBI Fair Practices
    Suitability failureSelling a product plainly unsuited to the customer's profile or needIRDAI, SEBI
    Non-disclosureOmitting charges, lock-in, exclusions, surrender terms or the free-look rightIRDAI, RBI, SEBI
    Pressure and inducementManufactured urgency, or benefits contingent on immediate actionAll, plus consumer protection

    The one that generates the most complaints in Indian lending is conditionality, and it is the hardest to catch because it is rarely stated outright. Nobody says "insurance is compulsory". They say "the file will not move", "approval is subject to the full package", "this is how the system processes it". Each of those is a conditionality claim in ordinary language, and a keyword search for "mandatory" or "compulsory" finds none of them.

    In insurance, the recurring pattern is guaranteed-return language applied to market-linked products, and non-disclosure of the free-look period, surrender charges, and premium payment term. A customer told a ULIP is "like an FD but better" has been mis-sold regardless of what the signed illustration says.

    In mutual fund distribution, the pattern is past performance framed as expectation, and risk disclosure delivered as a rushed disclaimer at the end of a call the customer has stopped listening to.

    What the regulators actually expect

    Four regimes touch a financial sales call in India. They overlap and compliance teams frequently conflate them.

    RBI, through the Fair Practices Code and its broader conduct expectations, governs how lenders and their agents deal with borrowers. The core obligations relevant to a sales call are transparency about terms and charges, no coercion, and no representation that an optional product is a condition of credit. Grievance redressal must be genuinely accessible and the internal ombudsman framework applies to covered entities. Our detailed treatment of the collections side of this sits in the RBI Fair Practices Code guide for AI collection calls; sales conduct is the mirror of it.

    IRDAI governs insurance solicitation and the protection of policyholders' interests. The operative expectations on a call are that the person soliciting is properly authorised, that the product is disclosed accurately including charges and exclusions, that the free-look right is communicated, and that there is a defensible basis for suitability. Recorded solicitation calls are a legitimate part of the evidentiary record. Where a bundled or cross-sold policy is involved, the optionality of that policy is the single most important thing the call must establish.

    SEBI governs investment product distribution, with risk disclosure, prohibition on assured-return representations, and suitability at its core.

    DPDP 2023 sits underneath all of it and governs the recordings themselves: purpose-bound consent, disclosure that the call is recorded, defined retention and the ability to honour deletion. Building a mis-selling audit system creates a large corpus of sensitive customer conversations, and that corpus is itself a regulated asset. We have covered the recording and consent mechanics in the DPDP versus TRAI consent audit trail playbook.

    The practical synthesis: your controls need to demonstrate, per call, that required disclosures were made, prohibited claims were not made, optionality was stated where relevant, and there was a suitability basis. Four questions. The problem is answering them across every call rather than 2% of them.

    Why 2% sampling structurally cannot work

    This is worth doing the arithmetic on, because most compliance functions have never done it.

    Assume a sales floor making 40,000 calls a month and a QA team sampling 2%, so 800 calls. Suppose 3% of calls contain a conditionality violation, which is a rate a compliance head would consider alarming.

    Random sampling finds roughly 24 of them. That sounds like detection working. It is not, for three reasons.

    It cannot localise. Twenty-four violations spread across 300 executives tells you almost nothing about which executives, which product, which script variant or which branch is generating them. You have a number, not a lead.

    It cannot detect a rate below its own noise floor. A violation occurring on 0.4% of calls, which on 40,000 calls a month is 160 genuine incidents, produces about 3 hits in the sample. Three is indistinguishable from zero. Yet 160 incidents a month is a systemic issue, and it is precisely the size of problem that generates a regulatory inspection finding.

    The sample is not random in practice. QA teams sample what is easy to sample: recent calls, complete calls, calls from executives already under review. Sales conduct violations concentrate at month-end under target pressure and among executives who are performing well on volume, which is exactly the population least likely to be pulled.

    Sampling was a reasonable answer when listening to a call required a human hour. It is no longer the only available answer, and continuing to rely on it is increasingly hard to defend when a regulator asks how you monitor conduct.

    Auditing 100% of calls: how it actually works

    The mechanism is transcription plus structured evaluation, and it is now cheap enough to run on every call rather than a sample. The pipeline has four stages.

    Transcribe with diarisation. Every call, speaker-separated, so you can distinguish what the executive said from what the customer said. Indian-language accuracy matters enormously here and it is where systems quietly fail. A sales floor in Coimbatore runs Tamil and English in the same sentence. Word error rate on the demo audio is not the word error rate on your floor, and the errors concentrate on numbers and product names, which are exactly the tokens a mis-selling check depends on.

    Detect the required elements. Did the executive state the interest rate and the processing fee. Did they say the insurance is optional. Did they mention the free-look period. Did they disclose the lock-in. These are presence checks and they are the easiest and most reliable part of the system.

    Detect the prohibited patterns. Harder, because the violation is semantic rather than lexical. "The file will not move without this" has to be recognised as a conditionality claim without the word "mandatory" appearing. This requires meaning-level evaluation rather than keyword matching, and it is where an LLM-based evaluator earns its cost.

    Score, route and sample the machine. Every call gets a structured score. Calls flagged above a threshold route to a human reviewer. And critically, a random sample of calls the machine passed also route to a human, because otherwise you have no measure of the machine's false negative rate and you have simply moved your blind spot.

    That last point is the one most implementations skip and it is non-negotiable. An automated audit system that is never itself audited is a worse control than honest sampling, because it produces confident coverage statistics that are unverified.

    What good detection looks like

    CheckTypeDetection reliability
    Interest rate and fees statedPresenceHigh
    Optionality of bundled product statedPresenceHigh
    Free-look period mentionedPresenceHigh
    Assured or guaranteed return claimedProhibited, lexical and semanticHigh
    Conditionality implied without stating itProhibited, semanticModerate, improving
    Manufactured urgencyProhibited, semanticModerate
    Suitability basis establishedReasoningLow to moderate, human review needed
    Customer confusion or non-comprehensionBehavioural signalModerate

    Be honest with your board about that right-hand column. Presence checks are close to solved. Semantic prohibition detection is good and getting better. Suitability judgement is not automatable today and should be routed to humans with the machine used to prioritise which calls they see. A vendor claiming automated suitability assessment is overselling.

    When the AI is the one making the sales call

    Everything above assumes human executives. When the outbound sales agent is itself an AI, the control problem inverts, and it becomes substantially easier.

    Claims become a controlled vocabulary. The agent can only say what it has been configured to say. Build the permitted claim set from approved product literature, and prohibited phrasings simply cannot be uttered. This is a genuinely stronger control than any amount of human training.

    Disclosures become deterministic. The free-look mention, the optionality statement, the rate and fee disclosure happen on every single call, in the same words, in a defined position in the conversation. Coverage is 100% by construction rather than by audit.

    The evidence is complete. Transcript, audio, configuration version and timestamp for every call. When a complaint arrives you can reconstruct exactly what was said and prove what the agent was permitted to say on that date.

    The risks are real and different. A configuration error propagates to every call instantly, so change control on the agent's script matters more than any individual executive's training ever did. Version every prompt and claim set, require compliance sign-off on changes, and keep the ability to reconstruct which version ran on which date. And the agent must handle the customer who says "so is this compulsory or not" with a clear, unambiguous answer rather than a deflection, which is a scenario worth testing adversarially before go-live.

    There is also a boundary question. An AI agent should not conduct suitability assessment for complex products. It can collect the inputs and route to a qualified human. It should not decide that a ULIP suits a particular customer.

    A 60-day implementation

    Days 1 to 10: define the standard. Write down, per product, the required disclosures and the prohibited claims. In actual sentences, not policy abstractions. Most compliance functions have this scattered across a product note, a training deck and a QA rubric that do not agree with each other. Reconciling them is the real work and it is worth doing regardless of whether you automate anything.

    Days 11 to 20: baseline honestly. Take 500 recent calls across products, executives and branches, and have humans score them against the new standard. This is your true violation rate. Expect it to be higher than your QA dashboard says. That gap is the point of the exercise.

    Days 21 to 35: build and calibrate detection. Run automated evaluation over the same 500 calls and compare to the human scores. Measure false positives and false negatives per check type. Tune. Do not deploy anything whose false negative rate you have not measured on your own audio.

    Days 36 to 50: run in shadow. Score 100% of calls, route flags to human review, but do not yet attach consequences for executives. Shadow mode surfaces the systemic patterns, which is where the value is, and it lets you fix detection errors before anyone is disciplined on a false positive.

    Days 51 to 60: operationalise. Attach the output to coaching, script revision and product design. Keep a standing random sample of machine-passed calls under human review, permanently. Report coverage and violation rate to the board monthly.

    One organisational warning. If the output of this system is used purely punitively, the sales floor will adapt to the detector rather than to the standard, and you will get compliant-sounding calls that mis-sell in new phrasings. The best-performing implementations route findings into script and product changes first, coaching second, and discipline last.

    What changes over the next 12 months

    Regulatory expectation on monitoring coverage will tighten. Once auditing every call is demonstrably affordable, "we sample 2%" becomes progressively harder to defend as adequate supervision. Firms that build this now are ahead of an expectation rather than reacting to a finding.

    Bundled product sales will get more scrutiny, particularly credit-linked insurance at origination. This is where complaint volumes concentrate and where the conditionality problem lives.

    Expect the evidentiary bar to move from documents to conversations. A signed consent form alongside a recording showing the customer was told the product was required is not a defence, it is an exhibit. Firms whose compliance rests on signed paperwork should assume the recording is what will be examined.

    Bottom line

    Mis-selling in Indian financial services is a detection failure before it is a conduct failure. The behaviour happens on voice calls, concentrates in bundled and cross-sold products, and expresses itself in ordinary language that keyword search cannot find. A 2% sample cannot localise a problem, cannot detect rates below its own noise floor, and is not random in practice. Automated evaluation of 100% of calls changes that: presence checks for required disclosures are reliable today, semantic detection of prohibited claims is good and improving, and suitability judgement still needs humans, whom the system should prioritise rather than replace. Where the sales agent is itself AI, claims become a controlled vocabulary and disclosure coverage becomes deterministic, which is a stronger control than training, provided you version the configuration and treat script changes as controlled changes.

    Define the standard in real sentences, baseline honestly against your own audio, measure your false negative rate before you trust anything, and keep auditing the auditor. If you want to see 100% call scoring run against your own recordings, including the ones your current QA rubric passes, talk to us.

    What to report to the board

    A mis-selling monitoring programme that reports "we reviewed 4,200 calls" is reporting activity, not risk. Five metrics actually inform a board.

    Coverage. Percentage of sales calls transcribed and scored, not percentage sampled. If this is not close to 100%, say so plainly and explain why.

    Violation rate by type. Broken out into the five categories, because they carry different regulatory consequences and different remedies. A rising conditionality rate is a product design problem; a rising non-disclosure rate is usually a script problem.

    Concentration. What share of violations comes from what share of executives, branches and products. Systemic issues and individual issues need entirely different responses, and only concentration data distinguishes them.

    Machine reliability. The false negative rate measured on the standing human sample of machine-passed calls. A board should never be shown a coverage figure without the accuracy figure beside it.

    Time to remediation. Days from detection to script change, coaching or product fix. This is the number that demonstrates supervision actually functions.

    Report the same five every month so trend is visible. Resist the temptation to change the rubric mid-year, because it destroys comparability exactly when you most need it.

    Where this sits in the wider compliance stack

    Mis-selling controls are one layer of a larger obligation set. Consent and recording disclosure sit underneath, DLT governs how you reach the customer at all, and sector rules govern what you may say once connected. Treating them separately is how gaps appear between them. Our BFSI industry page covers the sector view, and the voice AI India regulatory map shows which regulator applies to which use case, which is the question compliance teams get wrong most often when a workflow spans lending and insurance in the same call.

    The three objections you will hear internally

    "Our executives are trained, this is a solved problem." Training establishes what should happen. Monitoring establishes what does. Every organisation that has moved from 2% sampling to full coverage has found a violation rate higher than its QA dashboard reported, and the gap is not because the executives are dishonest. It is because incentives at month end are real and training decays.

    "This will destroy sales morale." It does if the output is used punitively first. It does not if findings route into script fixes and product design before coaching, and coaching before discipline. Executives generally welcome a system that can prove they said the right thing when a customer complains, and that defensive value is worth selling internally.

    "The false positives will bury us." They will, if you deploy without calibrating on your own audio and without a shadow period. That is precisely why the sixty-day plan spends fifteen days measuring detection accuracy before anyone sees a flag with their name on it.

    Frequently Asked Questions

    Kanan Richhariya

    Kanan Richhariya

    Other Blogs

    178.png
    Voice Automation Strategies

    Customer Not Available — A Business Continuity Plan for Last-Mile, Collections and Healthcare Operations in India 2026

    Kanan Richhariya

    Publish: Jun 16, 2026

    177.png
    Voice Automation Strategies

    Abandoned Cart Recovery via Phone Calls for Healthcare in India 2026: Diagnostics, Online Pharmacy and Tele-Medicine

    Kanan Richhariya

    Publish: Jul 10, 2026

    176.png
    Industry Solutions

    Voice AI for Education and Edtech in India 2026: Counselling, Fee Reminders, Attendance and Parent Calls

    Kanan Richhariya

    Publish: Jun 16, 2026

    175.png
    Voice AI & Voice Technology

    Best Vendors for AI Payment Reminder Calls in India 2026: A Borrower-Engagement Buyer's Guide

    Kanan Richhariya

    Publish: Jul 10, 2026

    173.png
    Voice AI & Voice Technology

    Voice AI for Indian Dental Chains 2026: Clove, Sabka Dentist, Apollo White Dental Playbook for Appointment Booking, RCT Follow-Up, Implant Recall and Aligner Programmes

    Kanan Richhariya

    Publish: Jul 10, 2026

    172.png
    Voice AI & Voice Technology

    Voice AI for Online Pharmacy and Diagnostic Labs in India 2026: 1mg, PharmEasy, Apollo Pharmacy, Dr Lal PathLabs Playbook for Order Verification, Sample Collection & Preventive Health Outreach

    Kanan Richhariya

    Publish: Jul 10, 2026

    171.png
    Voice AI & Voice Technology

    Voice AI for Chartered Accountants and CA Firms in India 2026: ITR Season, GST Filing, Audit Coordination, Client Document Collection Playbook

    Kanan Richhariya

    Publish: Jul 10, 2026

    170.png
    Voice AI & Voice Technology

    Voice AI for Indian Matrimony Platforms 2026: Bharatmatrimony, Shaadi, Jeevansathi Playbook for Profile Activation, Upsell, and Subscription Renewal

    Kanan Richhariya

    Publish: Jul 10, 2026

    169.png
    Voice AI & Voice Technology

    Voice AI for Gold Loan NBFCs in India 2026: Muthoot, Manappuram, IIFL Playbook for KYC, Auction Notice, Top-Up Upsell & Branch Operations

    Kanan Richhariya

    Publish: Jun 10, 2026

    168.png
    Voice AI & Voice Technology

    Voice AI for Jewellery Retail in India 2026: High-AOV Appointment Booking, Festive Campaigns & Tier-2 Store Launch Playbook

    Kanan Richhariya

    Publish: Jul 10, 2026

    Caller Digital

    © 2025 Caller Digital | All Rights Reserved