Methodology in context
Spotting vulnerable customers on calls
voice-only versus multimodal.
Speech analytics is how most UK lenders move beyond QA sampling, and it does real work. Here is what it catches, where it stops, and what adding more channels changes.
How can a UK lender automatically identify vulnerable customers during phone calls?
Most UK lenders do it with speech analytics, and for disclosed vulnerability it works. Calls are transcribed and language models flag phrases linked to the FCA's four drivers of health, life events, resilience and capability: a bereavement, a diagnosis, a lost job, a missed payment. Because it runs on every call instead of a QA sample, it turns Consumer Duty monitoring from a spot check into full coverage. Its limit is the customer who never says it. A transcript records words, not how they were said, so a customer who answers “I’m fine” in a flat voice after a long pause reads as fine. Adding voice prosody, meaning pitch, pause length and speech rate, catches some of that, and on video calls the face catches more. The strongest signal is disagreement between channels: reassuring words, strained delivery. Every flag should go to a trained colleague for review; none should decide an outcome by itself.
What speech analytics does well
It is worth being plain about this: for most lenders, deploying speech analytics is the single biggest step up from manual QA, and it does not need anything multimodal to be worth doing.
It covers every call. Manual QA typically reviews a small sample, so most calls are never assessed for vulnerability at all. Automated transcription and classification remove that ceiling, which is the precondition for any population-level outcome MI. The move from sampling to full coverage matters more than the choice of detection method.
It catches what is said. When a customer mentions a diagnosis, a bereavement, a redundancy, a carer’s role or difficulty paying, a well-tuned language model flags it reliably and can tag it to the FG21/1 driver it relates to. That is a large share of the vulnerability a lender needs to act on.
It checks the agent as well as the customer. Whether the disclosure was acknowledged, whether support options were offered, whether the right script was followed: these are agent-side outcomes, and transcripts evidence them well.
Speech analytics and conversation intelligence vendors active in UK financial services include Aveni, Voyc, CallMiner, Recordsure and Verint. Their approaches differ, and the practical test is the same for all of them, EchoDepth included: run the tool on your own recorded calls and compare what it flags with what your specialist team finds.
Where voice-only detection stops
Undisclosed vulnerability. Customers in difficulty often do not say so. Embarrassment, fear of consequences and simply not seeing their situation as relevant all keep it out of the transcript. Word-based detection cannot flag what was never said.
The reassuring answer. “I’m managing” transcribes identically whether it was said easily or after a long pause in a flat voice. Keyword and sentiment models read the words; the words say the customer is fine. This is the same failure that limits sentiment analysis.
Acoustic signals are noisy on their own. Pitch, pace and pauses carry information, but a poor line, an unfamiliar accent or a naturally flat speaker all move them too. Read in isolation, prosody produces the false positives the face produces when read alone.
Not every conversation is a phone call. Video affordability interviews, advice meetings and remote income-and-expenditure reviews carry visual information that an audio pipeline discards.
| Approach | Catches | Misses |
|---|---|---|
| Manual QA sample | Nuance, on the calls a reviewer hears. | Every call outside the sample. |
| Transcript + language model | Disclosed vulnerability and agent handling, on every call. | What the customer does not say. |
| Adding voice prosody | Hesitation, strain and flattening behind neutral words. | Line quality and speaker differences, when read alone. |
| Words, voice and face, fused | The gap between what was said and how it was said. | Cases where every channel agrees, which leave nothing to flag. |
What multimodal adds, and what it does not
Scoring several channels together changes the question. A single channel asks whether this customer sounds vulnerable, which no one channel answers reliably. Several channels ask whether the signals agree with each other. A customer who says the reassuring thing while the voice and, on video, the face say something different is the case word-based tools are structurally unable to see.
It does not replace speech analytics’ strengths. Disclosed vulnerability is still best caught from the words. It does not make a decision either: a divergence flag is a reason for a trained colleague to look again, recorded with the evidence behind it, and under Article 22 UK GDPR it should never be the decision itself.
For a lender, the practical sequence is: get to full coverage first, make sure disclosed vulnerability is caught and acted on, then add the channels that catch what is not disclosed. The Communication Risk Index is how EchoDepth reports that divergence per interaction and across a function.