Can facial expression analysis detect customer vulnerability in financial services?

Not on its own, and any vendor claiming otherwise is overselling it. The Facial Action Coding System measures which facial muscles moved, catalogued as Action Units, and a given Action Unit has no fixed emotional meaning: a brow raise accompanies surprise, concentration and scepticism alike. Read in isolation, facial data produces false positives at a rate no FCA-regulated firm should accept, and inferring a vulnerability classification from appearance alone would be indefensible under both Consumer Duty and Article 22 UK GDPR. What facial measurement does reliably is act as a second and third opinion against the words. Vulnerability signals are visible in the divergence: a customer states they are managing, while voice prosody shows hesitation and falling pitch and the face shows brow raise, lip press and gaze drop. The face is not the detector. The disagreement between channels is the detector, and the face is one vote in it.

The objection, stated properly

The academic critique of emotion recognition is substantially correct, and it is worth stating in full rather than working around it.

Facial expressions do not map cleanly onto emotions. The assumption that a given configuration of the face reliably indicates a specific internal state does not survive contact with the evidence. People smile when uncomfortable. People show very little when distressed. Context changes meaning more than the expression does.

Models carry demographic bias. Systems trained predominantly on one population score others less accurately. This is a documented and repeated finding, and a vendor that does not raise it unprompted should be asked why.

Vulnerability is not an emotion. It is a circumstance — health, life events, resilience, capability, in the FCA’s own framing. No facial measurement observes a circumstance. At best it observes a response to one.

All three objections hold. None of them is an argument against measuring the face. They are arguments against measuring only the face, and against converting a measurement directly into a decision.

What changes when three channels are scored together

A single channel asks “what is this person feeling?”, which is the question the evidence says cannot be answered reliably from appearance. Three channels ask a different and far more tractable question: do these signals agree with each other?

Channel What it measures Alone, it fails when
Face
44 Action Units
Which facial muscles moved, and when, independent of interpretation. The customer is composed, or the movement has several possible meanings.
Voice
prosody
Pitch contour, pause length, volume decay, vocal strain. The line is poor, the accent is unfamiliar, or the speaker is naturally flat.
Text
linguistic markers
Hedging, withdrawal, reduced commitment, loss of clarity. The customer says the reassuring thing, which is the common case.
All three, fused The distance between what was said and what the other channels indicate. Every channel agrees — in which case there is nothing to flag.

Each channel’s weakness is a different weakness. A composed face is caught by the voice. A flat voice is caught by the language. Reassuring language — the single most common presentation of unacknowledged financial difficulty — is caught by the other two. This is why the measurement is built on disagreement rather than on any one reading.

What it found that people did not

In a twelve-week deployment across a UK lender’s video affordability interviews, approximately 60% of the vulnerability flags raised were cases the interviewing staff had not identified. Those customers presented as composed, answered appropriately, and gave no verbal indication of difficulty.

That figure is the argument for multimodal measurement and, read honestly, also its limit: a signal that disagrees with a trained human three-fifths of the time requires a human to adjudicate it, not to defer to it. Every flag routes to qualified review. None of them decides anything.

Figures are from an anonymised deployment and have been adjusted to protect client confidentiality. EchoDepth’s models are trained across six countries and fourteen cultural cohorts. Outputs are advisory signals for qualified human review under Article 22 UK GDPR.

Read the methodology   How the score is built →

Article 22 UK GDPR — Human Review Required: EchoDepth vulnerability and distress signals are advisory outputs for qualified human review — not automated decisions. All vulnerability classifications, product suitability assessments and escalation decisions must be made by a qualified professional. Privacy policy · Terms