DEV Community

NeuPortal
NeuPortal

Posted on

Voice Is Not an Authentication Factor Any More. Here Is What the Detection Numbers Actually Say


A fraud in Italy this year is a useful forcing function for anyone building identity or payment flows. In late February about EUR 95 million left a private bank over three days. No system was reported breached. The attack was entirely social, and the interesting part for engineers is the authentication model it exploited.

The attack shape

  1. A text message impersonating the group's chief executive, from an unknown number: confidential acquisition, funds must route through you, tell nobody.
  2. Eleven documents on what appeared to be a law firm's letterhead.
  3. A phone call from a lawyer the target knew personally - voice reproduced with AI - confirming the text.
  4. A separate advance call to the payments manager, so the instructions would arrive expected rather than cold.

Note what is being attacked. Not a credential. Not a channel. The corroboration model in the target's head: two independent signals agreeing. The attacker supplied both, and the second one carried a biometric the target had always implicitly trusted.

Almost every English retelling got the detail backwards and said the CEO's voice was cloned. It was not - he was impersonated in text. The cloned voice belonged to the confirming third party, which is precisely why the scheme worked.

The numbers on human detection

If you are designing a flow that assumes a human can tell, here is the published evidence.

The largest peer-reviewed listening study is titled Warning: Humans cannot reliably detect speech deepfakes (PLOS ONE, 2023). Across 529 listeners: 73% accuracy on synthetic clips, 67.8% on genuine, about 70% overall. Training participants with examples first improved them by under four percentage points. The authors conclude that improving human detection is not a realistic strategy. Caveat worth knowing: the stimuli were generic synthetic speech, not clones of people the listeners knew.

A 2025 study (Scientific Reports) used an off-the-shelf commercial cloning product and 604 listeners:

  • 60.8% correctly flagged a clone as artificial
  • one in five performed at or below chance on the clones
  • asked whether a clone and a real recording were the same person, listeners said yes in a median 83.3% of trials
  • longer clips and scripted speech were MORE likely to be judged genuine

That last line should kill the "stay on the line and listen carefully" UX pattern wherever it exists.

The numbers on machine detection

Better, and still not a control you can lean on alone.

A 2025 Intel Labs system, designed specifically for cross-domain generalisation and which its authors say surpasses the top single system in the ASVspoof 5 challenge, reports:

test set equal error rate
in-domain 0.43%
in-the-wild audio 3.19%
ASVspoof 5 4.48%
hidden subsets 7.82% / 5.06%

Read the out-of-domain column as roughly one error in twenty, on clean research audio, before you add telephone codecs, packet loss and background noise. The authors present this as progress in generalisation, and it is - the point is that even the improved case degrades several-fold the moment the generator is unfamiliar.

Europol's 2022 deepfake report names the structural reason: detectors train on databases of known fakes, so performance against a new generator is unknown by construction, and a generator can be retrained specifically to stop emitting whatever a published detector keys on. It is an adversarial loop with an asymmetry that does not favour you.

And the enrolment cost keeps falling. The "three seconds" figure everyone quotes comes from a 2023 Microsoft paper synthesising personalised speech from "a 3-second enrolled recording of an unseen speaker" - on top of a model trained on 60,000 hours of other people. Treat any public audio of your users as compromised enrolment material.

What to build instead

Stop trying to classify the signal. Change what the signal has to prove.

Out-of-band callback, structurally enforced. The verification channel must be one the attacker does not control, and the destination must come from your own records, never from the inbound request. In a product, that means the callback number is rendered from your directory and is not editable in the flow that triggered it.

Challenge-response over something unresearchable. In July 2024 this same attack was run against Ferrari. It failed when an executive asked the caller the title of a book the CEO had recommended him days earlier. The call ended. He detected nothing - he requested something only the real principal could produce. There is research support for the pattern: a 2024 study found human evaluators at 72.6% alone, rising to 84.5% when combined with machine analysis of how the caller handled live challenges. It is a research prototype, not a product.

Two-party authorisation on state changes that move money, with the second party reached over the independent channel - not forwarded the same thread.

Treat urgency plus secrecy as a hard signal. It is present in every documented script here, and it is trivially machine-detectable in text channels. No legitimate process requires both.

The FBI published the callback and the second sign-off in 2017, when BEC had already cost $5.3 billion worldwide. Nothing about generative audio changes the control set. It only removes the last reason anyone had to skip it.

The thing nobody can tell you

How common is this? There is no answer. The FBI publishes exactly one figure specifically about cloning a family member's voice - "over $5 million in 2025" - and its headline "AI related" tally of 22,364 complaints is a tag meaning the report mentioned AI, not a finding that AI was involved. UK Finance's losses to criminals posing as police or bank staff fell 18% in 2025. Australia's 49-page national scam report has no voice-cloning category at all. Europol attaches no number to it.

The most careful survey found 6% of US corporate finance teams with a confirmed deepfake incident - and 40% who did not know whether they had been targeted.

Design for the 40%. The instrumentation that would tell you is the instrumentation nobody has built yet.

Full version with every source: https://neuportal.ai/blog/ai-voice-cloning-scam-95-million-bank

Top comments (0)