AI voice cloning attacks: how they work and how to stop them

Cloning a voice now takes seconds of audio and costs nothing. A practical breakdown of how these calls are built and the verification habits that defeat them.

7 min read

Voice has always carried unearned authority. We recognise a colleague in the first syllable and stop questioning identity entirely. That reflex was reasonable for most of human history. It stopped being reasonable about two years ago.

#What the attack actually requires

Less than people assume. The pipeline has four steps and none of them are exotic:

  • Collect audio. A recorded webinar, a podcast interview, a conference panel, a company video, an out-of-office voicemail greeting. Executives are the easiest targets precisely because their voices are public.
  • Clone. Commercial and open models produce a usable voice from a short sample. Quality is good enough for a phone line, where compression hides most artefacts anyway.
  • Research the pretext. Who reports to the person being cloned, what project is live, which supplier is mid-contract, who processes payments and who is on leave this week.
  • Call at the right moment. Late Friday, during a known travel window, or immediately after a genuine announcement that makes the request plausible.

The technology is the least interesting part. The pretext is where the effort goes, and it is built almost entirely from information you published yourself.

#The shape of a real call

It is short. The caller apologises for the rush, references something true and recent, and asks for one specific action that is unusual but not absurd: release a payment, confirm a code, add a new payee, share a document link. There is a plausible reason the normal channel is unavailable - a bad connection, a phone left behind, a confidential deal that cannot go through the usual approvers yet.

Crucially, the caller offers a reason not to verify, and makes verifying feel socially awkward. That is the payload. Everything else is packaging.

#Why detection tools are not the answer

Synthetic speech detectors exist and they help, but they are in a race they cannot durably win, and they are usually not present in the channel where the attack lands: an ordinary phone call to a mobile. Building a defence that assumes a detector will flag the call is building a defence that fails silently the day the models improve.

Do not try to determine whether a voice is real. Determine whether the request is real, through a channel the caller does not control.

#The controls that actually work

All of them are procedural, cheap, and independent of how good the clone is.

  • Callback on a known number. Never a number supplied during the call, never a number in the email signature that prompted it. The directory entry or the number already in your phone.
  • Out-of-band confirmation for anything that moves money or access. A second channel the attacker is unlikely to hold simultaneously: an internal chat message, a ticket, an in-person check.
  • A verification phrase for executives and finance. Simple, rotated, never spoken over email. Low technology, remarkably effective.
  • A hard rule that urgency never removes a step. If a request is genuinely urgent, following procedure costs ninety seconds. If ninety seconds breaks the deal, the deal was the attack.
  • Rehearsal. Run a consented vishing simulation against the roles that would be targeted, then debrief without blame. The goal is that the real call feels familiar.

#Reduce the sample material where you can

You will not remove your CEO's voice from the internet, and trying is usually the wrong fight. But you can decide whether every employee's voicemail greeting uses their full name and role, whether internal recordings are published externally by default, and whether your finance team's direct numbers are on the public site. Each removal raises the cost of research.

#What to do in the first ten minutes

If somebody believes they have taken a cloned-voice call: preserve everything, including call logs, timestamps and the number. Do not call the number back. Notify the person who was impersonated through a separate channel. If money moved, contact the bank immediately, because recall windows are measured in hours. Then check whether the same pretext has been tried elsewhere in the organisation, since these campaigns are rarely a single call.

Want this tested against your own organisation?

Book a free call and we will talk through which of these techniques would actually work against your team.

Book a free review