In 2024, an engineering firm named Arup lost $25 million in one virtual meeting. A finance worker for the company was tricked into a meeting with AI deepfakes of the chief financial officer and colleagues. 

While the video standpoint of the incident was infamous, the aspect that lost the $25 million was the audio. Today the advancement of generative AI has made voice cloning a dangerous threat that bypasses normal security measures.

Here is why traditional defenses are failing and how a modern voice cloning attack works.

The Transition from Robotic Voices

We used to easily be able to recognize an AI voice: weird pauses, robotic tone, and irregular rhythm. Modern voice cloning isn’t the old robotic text-to-speech anymore. 

Driven by deep learning models that can capture a speaker’s almost exact tone, pitch, rhythm, accent, and pronunciation. These models also only require a few seconds of someone’s voice—which could be scraped from the internet—to clone it. 

AI-generated voice file by ElevenLabs.

Real-Time, Low-Latency Attacks

People usually think that these voice cloning attacks are pre-recorded. If an attacker had to type up a response and generate audio, a live, undetectable conversation would be impossible.

Instead, now attackers use speech-to-speech pipelines for low delay.

Real-time attack pipeline architecture representation
Speech-to-Speech Pipeline Architecture Representation

    Real-time voice cloning attack process:

  1. The attacker speaks into the software they are using in real time.
  2. The voice cloning software morphs the audio by replacing the scammer’s vocal qualities with the victim's in milliseconds.
  3. The cloned audio plays directly on whatever platform the call is happening on.

This allows attackers to have a fluid, seamless conversation live and not raise any suspicion. 

Personalized Context and Emotional Manipulation

Technical precision is only half the attack; scammers often try to psychologically manipulate their targets. They don’t try to guess who your chief financial officer is; attackers use tools to scrape the internet to create personalized context around the attack.

By pairing context with emotional manipulation, for example, pretending to be a family member in a crisis or an executive demanding an emergency transfer, attackers demolish logical thinking.

Conclusion

If you or your company's current solution to detect deepfake voices is just listening to irregularities or calling back, you are exposed. Human ears can’t diagnose a modern voice clone reliably. Security teams must start treating AI voice fraud more seriously and consider having a detection tool.