In 2024, an engineering firm named Arup lost $25 million in one virtual meeting. A finance worker for the company was tricked into a meeting with AI deepfakes of the chief financial officer and colleagues.
While the video standpoint of the incident was infamous, the aspect that lost the $25 million was the audio. Today the advancement of generative AI has made voice cloning a dangerous threat that bypasses normal security measures.
Here is why traditional defenses are failing and how a modern voice cloning attack works.
The Transition from Robotic Voices
We used to easily be able to recognize an AI voice: weird pauses, robotic tone, and irregular rhythm. Modern voice cloning isn’t the old robotic text-to-speech anymore.
Driven by deep learning models that can capture a speaker’s almost exact tone, pitch, rhythm, accent, and pronunciation. These models also only require a few seconds of someone’s voice—which could be scraped from the internet—to clone it.
AI-generated voice file by ElevenLabs.
Real-Time, Low-Latency Attacks
People usually think that these voice cloning attacks are pre-recorded. If an attacker had to type up a response and generate audio, a live, undetectable conversation would be impossible.
Instead, now attackers use speech-to-speech pipelines for low delay.
- The attacker speaks into the software they are using in real time.
- The voice cloning software morphs the audio by replacing the scammer’s vocal qualities with the victim's in milliseconds.
- The cloned audio plays directly on whatever platform the call is happening on.
Real-time voice cloning attack process:
This allows attackers to have a fluid, seamless conversation live and not raise any suspicion.
Personalized Context and Emotional Manipulation
Technical precision is only half the attack; scammers often try to psychologically manipulate their targets. They don’t try to guess who your chief financial officer is; attackers use tools to scrape the internet to create personalized context around the attack.
By pairing context with emotional manipulation, for example, pretending to be a family member in a crisis or an executive demanding an emergency transfer, attackers demolish logical thinking.
Conclusion
If you or your company's current solution to detect deepfake voices is just listening to irregularities or calling back, you are exposed. Human ears can’t diagnose a modern voice clone reliably. Security teams must start treating AI voice fraud more seriously and consider having a detection tool.