How AI Voice Scams Work (And What Actually Stops Them)
A phone rings. The voice on the other end sounds exactly like your CFO, your bank manager, or your son. It is not. According to Group-IB's 2025 threat research, AI-generated voice attacks now target finance teams, executive assistants, and everyday families at scale. The voice is real enough to fool almost anyone. The person behind it is not.
August 4, 2026
Originally reported by Group-IB · Read the original article
- AI voice cloning requires as little as 30 seconds of audio to produce a convincing fake voice.
- Human detection accuracy for synthetic voices sits around 55 to 60 percent, barely better than guessing.
- Group-IB identifies finance teams, executive assistants, and help desks as the highest-risk targets in 2025.
- Shared rotating codewords give callers something to prove that no voice clone can provide.
- Trust Onion is free and works offline, with no server required.
The Attack Is Simpler Than You Think
Deepfake voice phishing, also called vishing, does not require a sophisticated hacker. It requires a voice sample and a goal.
Group-IB's research published in August 2025 breaks down exactly how these attacks unfold. An attacker gathers a short audio clip of the target's voice from social media videos, voicemails, earnings calls, or YouTube interviews. Thirty seconds is often enough. The AI clones the voice, the attacker calls someone who trusts that voice, and makes a request.
The request is almost always urgent: wire this money, give me your login, don't tell anyone yet.
Urgency is the real weapon. The cloned voice just makes it believable.
Who Gets Targeted
Group-IB's incident data points to three sectors absorbing the most damage: financial services, executive support, and remote-work help desks.
The pattern makes sense. These are environments where people regularly receive calls from authority figures they have never met face to face. A finance team member gets a call from the CEO asking for a wire transfer. An IT help desk gets a call from an executive requesting a password reset. A remote employee gets a call from HR asking for account credentials.
Nobody questions the voice because the voice sounds right.
Outside the corporate world, families face the same attack in a different form. A grandparent gets a call from a grandchild in trouble. A parent hears their teenager's voice in distress. The FBI reported that grandparent scams cost Americans over $74 million in 2023 alone, and that was before AI voice cloning became this accessible.
The Anatomy of the Attack
Group-IB breaks the attack into recognizable stages.
Voice collection. The attacker finds audio of the target. Public figures are easy. Private individuals take more effort, but social media has made most people findable.
Clone generation. AI voice synthesis tools, several of which are publicly available, process the sample and produce a cloned voice. The quality has improved sharply. Human detection accuracy for synthetic voices currently sits around 55 to 60 percent, barely better than a coin flip.
Call execution. The attacker calls the victim using the cloned voice, often spoofing the caller ID to show a familiar number. They apply pressure through time limits, emotional stakes, and confidentiality requests.
Harvest or transfer. The victim sends money, shares credentials, or takes an action they would not have taken knowing who was really calling.
The whole sequence can unfold in under five minutes.
Why Technical Defenses Alone Fall Short
Organizations are responding with call authentication systems, AI detection software, and staff training. These are worthwhile investments, but Group-IB's research is clear that technical mitigations cannot carry the full load.
AI voice quality improves faster than detection tools can keep up. A cloned voice that fools a trained security professional today was not possible two years ago. Waiting for a better detector is not a strategy.
The deeper problem is verification. On a phone call, you cannot see the person, check a badge, or verify an email signature. You rely entirely on the voice, and the voice is now fakeable.
The only reliable countermeasure is something the attacker cannot know, something absent from any audio clip, social media post, or public record.
What a Shared Codeword Changes
Families and close-knit groups have an advantage here that corporate environments are still working out.
If you and your family share a set of three rotating codewords, any call claiming to be from someone you love can be verified in seconds. You ask: "What are the words?" If the caller knows them, it's really them. If they hesitate, deflect, or give the wrong answer, the call is over.
AI can clone a voice. It cannot fake knowing three codewords that change every few hours and exist only on your family's phones.
Trust Onion is a free app built on exactly this idea. Three rotating codewords, no server, works offline. The words rotate on a schedule, so even if someone overheard yesterday's words, they are already expired. For deeper verification, Trust Onion lets you send a "Proofy": a selfie with the current three words overlaid and cryptographically signed, so the recipient can confirm it is genuinely you, at that moment, with today's words.
The question "What are the words?" costs nothing and takes three seconds. It ends a vishing attack before it starts.
The Lesson from Group-IB's Research
Group-IB's findings show that these attacks succeed because they exploit trust, not just technology. The attacker borrows the voice of someone you trust and uses it against you.
The countermeasure has to work at that same layer. You do not beat a fake voice by listening more carefully. You beat it by requiring proof that no voice clone can provide.
For security teams, that means building verification protocols that go beyond voice recognition: callback procedures, out-of-band confirmation, challenge questions that rotate regularly.
For families, it is simpler. Three words. Ask for them every time.
A Practical Checklist
For families:
- Set up a shared verification phrase with parents, grandparents, and kids
- Make it a habit: if someone calls claiming to be family and needs something, ask for the words
- Use Trust Onion to rotate the words automatically and share a Proofy when needed
For organizations:
- Establish a verbal verification protocol for any request involving money or credentials
- Brief executives and their assistants specifically, since they are the most targeted
- Train help desk staff to treat voice-only verification as insufficient for sensitive actions
The attack Group-IB documents is not theoretical. It is happening now, in real organizations, to real families. The voice on the phone sounds real because it is designed to. The only question worth asking is the one the attacker cannot answer.
What are the words?
Frequently Asked Questions
How do deepfake voice scams work?
Attackers collect a short audio sample of a target's voice, use AI to clone it, then call someone who trusts that person and make an urgent request for money or sensitive information. The whole attack can unfold in under five minutes.
Can you tell if a voice on the phone is AI-generated?
Not reliably. Research shows human detection accuracy for synthetic voices sits around 55 to 60 percent. The quality of AI voice cloning has improved faster than detection tools can keep up.
What is the best way to protect your family from voice cloning scams?
Use a shared verbal codeword your family all knows. If someone calls claiming to be a family member and cannot provide the correct words, hang up. Apps like Trust Onion rotate the words automatically so expired words are useless.
How much audio does it take to clone someone's voice?
Current AI voice synthesis tools can produce a convincing clone from as little as 20 to 30 seconds of audio. That is short enough to be lifted from a social media video or voicemail.
What is vishing?
Vishing is voice-based phishing. An attacker calls a victim while impersonating a trusted person, such as a bank, executive, or relative, and uses urgency and false identity to extract money or information.
Trust Onion is free and takes two minutes to set up. Give your family three rotating codewords that no AI can fake.
Protect Your Family FreeAugust 1, 2026
Senate Hearing: AI Is Targeting Seniors and Winning
A Senate aging committee heard how AI voice cloning and deepfakes cost seniors $7.7 billion in 2025....
August 1, 2026
A Reporter Cloned Their Voice and Called Their Mum
A Cybernews reporter cloned their own voice with AI and tried to scam their mother. It worked. Here'...
July 31, 2026
AI Voice Cloning Is Making Grandparent Scams Scarier
AI voice cloning lets scammers impersonate grandchildren using audio from social media. Here's how s...


