← Back to blog
Ai scamsVoice cloningScam typesAustralia

Three seconds of audio. That's all a scammer now needs to clone your daughter's voice.

AI voice-cloning cost Australians an estimated $25.8 million in the first half of 2025 alone. Here's how the technology actually works, why the old scam tells no longer apply, and the small number of things that still reliably beat it.

By Travis, Founder, FamilySentry··8 min read

The most disturbing thing about the current generation of voice-cloning scams isn't that the technology exists. It's how little audio the technology needs. Recent AI-cloning research suggests as little as three seconds of reference audio is enough to generate a synthetic voice good enough to fool most people over a phone line. Three seconds. That's less than the time it takes to answer a voicemail greeting. Less than a single social-media video caption. Less than the "hello, this is Sarah" that plays before you leave a message.

That's the scale of what your parent — or you — is now up against on an unknown-number call. And the Australian numbers show it's not theoretical. AI expert Dominique Carlon of Swinburne University estimates that Australians lost $25.8 million to AI voice-cloning scams in the first half of 2025 alone. ASIC reports that in the 2025–26 financial year, Australians lost $7.4 million to deepfake investment scams using AI-cloned voices of Prime Minister Albanese and nine other public figures.

This post walks through how the technology actually works, the three main scam patterns using it in Australia right now, why the old detection tricks don't work any more, and the small number of things that still reliably beat it.

How the technology actually works

Voice cloning is a subset of what researchers call generative AI — the same broad family of systems that produces synthetic images and text. A voice-cloning model is trained on a very large dataset of real human speech to learn how vocal characteristics — pitch, cadence, timbre, breath patterns, regional accent — combine to produce a recognisable voice.

Once the model is trained, it can take a small "reference sample" of a specific person's voice and generate arbitrary new speech in that voice. Early versions of the technology needed hours of clean recordings. Current versions need seconds. The scammer types in what they want the voice to say; the model produces audio; the scammer plays it down the phone line — either in real time (using a voice-conversion filter) or pre-recorded.

Sources of reference audio for any specific person are trivially easy to obtain: a TikTok video, an Instagram Reel, a wedding-speech clip on Facebook, a work presentation on LinkedIn, a voicemail greeting, a voice message left on a family WhatsApp group. Public figures have it worst — press conferences, interviews and podcasts provide unlimited training material — but the ASIC deepfake case involving Prime Minister Albanese used exactly the kind of publicly available footage that also exists, in much smaller quantities, for almost every social-media-using Australian.

The three patterns to know

1. The family-emergency call. The evolution of the "Hi Mum" scam we covered in our first post. Instead of a text message, the parent gets a phone call that sounds exactly like their child, grandchild, or sibling — panicked, in an accident, in police custody overseas, needing money urgently and quietly. In a US case in July 2025, a Florida mother sent $15,000 in cash to a courier after receiving a call from someone she believed was her daughter, in tears, asking for help. The daughter had not made the call.

2. The executive / authority call. Aimed more at working-age Australians than at retirees, but worth knowing about. A staff member gets a call from someone who sounds exactly like their CEO, their manager, or a trusted authority figure at their company, directing an urgent payment or a data transfer. A UK engineering firm lost the equivalent of more than $25 million to a single incident after an employee received video-call instructions from what appeared to be senior finance staff.

3. The celebrity / investment endorsement. Rather than calling one person, this pattern uses cloned voices at scale — typically in fake video ads on Facebook, YouTube and TikTok. The most prominent recent Australian example is the "Albanese investment" scam ASIC has been actively removing: authentic video footage of the Prime Minister is preserved, with only the audio track replaced by an AI-cloned voice pitching a fictitious investment scheme returning $40,000. Because the video is real, the standard advice about looking for "unnatural blinking" or "mismatched lip-sync" doesn't apply — those tells only work for fully synthetic video, not for what researchers now call audio-grafted deepfakes. ASIC reports removing more than 19,400 fraudulent sites in the past year — most tied to these campaigns.

For elderly Australians, pattern 1 is the one that matters most. It's targeted, personal, emotional, and specifically designed to bypass the normal "is this a scam" reflex by activating a much older reflex: a parent responding to a child in distress.

Why the old tells don't work any more

For twenty years, the standard scam-detection advice was to listen for the signals of a fake: broken English, robotic phrasing, background noise from an overseas call centre, a script the caller couldn't deviate from. Those signals mattered because they were correlated with the reality of who was on the line.

None of them apply to AI-cloned calls.

The voice sounds like the person it's pretending to be. The phrasing is generated on the fly, not read from a script, so the scammer can respond to what your parent says. There is no accent to give it away because the reference audio has the correct accent. And caller ID is trivially spoofable — the call can appear to come from your daughter's actual mobile number, because caller ID is a display, not a verification. What appears on the screen is what the caller tells the network to display.

The uncomfortable reality is that most of the "how to spot a scam" advice older Australians grew up on was correct for scams of the 2000s and now actively misleads. If your instinct is to detect a scam by listening for wrongness, AI voice cloning is the technology that renders that instinct obsolete.

Research bears this out. CommBank research from January 2026 tested Australians on their ability to identify AI-generated content: 89 per cent said they were confident they could spot it. When actually tested, they identified real versus AI-generated correctly only 42 per cent of the time — worse than random guessing. And the gap between over-65s and younger Australians was only six percentage points. This isn't a generational vulnerability. It's a human one.

What actually works

If detection is unreliable, the defence has to be structural rather than perceptual. In practice, three things work — and they work because they don't depend on your parent being able to tell a real voice from a fake one.

1. A family code word, agreed in advance. Something no scammer could learn from social media. The name of the family's first pet. A shared inside joke. A grandparent's middle name. Anything private. The rule is simple: any real family member calling with an emergency will be able to produce the word on demand. Any call that can't — no matter how convincing the voice — ends. As one recent Australian-focused guide put it, a cloned voice cannot answer a phone you dial, and a cloned voice cannot supply a word only your real family knows.

2. Hang up and call back on the number you already have saved. This is the single most robust defence against every voice-cloning attack. The scammer controls what comes down the line to your parent. The scammer does not control what your parent's phone dials outbound. If Mum thinks she has just spoken to her son, she calls her son back on his number in her contacts, not on whatever new number the call came from. In an emergency, if the son doesn't answer, she calls a second family member — again, on a saved number — to check.

3. Slow the call down. AI voice-cloning scams almost universally rely on urgency: the caller is crying, in police custody, mid-arrest, out of time. Every genuine emergency will survive being asked "can I ring you back in two minutes?" No real family member in real distress will refuse a two-minute callback. A scammer will always refuse, because in two minutes the whole operation falls apart.

The honest limit

I want to be careful here. These three defences — code words, callbacks, slowing down — do not stop every AI voice scam. They stop most of them, most of the time, if practiced. But no defence is perfect against a scammer with the right reference audio, a good pretext, and a target who is tired, frightened, or emotionally hooked before the call even started.

That's the reality of where the technology is now. And it's the reason the strongest protection is layered: personal habits (code words, callbacks) combined with a technical layer that can flag the call, while it's still happening, to family members who aren't on the line themselves.

How FamilySentry helps

FamilySentry sits between unknown callers and your parent's phone. Every unknown call is screened by AI configured to recognise the scam scripts being run in Australia right now — including the family-emergency script that voice-cloning scams almost universally follow, and the pattern signatures that are increasingly used to identify AI-generated speech itself.

If a caller starts walking through a familiar emergency-and-urgent-money pattern, your nominated family members get an SMS or push alert while the call is still happening, with a summary of what's being said. Even if your parent believes the voice, you get a second look from outside the call — from someone who isn't in the emotional grip of hearing their child in distress.

Known contacts — the GP, the real family, friends — ring through normally. Your parent doesn't have to learn anything, change anything, or admit anything.

FamilySentry is now live. See plans and pricing to set up protection for your parent's phone.

Further reading

Found this useful? Share it.

EmailX / TwitterFacebookLinkedIn

Related posts