Glossary · simply explained

Deepfake & vishing

Vishing (voice phishing) is fraud by phone call; deepfakes lift it to a new level: AI clones voices from a few seconds of audio and produces deceptively real videos. The familiar sound of management on the phone — or their face in the video call — is no longer proof of authenticity.

The attacks combine channels: first the preparatory mail, then the call with the cloned voice that builds pressure and confirms the transfer. Documented cases with losses in the millions show this is not theory.

Why deepfakes defeat classic vigilance

For decades the advice was: when in doubt about a mail, call back. Deepfakes turn that around — precisely the call, the voice message or the video can be forged. Telltale signs such as unnatural intonation or artefacts disappear with every model generation; your own hearing can no longer be trusted.

What holds are procedures instead of sensory impressions: callbacks exclusively via self-chosen, known numbers; approvals via defined second channels instead of on the phone; the four-eyes principle without exception — even if the boss personally seems to press. Exactly that pressure is the strongest warning sign.

Effective countermeasures

  • Never change payments or master data on someone’s say-so — always a second channel and the four-eyes principle.
  • Callback rule: only via known numbers from your own systems, never via numbers provided in the request.
  • Awareness: train urgency plus confidentiality plus authority as the alarm pattern.
  • Be sparing with voice and video material of key people where practical.

Frequently asked questions about Deepfake & vishing

How much audio does a voice clone need?

Frighteningly little — modern methods manage with seconds to a few minutes, as supplied by any podcast episode, webinar or voicemail greeting. For exposed people, their voice should be considered publicly compromised.

Can deepfakes be detected reliably?

Not in everyday life: artefacts disappear with every model generation, detection tools deliver probabilities rather than proof. Whoever ties decisions to a feeling of authenticity loses. Only processes that verify identity via independent channels are robust.

What is the difference between vishing and smishing?

Vishing runs via call or voice message, smishing via SMS or messenger. Both rely on immediacy and pressure; deepfake voices amplify vishing dramatically. The defence is identical: defined return channels instead of trusting the incoming contact.

Does a code word help in the company or family?

As an additional layer, yes: an agreed code word for unusual requests is cheap and effective — provided it never circulates in writing. In the company the formalised variant is better: approval processes that even real superiors cannot bypass on the phone.

Are video conferences affected too?

Yes — documented cases show entirely fake conferences with several deepfake participants that triggered transfers. Here too: visual contact in the video is not grounds for approval; critical instructions need the defined second channel.

From term to implementation: KAEMI supports you from the first assessment to the ongoing managed service.