A couple of months ago, I wrote the first article in this deepfake blog series and received positive feedback from readers in various industries. The initial post focused on the current state of deepfake threats and the extent to which detection tools can keep up. Today, I would like to shift our attention to how cybercriminals build deepfake campaigns. While statistics show that older people often lose more money and younger people are more susceptible to certain scams, recent incidents reveal attackers tricking older victims into believing they are speaking with relatives. At the same time, executives are being impersonated to pressure employees into making urgent transfers. Understanding how these attacks are crafted provides insight into why they are so effective and are spreading so quickly.
Their beginning point is always data. Publicly available material such as social media videos, podcasts, interviews, casual voice notes, serves as the basis for generating synthetic media. Attackers no longer need extensive or high-quality recordings. Modern voice-cloning tools can reproduce someone’s tone and style with just a few seconds of audio. For visuals, a few photographs or short video clips are enough to train face-swap or lip-sync models that produce convincing enough outputs to fool many observers.
After gathering that material, the fraudsters compose scenarios that exploit the target’s trust. A voice impersonation alone is weak unless embedded in a storyline, perhaps a family emergency or a sudden financial crisis, where urgency is emphasized. The synthetic media then acts as the evidence of that story. Victims are not only hearing or seeing something familiar but are also pressured by time or fear, which reduces the chance they will question what they perceive. Europol’s 2025 Internet Organised Crime Threat Assessment (IOCTA) reports that criminals are increasingly using generative tools to impersonate persons of trust in multi-lingual messages or voice-cloned calls, and these schemes are causing growing financial and reputational damage across both people and organizations.
Technical setup is refined to hide obvious signs. Attackers lower resolution, compress video, insert ambient noise, or slightly distort visuals so that detectors or human perception are misled. Recent research shows that detection tools that perform very well in controlled environments lose substantial accuracy when faced with content compressed by social platforms or in less ideal lighting. The Deepfake-Eval-2024 benchmark, for example, collected in-the-wild media (44 hours of video, 56.5 hours of audio, nearly 2,000 images from dozens of languages and websites) and found that open-source state-of-the-art detectors drop in performance by around 50% for video detection, 48% for audio, and 45% for image detection compared to their scores on older academic datasets.
Meanwhile, attackers also scale up. Deepfake-as-a-service platforms allow clients with little technical skill to order impersonations services. GPU time is rented, pre-trained models are shared, and simple applications enable custom synthetic voice or facial content. Thus, what once required expert work now becomes available to less-sophisticated cybercriminals. The result is that attacks are increasingly widespread, not only targeting high-profile individuals but regular people, small firms, and non-profit organizations.
To make the scam feel real, victims receive messages, calls, or fake visuals that seem familiar, mixed with emotional or urgent pressure that pushes them to act quickly. Even when doubts emerge, the presence of synthetic “proof” makes hesitation difficult. This entire process is not accidental but assembled by design: collecting authentic materials, crafting synthetic replicas, embedding them in persuasive narratives, and delivering them via channels the victim believes are reliable.
These campaigns have been successful lately because they exploit both the blind spots in detection technologies and the mental shortcuts people often rely on under pressure. Comprehending this complex design is necessary for creating defenses that are meaningful, whether through more stringent verification methods, better employee awareness training, or improvements in detection technologies. In the next blog, I will discuss these practical measures such as better verification, awareness training, and advanced detection tools that organizations and individuals can use to counter these threats.