Ever watched yourself on video and cringed?
Under normal conditions, experiencing our digital reflection can feel surreal or even uncomfortable. So first off, we commend our participating execs for allowing us to use their publicly available personal data to create live audio/visual doppelgangers – as we found out just how advanced, believable, and potentially malicious our identity cloning tools currently are.
To be clear, this is not the pre-rendered deepfakes used to create fake influencer accounts, corrosive disinformation campaigns (e.g. Slovakian corruption audio) or the concerning rise of deepfake pornography. The identity cloning frameworks we use enable real-time face re-enactment and voice conversion, which we set up, fine-tuned, and leveraged to convince our colleagues that we were actual JUMPSEC executives on internal video calls.
We dubbed this exercise ‘Project Havoc’ because the sophistication and accessibility of real-time identity cloning has the potential to fundamentally undermine trust in digital identity. As we’ll demonstrate, it can be used to manipulate employees, bypass assumptions of authenticity, and erode the basic premise that seeing and hearing someone is evidence of who they are.
How we did it
The setup combined the following components:
- Free open-source GitHub repos (VisoMaster, Deep-Live-Cam, Vonovox, RVC v4).
- A mid-to-advanced range gaming laptop equipped with Graphics Card that supports CUDA, 16GB RAM minimum and SSD Storage. Nothing out of this world.
- Other basics such a Microphone, Webcam, Virtual Audio Cable, cuDNN API, PyTorch and OBS Broadcaster with Virtual Webcam.
- A degree of music/video production skill to fine tune and sync the audio-visual components software and;
- A surprisingly concerted effort to dialect coach ourselves to adopt our speech to mimic the cloned executive to replicate unique intonation, accent, and general vibe.
All we needed was ~10 minutes of audio and a handful of static images. As accomplished business leaders, our executives aren’t camera shy and there was plenty of public audio and video material from which to build a convincing synthetic identity clone.
As synthetic identity cloning tools advance, encouraging executives to hide their identity on socials is a wise strategy from an attack surface reduction perspective. But as many organisations encourage leadership to take centre stage in external communications, images and audio/video recordings are often readily available.
To demonstrate the degree of believability at play we’ve cloned a more recognisable but no less charismatic executive. Who better to add another layer of futuristic surrealism than Elon Musk.
This framework combines several social engineering capabilities:
- Face swapping, where the cloning subject’s likeness is mapped onto the attacker’s facial movements.
- Facial re-enactment, where expressions, head pose, and eye movement are translated to a synthetic model.
- Voice conversion, transforming the attacker’s voice into that of the target while preserving speech timing.
- Lip synchronisation, ensuring audio and facial movements remain believable in real time.
- Virtual camera injection, allowing the synthetic output to appear as a legitimate webcam feed in collaboration platforms.
Unlike traditional deepfakes that require extensive rendering time, this real-time system prioritises latency over perfection. This means we can achieve a slightly imperfect but interactive clone which is more persuasive than a flawless pre-recorded video because the victim can ask questions and receive immediate responses.
If the realism of facial movements, the audio-visual synchronicity, or lack of lag time has you mentally projecting high-impact scenarios, let’s walk through the known and hypothetical attack paths.
Known attack paths
In early 2025, global engineering firm Arup lost $25 million after finance staff were manipulated into authorising a fraudulent transfer during a video call in which other participants were deepfakes. Around the same time, widespread reporting emerged of North Korean nationals using AI-assisted identity fraud to secure remote technical work, retaining insider access to sensitive systems while funnelling salaries back to the regime.
And from ransomware and data extortion perspective, as many organisations have hardened phishing defences and MFA, attackers have pushed harder into the human layer with IT helpdesk vishing, MFA fatigue, and voice-based pretexting. This means groups who already successfully leverage social engineering techniques, like Scattered Spider (now more broadly Scattered LAPSUS$ Hunters), may use synthetic media as a natural addition to their toolset.
The three scenarios below represent the attack paths we therefore consider most immediately viable against Financial, IT/Security Operations and HR business functions.
The paths illustrated above represent what we assess to be the most immediately executable scenarios. However, it would be naive to treat this as an exhaustive list. For exmaple, supply chain impersonation targeting MSP relationships and executive cloning deployed against counterparties in M&As or legal proceedings are also plausible high-impact vectors.
What we learned
Since early 2026, tooling has progressed from modelling and re-projecting underlying facial geometry from 5-10 frames per second to ~30 frames per second. Currently, the minimal lag is due more to GPU/graphics restrictions than the tooling itself and, at this rate, the lag will likely become undetectable.
The tools required to believably clone anyone with a decent LinkedIn profile picture and 10 mins of audio are available, low- no- cost, and require minimal amount of preparatory open-source research to manipulate colleagues into carrying out unauthorised malicious actions.
However, an attacker’s ability to convince a subject does require certain conditions:
- How familiar the victim is with the cloned individual is significant. In successful impersonation instances, the relationship between the target and the cloned person was >1 year – a few months. Where the relationship was long standing, for example a +1-year close relationship, attempts to impersonate were unsuccessful due to personality traits and mannerisms.
- Seniority is key. In cases where the target of impersonation was convinced of the clone’s authenticity, and subsequently agreed to undertake requested actions, seniority was a persuasive factor. In essence, employees are predisposed not to push back on direct requests from senior personnel, and thus directly assigned tasks do not need excessive explanation or persuasion to elicit compliance.
- Likeness to the impersonation target’s head shape, facial structure, and hair (or lack thereof) are limitations. As are idiosyncratic accents and speech patterns. The more unique you are, the harder you will be to clone convincingly.
- Social engineering skills and preparation are still required. As an impersonator, you cannot hear yourself as you speak, and therefore without live feedback it is not easy to match the impersonated person’s mannerisms, speech patterns, and intonation. This is a new addition to the social engineer’s toolset, and not a replacement for social engineers by any means.
Mitigations
The first step of individual’s awareness and preparation to combat the threat of is in its infancy. Real cases such as the North Korean IT workers and Arup financial fraud should be referenced to clarify that the risk is not hypothetical.
Basic mitigations discussed in 2025 to simply ask a suspicious person to wave their hand in front of their face no longer distorts the face-swap. Out-of-band mechanisms are required. Every financial authorisation, credential reset, or sensitive action should incorporate an out-of-band verification mechanism that cannot be reproduced in real time by AI. This may include:
- Secondary approval through a separate communication channel.
- Hardware-backed authentication using FIDO2 security keys.
- Cryptographically signed requests.
- Callback procedures using pre-established contact details.
- Four-eyes approval workflows requiring multiple individuals to authorise high-risk actions.
- Delayed execution periods for exceptional requests.
Ultimately, organisations should shift away from treating voice and video as evidence of identity. In an era of synthetic media, “seeing is believing” is no longer a security control. Nonetheless, encouraging employees in high-privileged or executive roles to rethink their online presence as a personal and professional ‘attack surface’ will help to reduce the volume of images and video attackers require to create convincing campaigns.
What’s next?
Tooling, particularly for voice impersonation is already effective, and has already proven useful on recent social engineering engagements. Executives or high privileged technical personnel, such as IT support, security and DevOps engineers, will likely be high value cloning targets.
It is important to monitor emerging vulnerabilities within productivity and collaboration platforms. In late 2025, researchers demonstrated multiple Microsoft Teams weaknesses capable of enabling message spoofing, notification manipulation, and executive impersonation scenarios (i.e. Check Point’s Microsoft Teams spoofing research). Although patched, such vulnerabilities illustrate how trust assumptions within collaboration platforms can be exploited to amplify social engineering attacks. Threat actors have increasingly abused external Teams collaboration and impersonated IT support personnel to convince users to grant remote access, blurring the distinction between traditional phishing and trusted internal communications.
Ethical and Legal Considerations
Rest assured, as part of this project we gained preparatory permissions and conducted after-call feedback, taking steps to clarify the experience as a training and research. With such considerations, cloning can be used for social engineering within adversarial simulations (e.g. Red Teaming) by ensuring appropriate permissions are aligned to company policies on security assurance training.
In both deepfake likeness and live identity cloning, the full ethical and legal considerations are not fully developed. Depending on the jurisdiction, legal considerations need to be updated to mitigate the dangers of current capabilities, and prosecute individual who utilise advancing technologies to maliciously manipulate others.
Within the European Union, the AI Act introduces transparency requirements for synthetic content, requiring certain AI-generated audio, image, and video material to be labelled. However, the Act primarily addresses disclosure rather than criminal misuse itself. In the United Kingdom, existing legislation such as the Fraud Act 2006, Computer Misuse Act 1990, and Online Safety Act provide mechanisms to prosecute malicious uses of synthetic identities, though none were designed specifically with real-time identity cloning in mind.
As an emerging threat vector, awareness is critical at individual, organisational, legal and political levels. Defences will depend less on improving the realism of detection signals (such as visual or vocal cues) and more on strengthening identity verification processes, particularly for high-risk actions. Currently, this includes shifting reliance away from real-time audio/visual confirmation and toward cryptographically verifiable authentication, out-of-band approvals, and robust internal authorisation workflows.
About JUMPSEC
JUMPSEC is a specialist cybersecurity consultancy delivering advanced threat intelligence, offensive security, and cyber risk management services to organisations operating in complex and high-risk environments.
For more information on JUMPSEC visit: Leading Cyber Security Services Company, UK | JUMPSEC
Sean Moran
Sean is Head of Research & Enablement at JUMPSEC, where he works on cyber threat intelligence research, ransomware and extortion analysis, and the production of technical CTI reports.
Jack Lewis
Jack is a security researcher with a strong focus on malware analysis, tracking new threat actors and campaigns, reverse engineering, patch diffing, and proactive threat hunting.
