SponsoredHorizon3.ai

Proactive Security for the AI Era

NodeZero continuously and autonomously pentests infrastructure, identity, cloud, and now web applications, chaining weaknesses across every domain the way real attackers do. Every finding ships with replayable proof showing exploitable business impact, not theoretical risk.

See NodeZero WebApp in action

In January 2024, a finance employee at the engineering firm Arup joined a video call with a person who looked and sounded exactly like the company's UK-based CFO, alongside several other colleagues the employee recognized. Over the course of that call, the employee was instructed to process a secret transaction. Every person on the call except the employee was an AI-generated deepfake, built from footage the real executives had already put in public online conferences and company videos. The employee ended up authorizing 15 wire transfers totaling roughly $25.6 million in Hong Kong dollars before anyone at Arup's actual headquarters was contacted to confirm the transaction. No systems were breached. No malware was involved. The entire attack was a live impersonation of people the victim already trusted, executed convincingly enough to survive a real-time video call. This piece is a framework for governance, security, and executive protection teams who need to answer a specific question: what verification protocol actually holds up when the person asking for money, wire approval, or confidential information on a call sounds and looks exactly like the executive it claims to be. It is not a general deepfake-awareness primer. For that broader threat landscape, see our coverage of deepfake fraud enterprise defense and the specific AI voice cloning vishing pattern used against CFOs. This article is about the protocol layer: what to put in place before the call happens, so the outcome does not depend on an employee's ability to spot a fake in the moment.

What actually happened, and what did not

The Arup case is the most documented example of a live deepfake video call producing a real financial loss, and it is worth being precise about what made it work, because the honest version of the story is less exotic than the headlines suggest. Hong Kong police confirmed the attackers built the deepfake participants from publicly available footage: the real CFO and colleagues had appeared in earlier video conferences and company-produced videos, and that footage was harvested and repurposed to drive both the voice and the face on the call. Arup's CIO Rob Greig described the incident afterward as technology-enhanced social engineering rather than a systems intrusion; no corporate network or data was compromised. The fraud was only caught because the employee later contacted Arup's real head office to follow up on the 'secret transaction,' well after the money had already moved.

A second, less-cited case shows the same pattern failing against a prepared target. In May 2024, scammers targeting WPP set up a Microsoft Teams meeting using a WhatsApp account and a voice clone plus YouTube footage of CEO Mark Read, attempting to solicit money and personal details from an agency leader under the pretext of setting up a new business. WPP's own account credits the attempt failing to the vigilance of the people on the call, including the executive being impersonated, not to any technical detection. Two data points, same attack pattern, opposite outcomes, and the difference in both cases was whether a human on the call had a reason to independently verify identity rather than trust the video and audio in front of them. That is the entire premise of this framework: the protocol has to make verification mandatory and procedural, not dependent on someone happening to feel suspicious.

The real state of voice cloning: how fast, and how good

Getting the threat model right matters more than getting it scary. Commercial voice cloning has genuinely crossed the threshold needed for a live phone call, not just a pre-recorded message. ElevenLabs' Flash v2.5 model, built specifically for conversational and real-time use, documents first-audio latency of approximately 75 milliseconds of model inference time, which the company positions as comfortably inside the roughly 500 millisecond budget for natural-feeling conversation and well under the 800 millisecond point where a phone caller starts to notice dead air. Real end-to-end latency on an actual call will run higher once network transport and any live conversational logic are added, but the underlying generation step is no longer the bottleneck it was even two or three years ago. Voice cloning services in this category typically need only a short reference sample of a target's voice, and for senior executives, that reference material is often sitting in public earnings calls, conference keynotes, podcast interviews, and YouTube video, exactly the kind of footage Hong Kong police say was used to build the Arup deepfake.

What this means practically: an attacker does not need a large research budget or custom infrastructure to clone an executive's voice for a live call. They need publicly available audio of the person speaking and access to a commercial or open-source voice cloning tool. That is a meaningfully lower bar than most organizations' incident response planning assumes, and it is why the verification protocol below is built around defeating the attack regardless of how convincing the voice or face is, rather than betting that employees will detect a subtle synthetic tell.

Free daily briefing

Briefings like this, every morning before 9am.

Threat intel, active CVEs, and campaign alerts, distilled for practitioners. 50,000+ subscribers. No noise.

The real state of real-time video deepfakes: what the research actually shows

Live video deepfakes are real and have caused real losses, but the honest technical picture is more constrained than 'any face can be flawlessly faked in real time on any call.' Academic work on real-time video deepfake detection gives a useful window into the actual limits. Researchers behind the Gotcha challenge-response detection approach note that real-time deepfake generation operates under a tight per-frame latency budget, commonly cited in the 30 to 50 millisecond range, which forces attackers to use smaller, distilled, and quantized models running at lower resolution than what offline, pre-rendered deepfakes can use. Those same compromises are what leave behind the temporal jitter, blink irregularities, and gaze inconsistencies that detection research targets. Separate work on deepfake CAPTCHA-style defenses (DF-Captcha) is built entirely around the same premise: ask a live participant to perform an unexpected physical action, a specific head turn, a hand gesture crossing the face, rapid unscripted speech, and real-time face-swap and voice-conversion pipelines measurably degrade or lag under that kind of unscripted challenge in a way a genuine video feed does not.

The practical takeaway for a governance framework is not 'video deepfakes are unconvincing.' Arup's case proves the opposite under normal conversational conditions. The takeaway is that real-time video deepfakes are measurably more fragile under unscripted, physically demanding, or unexpected interaction than they are in a normal, cooperative conversation where the target is not actively trying to break the illusion. That fragility is exploitable, and it is the basis for the human-only verification questions covered later in this piece. It is also why callback verification through a completely separate channel is more reliable than any in-call test: it does not depend on the deepfake's technical limits holding up in the moment at all.

What ElevenLabs and Resemble AI actually publish about detecting their own output

Both companies named as reference points for this threat model also publish real detection and provenance work, and it is worth being precise about what that work does and does not solve, since neither should be read as a reason to relax verification protocols.

ElevenLabs has partnered with Google DeepMind to embed SynthID, an inaudible digital watermark, directly into audio generated through its platform, and has launched a free Audio Detector tool that checks for that watermark before falling back to a legacy AI Speech Classifier model when no watermark is found. ElevenLabs states the SynthID watermark is designed to survive trimming, speed changes, and format conversion. Resemble AI takes a comparable approach with PerTh, an open source, MIT-licensed watermarking model that embeds an imperceptible, psychoacoustically masked signal into generated audio at the point of creation, designed to survive compression, re-encoding, and typical editing. Resemble also offers a separate real-time detection product, Resemble Detect, aimed at flagging synthetic audio frame by frame independent of whether a watermark is present.

The honest limitation, and the reason this matters for a verification framework rather than a vendor comparison, is twofold. First, watermarking only identifies content generated through the watermarking vendor's own platform; it does nothing against a clone made with a different tool, an older model version, or a fine-tuned open source pipeline that never embeds a mark. Second, recent academic research ("The Watermark Shortcut") found that detectors trained on a mix of watermarked synthetic audio and unwatermarked real audio can learn to key off the watermark itself as a shortcut rather than genuine synthetic artifacts, producing failure modes where a watermarked fake can be stripped of its mark to evade detection, or where real human speech gets misclassified as fake once it happens to pick up a similar-sounding artifact. None of this makes ElevenLabs' or Resemble AI's provenance work worthless; it makes it one signal among several, not a replacement for a verification protocol that does not depend on any vendor's detector running correctly on a live call. For a broader comparison of dedicated detection platforms built specifically for live meeting protection, see our enterprise deepfake detection platform comparison.

Protocol 1: pre-established out-of-band verification codes

The single most effective control against both the Arup and WPP attack patterns is a shared secret that never travels through the same channel as the call itself, and that is established before any high-stakes call, not improvised during one.

Issue rotating codes through a separate, pre-vetted channel

Distribute a short verification phrase or numeric code to the small group authorized to approve high-value transactions or attend sensitive board discussions, delivered through a channel that is never the same medium as the call requesting verification (for example, a physical card, a hardware token, or a message in a platform the call itself cannot access). If the code arrives over the same video platform or phone line the suspect call is using, it is not out-of-band and does not count.

Rotate codes on a fixed schedule, not on demand

A code requested and delivered in response to a specific call is worthless if the attacker controls or can intercept that request. Codes should be pre-issued on a monthly or quarterly cadence to the relevant executives and their verified deputies, so that no single call can trigger the code's creation or delivery.

Make code presentation mandatory, not optional, for defined trigger conditions

Define in writing which categories of request require code verification regardless of how well the requester is recognized: any wire transfer or payment authorization above a set dollar threshold, any request framed as urgent or confidential, any request to bypass a normal approval chain, and any board or M&A discussion involving non-public financial terms. Make the policy explicit that failing to present the code ends the call's authority to approve anything, with no exception for 'I recognize your voice.'

Protocol 2: callback to a known number

Callback verification is older than deepfakes and remains the most reliable control specifically because it does not depend on judging the call in progress at all; it moves verification to a second channel entirely.

Maintain a callback directory that is never sourced from the call itself

Every executive, board member, and finance approver who can authorize high-value actions needs a documented phone number on file in a system controlled by the organization, not a number provided by the caller, texted during the call, or listed in an email signature attached to the request. If the callback number comes from the same interaction being verified, it verifies nothing.

End the original call before calling back

The employee should disconnect from the video call or phone line entirely before dialing the known number, rather than being transferred, conferenced in, or asked to 'stay on the line while we connect you.' A request to remain on the original call during verification is itself a red flag, since it is a straightforward way for an attacker to control or spoof the callback.

Require the callback to reach a person, not just ring through

Voicemail, an auto-attendant, or a callback that is answered by someone claiming to be an assistant confirming the request on the executive's behalf does not satisfy the protocol. The policy should require the callback to reach the actual authorizing individual directly, or the transaction is held until it does.

Protocol 3: human-only verification questions

In-call questions have a real but narrow role: they are a supplementary check during the live interaction, useful precisely because they exploit the technical fragility of real-time cloning described above, not a replacement for out-of-band verification.

Ask about shared context that was never recorded or published

A question referencing a private conversation, an internal joke, a detail from an unrecorded meeting, or something only discussed off-camera resists cloning specifically because voice and video cloning are built from existing recorded material. A generative model cannot invent a factually correct answer about an event it was never trained on; it can only guess or deflect.

Request an unscripted physical action on video calls

Consistent with the DF-Captcha research on real-time deepfake fragility, ask the person to perform something a live, unscripted human does naturally but a real-time face-swap pipeline handles poorly: turn fully to one side and back, hold a hand up in front of part of their face, or read a short string of numbers you provide in that moment rather than a rehearsed phrase. Sudden, specific, physically demanding requests are harder for real-time synthesis to track cleanly than a normal head-on conversational pose.

Never use questions with answers available in public materials

Avoid questions whose answers appear in earnings calls, press releases, LinkedIn, or the same public video and audio libraries attackers use to build the clone in the first place, since that is exactly the material a well-prepared attacker will have studied. The Arup and WPP cases both involved attackers using footage the real executives had already made public.

Protocol 4: governance for who can approve a high-value transaction on a call alone

The most important control is organizational, not technical: a written policy that a voice or video call, no matter how convincing, is never sufficient by itself to authorize a high-value transaction or a material disclosure.

Set a mandatory dual-channel rule for anything above a defined threshold

Any wire transfer, payment change, vendor bank detail change, or confidential disclosure above a set dollar or sensitivity threshold requires confirmation through at least two independent channels, one of which cannot be a live call, before it is executed. This should be a board-approved policy, not an informal finance team habit, so it survives leadership turnover and cannot be waived by a single requester regardless of seniority.

Remove single-approver authority for urgent or confidential requests

Urgency and confidentiality are the two framing devices used in nearly every documented executive impersonation case, including Arup, because they discourage the target from checking with anyone else. Policy should explicitly state that a request framed as urgent or secret is not exempt from standard approval chains; if anything, that framing should trigger stricter review, not less.

Train the specific employees who can move money, not just general staff

Awareness training that treats deepfake fraud as a generic phishing topic misses the point for this scenario. Finance staff, treasury, executive assistants, and anyone with wire authority need protocol-specific training: what the callback procedure is, where the callback directory lives, what conditions trigger mandatory out-of-band code verification, and explicit permission to end a call and verify even when the person on screen outranks them.

Document and rehearse the protocol like an incident response plan

A verification protocol that exists only as a policy document is unlikely to be followed under the social pressure of a live call with someone who appears to be the CFO. Run the callback and code-verification procedure as a tabletop exercise at least annually with the actual employees who would execute it, the same way an organization rehearses an incident response plan rather than just publishing one.

The bottom line

The Arup and WPP cases were not defeated or enabled by detection software. Arup's employee had no separate verification step to fall back on and lost $25.6 million; WPP's target had colleagues alert enough to question an unusual request and the attempt failed with no loss. Voice cloning latency is now fast enough for live conversation, and real-time video deepfakes are convincing enough under normal, cooperative call conditions to fool a trained finance professional, so the defense cannot depend on someone spotting a synthetic tell in the moment. It has to depend on a written, rehearsed protocol that makes out-of-band verification, callback to a known number, and dual-channel approval mandatory for high-value or high-sensitivity requests, regardless of how convincing the person on the call sounds or looks.

Frequently asked questions

What was the Arup deepfake board call fraud and how much was lost?

In January 2024, an Arup finance employee joined a video call with AI-generated deepfakes of the CFO and several colleagues, built from publicly available footage, and was instructed to process a secret transaction, resulting in 15 wire transfers totaling roughly $25.6 million before the fraud was discovered.

Can AI voice cloning really work fast enough for a live phone call, not just a recording?

Yes. ElevenLabs documents its Flash v2.5 model producing first audio in approximately 75 milliseconds of inference time, well inside the roughly 500 millisecond budget for natural-feeling conversation, and commercial voice cloning tools generally need only a short public reference sample of a target's voice.

Are real-time video deepfakes as good as pre-recorded deepfake videos?

No. Academic research on real-time detection notes that live deepfake generation operates under a tight per-frame latency budget, commonly cited around 30 to 50 milliseconds, forcing smaller and lower-resolution models than offline deepfakes can use, which is why unscripted physical challenges expose real-time fakes more reliably than passive observation does.

Do ElevenLabs and Resemble AI provide tools to detect their own AI-generated voices?

Yes. ElevenLabs offers a free Audio Detector built on SynthID watermarking developed with Google DeepMind, and Resemble AI offers PerTh, an open source watermarking model, alongside a separate real-time detection product, though both approaches only catch content generated on that specific vendor's platform and academic research has identified ways watermark-based detection can be evaded or produce false positives.

What is the single most effective protocol against a deepfake board call scam?

Callback verification to a pre-documented phone number that was never sourced from the suspect call itself is the most reliable control, because it moves verification to a separate channel entirely rather than depending on judging the authenticity of the call in progress.

Should any single executive be able to authorize a wire transfer based on a video call alone?

No. Organizations should adopt a board-approved policy requiring dual-channel confirmation, one of which cannot be a live call, for any wire transfer, payment change, or confidential disclosure above a defined threshold, with no exception made for requests framed as urgent or secret.

Sources & references

  1. Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee, CNN Business
  2. Incident 983: Scammers Reportedly Used AI Voice Clone and YouTube Footage to Impersonate WPP CEO, AI Incident Database
  3. Criminals Use Generative Artificial Intelligence to Facilitate Financial Fraud, FBI IC3 PSA I-120324
  4. Detecting audio generated by ElevenLabs with SynthID, ElevenLabs
  5. AI Speech Classifier: detect ElevenLabs-generated audio, ElevenLabs
  6. Latency optimization, ElevenLabs Documentation
  7. PerTh Watermarker model, Resemble AI
  8. Multimodal, Real-Time Deepfake Detection at Enterprise Scale, Resemble AI
  9. Gotcha: Real-Time Video Deepfake Detection via Challenge-Response, arXiv
  10. DF-Captcha: A Deepfake Captcha for Preventing Fake Calls, arXiv
  11. The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection, arXiv

Free resources

25
Free download

Critical CVE Reference Card 2025–2026

25 actively exploited vulnerabilities with CVSS scores, exploit status, and patch availability. Print it, pin it, share it with your SOC team.

No spam. Unsubscribe anytime.

Free download

Ransomware Incident Response Playbook

Step-by-step 24-hour IR checklist covering detection, containment, eradication, and recovery. Built for SOC teams, IR leads, and CISOs.

No spam. Unsubscribe anytime.

Free newsletter

Get threat intel before your inbox does.

50,000+ security professionals read Decryption Digest for early warnings on zero-days, ransomware, and nation-state campaigns. Free, daily, no spam.

Unsubscribe anytime. We never sell your data.

Eric Bang
Author

Founder & Cybersecurity Evangelist, Decryption Digest

Cybersecurity professional with expertise in threat intelligence, vulnerability research, and enterprise security. Covers zero-days, ransomware, and nation-state operations for 50,000+ security professionals every morning.

Giveaway: InfoSec World 2026 All Access Pass ($3,895 value)

Details →
Daily Briefing

Subscribe to enter the giveaway

Every subscriber is automatically entered. You also get daily threat intel every morning: zero-days, ransomware, and nation-state campaigns. Free. No spam.

Already subscribed? You're already entered.

Giveaway

Win a $3,895 InfoSec World 2026 pass.