The Arup deepfake: $25M lost on a synthetic video call
By Elias Lankinen10 min read
I have enough well-sourced material and four verified Wikimedia image URLs. Writing the post now.
The meeting that never happened
Early in 2024, a finance employee in the Hong Kong office of Arup, the British engineering firm behind the Sydney Opera House and the Beijing Olympic "Water Cube," sat down for a video call. On screen was the company's chief financial officer, dialing in from the United Kingdom, along with several other colleagues the employee recognised. The CFO explained that a confidential transaction needed to happen, quickly and quietly. Over the course of the call and the days around it, the employee made 15 separate transfers to five different Hong Kong bank accounts, totalling roughly HK$200 million, about US$25 million. Every other person on that call was a fabrication. The CFO was not the CFO. The colleagues were not colleagues. According to the Hong Kong police account reported by The Register, the only real human in the meeting was the victim. The rest were AI-generated forgeries of people who existed, saying things they never said, in a meeting that, for everyone but the employee, never took place.
The case became, almost overnight, the reference example for a new category of fraud: not a hacked account or a stolen password, but a meeting populated by synthetic people. It is worth slowing down and looking at exactly what happened, because the details are more instructive, and more unsettling, than the headline number suggests.
The call that cost $25 million
The sequence began, as most of these frauds do, with a message. In January 2024 the Hong Kong employee received communication purporting to come from Arup's UK-based CFO, describing the need for a secret transaction. As Arup later confirmed to Dezeen, the staff member was initially suspicious. The request had the classic shape of a phishing attempt: an executive, a confidential deal, an unusual payment. What dissolved that suspicion was the video call. Seeing the CFO's face and hearing a familiar voice, alongside other recognisable colleagues, was enough to convince the employee that the request was genuine. The impersonated CFO asked for the transfers, and the employee complied, executing 15 transactions to five local accounts. The fraud unravelled only afterward, and through an ordinary act rather than a clever one. The employee eventually contacted Arup's actual headquarters to follow up on the "secret transaction." The real executives had authorised nothing, held no such meeting, and knew nothing about any video conference. The money was already gone. There is a detail here that most retellings flatten, and it matters. Hong Kong senior superintendent Baron Chan Shun-ching, describing the case to reporters, said: "I believe the fraudster downloaded videos in advance and then used artificial intelligence to add fake voices to use in the video conference." In other words, the police working theory was not necessarily a fully interactive, real-time puppet responding live to the employee's questions. It may have been prerecorded video footage of the executives, scraped from public sources, with AI-generated speech layered over it, stitched into something that looked like a live meeting. The scammers reinforced the illusion through other channels too, using WhatsApp, email, and one-on-one video to build credibility. Whether the call was a genuine real-time deepfake or a cleverly assembled recording remains, to some degree, uncertain. What is not uncertain is that it worked.
What a deepfake actually is, and is not
The word "deepfake" is a compression of "deep learning" and "fake." Deep learning is the branch of AI that trains large neural networks on huge quantities of examples until they can reproduce the patterns in that data. Point such a system at thousands of images and audio clips of a specific person, most of which, for a public-facing executive, are sitting on the open internet, and it can learn to generate new footage of that person's face and new speech in that person's voice.
There are a few distinct techniques bundled under the single word, and the distinction matters for understanding both the Arup attack and the defences against it. Face swaps graft one person's face onto another's body in video, and are the variety most capable of defeating facial-recognition checks. Voice cloning reproduces a target's speech from a sample that can now be as short as a few seconds. Real-time or "live" deepfakes run inside a video call as it happens, feeding synthesised frames into a virtual camera, a piece of software that presents itself to your computer as if it were a physical webcam, so that Zoom or Teams simply receives the fake feed and treats it as real. Live deepfakes are the hardest to pull off convincingly, because everything has to be computed fast enough to hold a natural frame rate, which forces compromises in resolution and detail. This is where the most common misconception needs correcting. Many people assume a fraud like Arup's must have involved "hacking," some breach of the company's systems. It did not. Arup's global chief information officer Rob Greig was explicit on this point, telling reporters that "none of our internal systems were compromised" and that the firm's financial stability and operations were untouched. He described the attack as "technology-enhanced social engineering." No firewall was breached. No malware was installed. The technology did not attack the network. It attacked the person, by manufacturing the single thing humans have always used to decide whom to trust: the sight and sound of a familiar face.
Why it worked
It is tempting to treat the victim as gullible. That instinct is worth resisting, both because it is unfair and because it is the exact complacency that makes the next victim. Consider the specific combination of pressures assembled here. First, authority: the request came from the CFO, one of the few people in any company whose instruction to move money is routine and unquestioned. Second, urgency: the transaction was framed as time-sensitive, which is a deliberate device to short-circuit the deliberation that would otherwise catch a fraud. Third, secrecy: labelling it confidential discouraged the employee from doing the one thing that would have exposed it immediately, which is asking a colleague, "did you hear about this deal?" And fourth, multisensory confirmation: the employee's initial, correct suspicion was overturned not by argument but by evidence, the apparent evidence of their own eyes and ears. For most of human history, seeing a person's face and hearing their voice in a live conversation was a reliable proof of identity. Forging a written signature was possible; forging a moving, speaking human in real time was not. Our institutions, and our instincts, are built on that assumption. The video call felt like the verification step, the thing you do to be sure. That is precisely why it was the vector of attack. The fraud did not defeat Arup's controls so much as impersonate them. The real weakness, then, was not a person but a process. There was, at the decisive moment, no channel of verification independent of the one the attackers controlled. Everything the employee used to confirm the request, the video, the voices, the follow-up messages, came through the fraudsters' own pipe. A single out-of-band check, a phone call to a known number, a message on a separate system, a face-to-face question, would have collapsed the whole thing. The employee eventually made exactly that check. It just came after the money left.
Not the first, and nowhere near the last
The Arup case felt novel, but its lineage runs back years. In 2019, criminals used AI to clone the voice of the chief executive of a German energy company's parent firm and phoned the head of a UK subsidiary, demanding an urgent transfer of €220,000 (about $243,000) to a Hungarian supplier. The employee later told investigators he recognised the "slight German accent and the melody" of his boss's voice. He paid. The money was swept from Hungary to Mexico and dispersed. When the fraudsters called back for a second and third transfer, the growing implausibility finally raised suspicion, and the later requests were refused. That case, widely reported as the first known audio-deepfake fraud, involved only a voice on a phone. Five years later, Arup's attackers had a boardroom full of faces. What changed between those two cases is not the trick, which is ancient, but the cost and quality of the forgery. And that curve is steep. Deloitte's Center for Financial Services projects that generative AI could push fraud losses in the United States to $40 billion by 2027, up from $12.3 billion in 2023, a compound annual growth rate of about 32 percent. Deloitte attributes part of this to an "entire cottage industry on the dark web" selling scam software for anywhere from $20 to several thousand dollars, which puts capabilities that were once the preserve of well-resourced operations into far more hands. The aggregate figures around deepfakes specifically should be read with some caution, because definitions vary between reports and vendors have an incentive to alarm. But even the conservative direction of travel is clear, and it points one way: up, and fast.
The uncomfortable economics
The most sobering part of the Arup story is not the loss. It is what Rob Greig did afterward. According to his account to the World Economic Forum, he decided to see how hard the underlying technology actually was. Using freely available open-source software, he made a passable deepfake video of himself. It took him about 45 minutes. That number reframes everything. The barrier to this kind of attack is no longer technical skill or money. Greig noted that the tools are "freely available to someone with very little technical skill to copy a voice, image or even a video." The raw material, footage and audio of a company's executives, is already public for any firm with a website, a conference talk, an earnings call, or a LinkedIn video. The scarce ingredient in the fraud is not the deepfake. It is a target organisation without an out-of-band verification habit. Greig has been candid that Arup is not an outlier in being attacked, only in being attacked successfully and then talking about it. "Like many other businesses around the globe, our operations are subject to regular attacks, including invoice fraud, phishing scams, WhatsApp voice spoofing, and deepfakes," he said, adding that "the number and sophistication of these attacks has been rising sharply." His blunter summary to the WEF: "We are attacked every day." Part of why Arup went public, having initially been described only as an anonymous multinational when Hong Kong police disclosed the case in February 2024 before the firm named itself in May, was to make exactly this point. Silence protects reputations in the short term and leaves everyone else undefended.
What actually stops this
The instinctive response to a deepfake problem is a deepfake-detection solution: software that scans a video feed for the tell-tale artefacts of synthesis. Detection tools exist and are improving, and they have a role. But relying on them as the primary defence is a losing game, because it is an arms race between generators and detectors in which the generators keep getting better, and because live detection has to work inside the same brutal time budget that constrains the fakes themselves. The more durable defences are procedural, and they are unglamorous:
- Out-of-band verification for money movement. Any transfer above a threshold requires confirmation through a separate, pre-established channel, calling the requester back on a number you already have, not one supplied during the request. This single control would have stopped the Arup fraud outright.
- Break the urgency-and-secrecy pattern. Policies should treat "urgent and confidential, do not tell anyone" as a red flag in itself, not a reason to comply faster. Legitimate executives can tolerate a verification delay. Fraudsters depend on you skipping it.
- Dual authorisation. Large payments should require two people, so no single deceived employee can complete a transfer alone.
- Shared verification cues. Some organisations have adopted agreed pass-phrases or challenge questions for sensitive requests, a low-tech check that a scraped video cannot answer. None of this is exotic. Most of it predates AI entirely, because the vulnerability predates AI entirely. The deepfake did not invent the con of the urgent, confidential, authority-backed payment request. It only supplied a more convincing costume. Arup itself has emphasised that its systems were never breached, which is precisely the point: the fix lies in process and training, not in a better firewall. The harder question is what happens as the costume becomes flawless. For now, verification still works because there remains a real person, on a real known number, who can say "I never asked you to do that." But we are watching the erosion of a very old assumption, that a face and a voice, live and in conversation, are proof of who is speaking. When that assumption is gone, trust has to be rebuilt on something the camera cannot show. The organisations that come through this well will be the ones that stopped trusting their eyes before their eyes gave them a $25 million reason to.
Sources
- The Register, "Deepfaked 'CFO' tricks British engineering firm Arup into paying out $25 million", 2024
- Dezeen, "Arup victim of multimillion-dollar deepfake video scam in Hong Kong", 2024
- World Economic Forum, "Cybercrime: Lessons learned from a $25m deepfake attack", 2025
- CNN Business, "Arup revealed as victim of $25 million deepfake scam involving Hong Kong employee", 2024
- Sophos Naked Security, "Scammers deepfake CEO's voice to talk underling into $243,000 transfer", 2019
- Deloitte Center for Financial Services, "Deepfake banking and AI fraud risk", 2024
- Biometric Update, "Deloitte predicts losses of up to $40B from generative AI-powered fraud", 2024
- AI Incident Database, "Incident 634: Alleged Deepfake CFO Scam Reportedly Costs Multinational Engineering Firm Arup $25 Million", 2024
- Adaptive Security, "Types of AI Deepfakes: A Complete Guide to Every Synthetic Media Threat", 2025