From Shubhi K | Product & Market Analysis

Deepfake CFO Fraud: Detection Is Losing, So Fix the Payment Process

On this page

Deepfake CFO fraud is not a detection problem. Open source detectors lose about half their measured accuracy on deepfakes collected from the wild, so the tool that spots a synthetic face is the wrong first line. Business email compromise still cost $3.05 billion in reported United States losses during 2025. The money leaves through your payment process, so that is where the fix belongs.

Key takeaways

  • Detection accuracy collapses outside the laboratory. Open source video detectors lost 50% of their AUC on in-the-wild 2024 deepfakes, audio models 48% and image models 45%, measured against the benchmarks they were tuned on.
  • The reported loss line is still email, not video. The FBI logged 24,768 business email compromise complaints and $3.05 billion of losses in 2025, alongside $893 million across its first standalone AI fraud category.
  • Arup's finance worker was already suspicious. The deepfake video call was not the attack. It was the verification step, and it passed, which is the part most write-ups skip.
  • Ferrari's attempt failed on a question, not a tool. An executive asked the caller to name a book the real chief executive had recommended days earlier, and the call ended.
$3.05BBusiness email compromise losses reported to the FBI in 2025, across 24,768 complaints. Source: FBI IC3 2025 Internet Crime Report, April 2026.
50%The AUC open source video deepfake detectors lose on real 2024 deepfakes. Source: Deepfake-Eval-2024, arXiv 2503.02857.
76%US organisations that faced attempted or actual payments fraud in 2025. 74% were hit by BEC. Source: AFP, April 2026.

What deepfake CFO fraud actually is now

Business email compromise used to be one channel. An attacker spoofed or took over a mailbox, sent a payment instruction, and hoped nobody picked up the phone.

Picking up the phone was the control that worked. Deepfake CFO fraud is the same crime with that control removed, because the attacker can now answer the phone in your chief financial officer's voice.

Read that as an escalation in the verification layer, not a new category of attack. The wire instruction has not changed. The step you use to check it can now be faked in real time.

Three shapes, one payment

The first shape is a synthetic voice call confirming an emailed instruction. The second is a synthetic video conference, where several apparent colleagues corroborate the request. The third targets onboarding rather than payments, using generated identity documents to open accounts that later receive the money.

The United States Treasury has been explicit about that third shape. In November 2024 FinCEN issued an alert on deepfake media in fraud schemes, after a rise in suspicious activity reports describing generated documents used to get past customer identification controls. Banks filing on this pattern tag the report FIN-2024-DEEPFAKEFRAUD.

The cost of cloning is not the constraint

Producing a convincing clone of a named executive is no longer expensive, slow or technical.

Arup's chief information officer said so plainly after his firm lost money to this exact attack. He told the World Economic Forum that copying a voice, an image or a video is freely available to someone with very little technical skill. He added that this happens more often than people realise.

That matters for budget. If the attacker's cost per attempt is near zero, you cannot win by making each attempt slightly harder to produce. You win by making the payout impossible to collect.

The loss numbers, and what they do not say

The FBI's Internet Crime Complaint Center recorded 1,008,597 complaints in 2025 and $20.9 billion in reported losses. Business email compromise accounted for 24,768 complaints and $3,046,598,558, the second largest loss category behind investment fraud.

The report also carried a standalone artificial intelligence category for the first time, at 22,364 complaints and roughly $893 million. Treat that as a floor rather than a measurement. A victim only files under it if they realised AI was involved, and the point of a good clone is that they did not.

Length of record matters too. The FBI's public service announcement on business email compromise put exposed losses at $55.5 billion across 305,033 incidents between October 2013 and December 2023. This is an industrialised crime with a decade of loss data behind it, not an emerging one.

Survey data fills in what complaint data cannot. The 2026 AFP Payments Fraud and Control Survey found 76% of US organisations faced attempted or actual payments fraud during 2025, with 74% affected by business email compromise. The sample is 465 treasury practitioners, surveyed in January 2026.

Two findings in that survey deserve more attention than they get. Treasury discovered 83% of attempted fraud, which tells you where the effective control already sits. Only 17% of organisations use AI anywhere in fraud mitigation, which tells you the technology arms race everyone writes about is not the one being fought.

How much accuracy detectors lose when the deepfakes are real. Fall in AUC for open source models, against in-the-wild 2024 media. Video models-50% Audio models-48% Image models-45% Sample: 45 hours of video, 56.5 hours of audio, 1,975 images, 88 sites, 52 languages. Commercial detectors scored better, and still below human forensic analysts. Source: Deepfake-Eval-2024, arXiv 2503.02857.
The number to notice is not the size of the drop. It is that the drop happens between the test set and the world, which is the gap a vendor demo lives in.

Why detection technology is the wrong first line

The instinct is to buy a detector. There is now a measured reason to call that the wrong instinct, rather than an opinion.

Detectors lose about half their accuracy outside the lab

The Deepfake-Eval-2024 benchmark assembled deepfakes that actually circulated during 2024, rather than deepfakes generated to order for a research dataset. It covers 45 hours of video, 56.5 hours of audio and 1,975 images, drawn from 88 websites in 52 languages.

Open source detection models scored dramatically worse on it. AUC fell by 50% for video models, 48% for audio and 45% for image, against their performance on the academic benchmarks they were built for.

The authors are careful, and so should you be. Commercial detectors and models fine tuned on the new data beat off-the-shelf open source ones. They still did not reach the accuracy of human forensic analysts.

So the realistic position is this. A detector will catch some attempts and miss others, and it cannot tell you which is which while a payment is waiting. That is a supporting control, not a gate.

The visual tells people were taught to look for have gone too. Blink rate, edge artefacts and unnatural lip sync were 2019 problems, and training a finance team to spot them now teaches false confidence. Confidence is exactly what the attacker is farming.

Two calls, two outcomes

The clearest argument for process over detection is two real incidents where the technology was comparable and the outcomes were not.

Why the Arup video call worked

In January 2024 a finance employee at the Hong Kong office of the engineering firm Arup received an email, apparently from the group's United Kingdom based chief financial officer, about a confidential transaction.

The employee was suspicious. That detail gets lost in most retellings. The human control fired exactly as designed, and the attacker had planned for it.

A video conference followed, populated by the chief financial officer and several recognisable colleagues. Every participant was synthetic, built from footage the company had published itself. The employee then made 15 transfers totalling about $25.6 million to five bank accounts.

Arup confirmed no systems were breached and no data was taken. Rob Greig called it technology enhanced social engineering, which is the correct classification and a hard one for a security budget to absorb, because none of the money should go to the security stack.

Read the sequence again and the lesson is uncomfortable. The video call was not the attack. It was the verification step, and it passed. Any control that consists of moving to a richer channel is now a control the attacker wants you to use.

The question that could not be cloned

In July 2024 a Ferrari executive began receiving messages that appeared to come from chief executive Benedetto Vigna, referring to an unannounced acquisition and an urgent non-disclosure agreement. A call followed, in a convincing imitation of Vigna's southern Italian accent.

The executive did not run a detector. He asked the caller to name the title of a book Vigna had recommended to him days earlier. Bloomberg reported that the call ended immediately.

That question cost nothing, took four seconds, and worked against a technology no product reliably detects. It worked because the answer existed only in a private exchange between two people, and no amount of scraped conference footage contains it.

My own view, stated plainly: that single question is worth more than any detection licence a mid-market finance team will be sold this year. The reason nobody sells it to you is that nobody can invoice for it.

Where a control had to sit to stop the Arup loss. Sequence as disclosed by Hong Kong Police and confirmed by the firm, January 2024. Email from the "CFO" Employee is suspicious. The human control fires correctly. Deepfake video call Suspicion resolved. The verification step is the attack. 15 transfers, 5 accounts HK$200m, about $25.6m, moves. No recovery. Discovery, after the fact Systems were never breached. Nothing for a tool to find. The only place to intervene. Callback originated by finance, on a number held before the request arrived. Source: Hong Kong Police disclosure and Arup statements, 2024.
Three of the four steps offer a security team nothing to act on. The payment step is the only one your organisation fully controls.

Five process controls that actually hold

These are ordered by how much fraud they stop per unit of friction added. Implement them in this order.

Payment controls against synthetic-media authorisation fraud
ControlWhat it stopsWhere it fails
1. Callback on a number you already heldEvery request that arrives carrying its own contact details, which is most of themFails on a stale directory, or if the attacker changed the record first. Lock the directory separately.
2. Dual authorisation on bank detail changes, not only on paymentsThe beneficiary change, which is where the money actually redirectsFails when both approvers sit in one email thread. Require different systems, not different names.
3. A cool-off window before the first payment to a new beneficiaryUrgency, which is the attacker's only real pressure pointFails against genuine same-day obligations. Needs a documented override with a named approver.
4. A shared verification phrase held outside your systemsReal-time voice and video clones, including a full synthetic meetingFails if the phrase lives in email, a shared drive or a password manager the attacker reached.
5. A named right to refuse, rehearsed once a quarterSeniority pressure, which is what converts suspicion into complianceFails without visible executive backing the first time someone stops a real payment.

Controls 1 and 2 do most of the work. The FBI's own guidance has said as much for years, in less detail: use secondary channels or two-factor authentication to verify any request to change account information. The word doing the work there is secondary. A channel the requester chose is not secondary.

Control 4 is the Ferrari control, generalised. A phrase agreed in person and never typed into a system is the only authentication factor a generative model cannot reconstruct from public material.

Control 5 is the one most firms skip and the one Arup argues for hardest. That employee was suspicious and proceeded anyway. Suspicion without a rehearsed permission to stop the payment is not a control, it is a feeling. The same design principle decides whether a human approval step in an automated workflow is meaningful or theatrical: whether the human may say no without consequence.

The unglamorous objection to all five is friction. A cool-off window will, at some point, delay a genuine payment that mattered. Treat that as a priced trade rather than a failure. Write down the override path before you need it, name who authorises it, and log every use. A control with no override gets bypassed informally within a quarter, and an informal bypass is worse than none, because you now believe you have one.

Where these five controls fail

They fail together, in one scenario. If the attacker holds a mailbox inside your finance function, they can watch the callback being arranged and intercept it.

That is an access problem, not a payments problem, and it needs the boring answers: phishing-resistant multi-factor authentication, alerting on inbox rules, and a review of who can edit vendor master data. Service accounts and machine identities are often the weakest link there, a pattern examined in the piece on non-human identity and access.

They also fail on unmanaged tooling. A finance team routing approvals through a chat app or an AI assistant nobody has inventoried has built a channel outside all five controls, which is one reason unsanctioned AI usage carries a measurable breach cost premium.

The channel that loses the most money is not the futuristic one. Share of US organisations reporting fraud by payment method, 2025. Cheques58% ACH debits30% Wire transfers25% Only 17% of these firms use AI anywhere in fraud mitigation. Treasury finds 83% of attempts. Source: 2026 AFP Payments Fraud and Control Survey, 465 practitioners, January 2026.
Deepfakes get the coverage and cheques get the money. A control programme built only against synthetic media is defending the smaller bar.

Where regulators moved first, and what they left out

Two regulatory moves are worth tracking, because both act on the payment rather than on the media.

The first is Verification of Payee. Under the European Union's Instant Payments Regulation, euro area payment service providers have had to offer a free payee name and IBAN consistency check since 9 October 2025. Providers outside the euro area have until 9 July 2027. The payer sees a match, close match or no match response before authorising.

That is the correct shape of intervention. It does not try to judge whether a video call was real. It checks whether the account you are about to pay belongs to the party you think you are paying, which is the fact the fraud depends on being wrong.

The second is reimbursement. Since 7 October 2024 United Kingdom payment firms have had to reimburse authorised push payment scam victims up to £85,000, and the Payment Systems Regulator reported that 88% of money stolen in the first nine months of the regime was returned.

Do not read that 88% as protection for your treasury function. The United Kingdom regime covers consumers, micro-enterprises and small charities, so a mid-sized company wiring $500,000 on a deepfaked instruction sits outside it. In the United States there is no equivalent for commercial accounts at all. A business wire authorised by an employee is, in most cases, an authorised payment, which is why the internal control is the whole defence and why the clauses a CFO negotiates into AI and payments contracts carry more weight than the security marketing attached to them.

Where this argument is weakest

Two honest objections, including one about the evidence in this post.

The statistics circulating about this topic mostly do not trace

While researching this piece I found the same striking claims repeated across a dozen vendor blogs. Deepfake-enabled business email compromise up 312% year on year. Forty per cent of BEC attacks now containing synthetic audio or video. A 3,892% surge in deepfake fraud.

None of them traces to a document you can open. Several are attributed to the FBI's 2025 report, which does not contain them. Six restatements of an unsourced figure is one unsourced figure, and the correct response is to leave the number out and say why.

So this post makes a narrower claim than most coverage does. The published data proves business email compromise is enormous and that detection generalises poorly. It does not prove what share of BEC now involves a deepfake, because nobody has published that with a method attached. Anyone quoting a precise percentage there is quoting a vendor.

The honest case for buying detection

There is a real case, and it applies where process controls have nothing to grip.

High volume remote onboarding is the clearest example. If ten thousand people a month open an account and no human speaks to any of them, there is no callback to make and no shared phrase to ask for. That is the scenario FinCEN's alert describes, and automated liveness and document checks are the only available control.

Detection also improves. The same benchmark that embarrasses open source models shows commercial ones doing better, and fine tuned models better still. A detector honest about its false negative rate is a reasonable second layer. It is a bad first one.

Frequently asked questions

What is deepfake CFO fraud?

Deepfake CFO fraud is business email compromise with a synthetic voice or video layer added. An attacker sends a payment or bank detail change request impersonating a senior finance executive, then uses AI generated audio or video to pass the verification step when staff try to confirm it. The best documented case is Arup, which lost about $25.6 million in January 2024 after a finance employee joined a video call where every other participant was synthetic.

How much does business email compromise cost businesses?

The FBI's Internet Crime Complaint Center recorded 24,768 business email compromise complaints in 2025 with $3.05 billion in reported losses, making it the second largest loss category after investment fraud. Over the longer record, the FBI put exposed losses at $55.5 billion across 305,033 incidents between October 2013 and December 2023. Both figures rest on voluntary complaints, so they understate the real total.

Can deepfake detection software stop CFO fraud?

Not reliably on its own. The Deepfake-Eval-2024 benchmark found open source detectors lost 50% of their AUC on video, 48% on audio and 45% on images when tested against deepfakes actually circulating in 2024 rather than laboratory datasets. Commercial tools performed better but still below human forensic analysts. Use detection as a second layer behind payment process controls, never as the gate on a transfer.

How do you verify a payment request from your CEO?

Call back on a number you already held before the request arrived, never a number supplied in the request itself. Add a verification phrase agreed in person and stored outside your email and file systems. Require two approvers in two different systems for any change to bank details. A richer channel is not verification, because a video call can be fully synthetic, as the Arup case demonstrated.

Does insurance or the bank cover deepfake wire fraud losses?

Usually not automatically. The United Kingdom reimbursement rules that took effect on 7 October 2024 cap payouts at £85,000 and cover consumers, micro-enterprises and small charities, not mid-sized or large companies. In the United States there is no equivalent protection for commercial accounts, and a wire authorised by your own employee generally counts as authorised. Check whether your crime policy names social engineering explicitly.

What is Verification of Payee and does it stop deepfake fraud?

Verification of Payee is a free check that compares the payee name you enter against the name registered to that account, returning a match, close match or no match. Euro area payment providers have been required to offer it since 9 October 2025 under the Instant Payments Regulation. It does not detect a deepfake. It flags that the destination account does not belong to the party you believe you are paying.

Where to start this week

Two moves, both cheap, both doable before your next payment run.

First, agree a verification phrase with each person authorised to approve payments above your threshold. Do it in person or on a call you initiated, write it nowhere digital, and use it on the next request that feels urgent. That is the Ferrari control and it costs nothing.

Second, audit one thing: how your finance team gets the callback number. If it comes from the request, the signature block, or a recent email, you do not have a callback control. You have a formality the attacker already owns, and fixing it is a directory change rather than a purchase.

Related on controls and exposure

If your approvals run through automated workflows, read what makes a human-in-the-loop step real rather than decorative. For AI's share of reported crime by region, see the analysis of AI enabled cybercrime.

References

  1. FBI Internet Crime Complaint Center, 2025 Internet Crime Report, April 2026. Complaint volumes, total losses, BEC losses, AI category.
  2. FBI Internet Crime Complaint Center, Business Email Compromise: The $55 Billion Scam, 11 September 2024. The 2013 to 2023 exposed loss total and the secondary channel guidance.
  3. Chandra and others, Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024, arXiv 2503.02857. Every detector accuracy figure and the dataset composition.
  4. Association for Financial Professionals, 2026 Payments Fraud and Control Survey, April 2026. The 76%, 74%, payment method and AI adoption figures. Sample 465 treasury practitioners.
  5. FinCEN, Alert on fraud schemes involving deepfake media, FIN-2024-Alert004, 13 November 2024. The identity document typology.
  6. Bloomberg, Ferrari narrowly dodges deepfake scam simulating deal-hungry CEO, 26 July 2024. The Ferrari attempt and the book question.
  7. World Economic Forum, Lessons learned from a $25m deepfake attack, February 2025. The Rob Greig quotes and the Arup description.
  8. European Commission, New EU rules make instant euro payments faster and safer, 10 October 2025. Verification of Payee dates and scope.

Weakest part of this source base: the Arup transfer detail rests on police disclosure relayed through press reporting rather than a filing, and no primary dataset exists on what share of business email compromise now involves synthetic media. Figures are current as of 26 August 2026.

RR
Ritu Raj
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading