VoiceGuard · bank voice channel

Twenty-one ways in.
Most teams
defend twelve.

A real-time deepfake and vishing detection layer for the bank voice channel. Passive on the line, explainable in the verdict, evidence-grade in the audit trail.

Book a 15-minute call Try the Audio Challenge →
Shipped detector · public benchmark
3.05% Equal Error Rate — ASVspoof 5 dev track_1
AUC · 0.9943
N · 20,000
MODEL · v11_doomsday
MEASURED · 2026-07-21
Live verified · v11.detector.eer
Delivery modelOn-premise, customer-controlled
PositionPassive on the call
Runs today in our own development and demo environment on NVIDIA Triton. No customer production deployment.
01
The threat surface

Published cases
Not hypotheticals

Twenty-one vectors, each with a documented case.

Every vector below maps to a published case or a regulator-recorded trend. This is the fraud map of the bank voice channel — not a threat landscape slide.

Against the customer

Group A · 7 vectors
A1Bank or authority impersonation (vishing)
A2Authorised push payment fraud
A3Customer calls the bank on a scammer's script
A4Cloned-relative “grandparent” scam
A5SIM-swap into account takeover
A6One-to-one deepfake voice call
A7Multi-party deepfake conference

Against the bank

Group B · 8 vectors
B1Voice clone through IVR Voice ID
B2Voice clone into internal or treasury fraud
B3Human social engineering against the agent
B4SIM-swap into OTP bypass
B5Insider or rogue agent
B6Synthetic-identity or voice-KYC bypass
B7Account-recovery hotline engineering
B8IVR reconnaissance

Infrastructure

Group C · 3 vectors
C1Caller-ID spoofing — including the regulator's own number
C2Call-forwarding exploit
C3Channel-hopping to an OTT app

Asynchronous

Group D · 3 vectors
D1Smishing escalating into vishing
D2Recovery-room scam and re-victimisation
D3Voicemail or voice-note deepfake
21 documented vectors · case library on request VoiceGuard defends the bank-side surface passively on the call. For the rest, the boundary is stated below rather than blurred.
02
What it does

One pass
Two threats

A perfect clone passes the match.
So we ask the other question.

Voice biometrics asks who is this. When the clone is good enough, the answer is wrong and confident. VoiceGuard runs the orthogonal question on the same audio.

Pillar 1 · Deepfake

Is this a live human voice, or generated?

A four-stream KAN ensemble scores every inbound caller automatically — no enrolment, no voiceprint, nothing for the customer to do. The training corpus spans 33 synthesis engines.

Measured · ASVspoof 5 dev track_1
Pillar 2 · Vishing & social engineering

Is a manipulation pattern being spoken?

Speech recognition plus a fraud-keyword engine runs on both legs, caller and agent, and fires even when there is no deepfake at all: safe-account pretext, OTP coaching, channel-hopping, or an agent policy violation.

Deterministic rules · repo verified

The two pillars are not averaged into one number. The system returns which of them fired:

VISHING_SCAMDEEPFAKE_DETECTEDDEEPFAKE_VISHING
03
How it connects

Observes
Does not carry

Nothing of ours sits in the media path.

If VoiceGuard fails, the call does not. That is an architectural property, not an operational promise.

THE CALL — UNTOUCHED Caller Capture point Agent MIRRORED COPY VoiceGuard Your fraud systems NO RETURN PATH INTO THE CALL — VOICEGUARD NEVER BLOCKS, HOLDS OR ROUTES
D1 · Topology — observation branch only
TodayIn-app and conference-join capture. This is what is shipped and what a first deployment uses.
Shipped
Enterprise-scale targetA carrier-side SIPREC/SBC media fork at the contact-centre boundary and the internal PBX is the enterprise-scale integration target, not a shipped capability.
Integration target — not shipped
What we returnStructured risk telemetry with the reasoning attached. Your systems and your policy decide and act.
Decision support
04
The verdict path

Mechanism only
Threshold withheld

It does not shout on one suspicious moment.

Several consecutive segments have to agree before the system speaks. That temporal rule is the difference between a detector and something an agent can live with.

AUDIO SEGMENT PARALLEL STREAMS KAN ENSEMBLE TEMPORAL RULE segment spectral prosodic / temporal transient / phase codec interaction glass-box fusion ONE VERDICT PER SEGMENT consecutive agreement → session verdict EXACT COUNT WITHHELD
D2 · Verdict mechanism — the exact segment count is deliberately not published

Each model's contribution to the fused score is exact arithmetic and can be read back per decision, together with the rule chain that produced the alert. That is what makes the verdict defensible in a dispute — not a confidence percentage.

Threshold withheld by design
05
Why it's different

Four properties, each one checkable.

Explainable by construction

A KAN places learnable functions on the edges. Interpretability is an architectural property, not a post-hoc approximation bolted on for the audit.

No load on the agent

No caller enrolment, no friction in the conversation, nothing the agent has to operate. It runs beside the call.

Audit-grade trail

Every signal is timestamped and structured — the documented-diligence record a dispute or a regulator asks for.

It de-risks your stack

It does not replace Voice ID, your fraud engine or your CRM. It adds the orthogonal live-or-generated signal that those systems cannot produce.

06
Why now

External sources
Register-backed

The liability is moving towards the bank.

Not a forecast — the reimbursement line regulators already publish.

£316M reimbursed by UK banks since mandatory APP reimbursement began
SOURCE · UK Payment Systems Regulator
PERIOD · 2024-10-07 → 2026-03-31
Live verified
88% reimbursement rate on in-scope claims — up from 61%
SOURCE · PSR APP dashboard
PRIOR · UK Finance 2024
Live verified
57% of adults worldwide hit by a scam attempt in twelve months
SOURCE · GASA 2025
N · 46,000 adults · 42 markets
Live verified

With DORA and PSR raising the operational bar, an undetected synthetic-voice transfer increasingly lands on the bank — and “we had no way to check” becomes a documented position in the dispute file.

07
Honest scope

Where it helps
Where it does not

We would rather you heard the boundary from us.

In scope

Synthetic-voice detection and vishing-pattern detection on live call audio, delivered through in-app and conference-join capture. Contact-centre SBC and internal-PBX capture is the enterprise integration target.

Shipped · measured
Out of scope, stated plainly

Video deepfakes. End-to-end-encrypted audio after a mid-call switch to an OTT app. Pure DTMF IVR reconnaissance, which is a bank-side telephony control. Some customer-side vectors depend on mobile-operator API availability.

Named gaps

Lossy telephone-band compression is measured separately and is exactly what a first deployment calibrates against your own traffic.

Hear it on your own audio.

The Audio Challenge runs a structured detection test on your own anonymised samples and sends back the report — verdict by verdict, with the reasoning. No integration, no personal data required, no commitment.

Book a 15-minute call Try the Audio Challenge →

Gyula Németh — Founder & CTO, BlueWave AI
gyula.nemeth@bluewaveai.co