A real-time deepfake and vishing detection layer for the bank voice channel. Passive on the line, explainable in the verdict, evidence-grade in the audit trail.
Every vector below maps to a published case or a regulator-recorded trend. This is the fraud map of the bank voice channel — not a threat landscape slide.
Voice biometrics asks who is this. When the clone is good enough, the answer is wrong and confident. VoiceGuard runs the orthogonal question on the same audio.
A four-stream KAN ensemble scores every inbound caller automatically — no enrolment, no voiceprint, nothing for the customer to do. The training corpus spans 33 synthesis engines.
Speech recognition plus a fraud-keyword engine runs on both legs, caller and agent, and fires even when there is no deepfake at all: safe-account pretext, OTP coaching, channel-hopping, or an agent policy violation.
The two pillars are not averaged into one number. The system returns which of them fired:
VISHING_SCAMDEEPFAKE_DETECTEDDEEPFAKE_VISHING
If VoiceGuard fails, the call does not. That is an architectural property, not an operational promise.
Several consecutive segments have to agree before the system speaks. That temporal rule is the difference between a detector and something an agent can live with.
Each model's contribution to the fused score is exact arithmetic and can be read back per decision, together with the rule chain that produced the alert. That is what makes the verdict defensible in a dispute — not a confidence percentage.
A KAN places learnable functions on the edges. Interpretability is an architectural property, not a post-hoc approximation bolted on for the audit.
No caller enrolment, no friction in the conversation, nothing the agent has to operate. It runs beside the call.
Every signal is timestamped and structured — the documented-diligence record a dispute or a regulator asks for.
It does not replace Voice ID, your fraud engine or your CRM. It adds the orthogonal live-or-generated signal that those systems cannot produce.
Not a forecast — the reimbursement line regulators already publish.
With DORA and PSR raising the operational bar, an undetected synthetic-voice transfer increasingly lands on the bank — and “we had no way to check” becomes a documented position in the dispute file.
Synthetic-voice detection and vishing-pattern detection on live call audio, delivered through in-app and conference-join capture. Contact-centre SBC and internal-PBX capture is the enterprise integration target.
Video deepfakes. End-to-end-encrypted audio after a mid-call switch to an OTT app. Pure DTMF IVR reconnaissance, which is a bank-side telephony control. Some customer-side vectors depend on mobile-operator API availability.
Lossy telephone-band compression is measured separately and is exactly what a first deployment calibrates against your own traffic.
The Audio Challenge runs a structured detection test on your own anonymised samples and sends back the report — verdict by verdict, with the reasoning. No integration, no personal data required, no commitment.
Gyula Németh
— Founder & CTO, BlueWave AI
gyula.nemeth@bluewaveai.co