StreamFake: Word-level Streaming Speech Deepfake Detection

The demo has two parts: real-time voice-cloning attacks and deepfake detection, followed by word-level streaming detection during live communication.

Teaser: Gemini 3.1 Flash TTS · English · Female · Google Meet · 1920 ms streaming chunk size.

01 / Paper overview

Abstract

Speech synthesis technologies have enhanced the convenience of voice interaction while also intensifying security risks such as impersonation and voice fraud. These risks are particularly serious in real-time communication scenarios, such as online meetings, where speech deepfakes may directly interfere with real-time decision-making and cause financial losses. However, existing detection methods typically rely on complete utterances and require the additional deployment of large-scale pretrained feature extractors, limiting their practical applicability.

In this paper, we propose StreamFake, a low-overhead and robust word-level streaming speech deepfake detection system for real-time communication. We observe that streaming automatic speech recognition has been widely deployed in voice interaction and that its encoder provides incremental acoustic representations containing rich discriminative information. Further analysis demonstrates that speech recognition and deepfake detection do not exhibit substantial optimization conflict. Based on these insights, we integrate speech recognition and deepfake detection into a unified streaming framework. Specifically, StreamFake reuses the streaming encoder to extract representations, aggregates word-level acoustic representations according to the temporal boundaries produced by the content decoder, and employs a lightweight detection decoder to synchronously produce authenticity predictions. This design introduces only 0.22M additional parameters without degrading the original transcription performance. Extensive experiments demonstrate that StreamFake achieves state-of-the-art streaming detection performance across different detection latencies. Evaluations involving black-box voice-cloning APIs and real-world communication environments further validate its robustness and practical applicability, consistently outperforming baseline methods.

02 / Method

Word-level streaming deepfake detection

StreamFake architecture showing the streaming speech encoder, incremental content decoder, representation aggregation, and deepfake detection decoder.
Incremental decoding supplies content events; the detector aggregates the corresponding streaming speech representations and produces word-level authenticity decisions.

03 / Attack demo

Real-time voice-cloning attack and detection

StreamFake detects speech from real speakers and real-time voice clones as it unfolds. Samples span Mandarin and English, male and female speakers, and four streaming chunk sizes.

Selected attack sample · Gemini 3.1 Flash TTS

English · Female · 160 ms

Sample 01 of 16
Language
Voice
Streaming chunk size
Preparing selected sample…
A
Microphone inputOriginal captured speech
B
Cloned outputComplete generated speech

Attack model: Gemini 3.1 Flash TTS · Communication platform: XX · Streaming chunk size: 160 ms. Track A contains the original captured speech; Track B contains the complete real-time voice-cloned output.

04 / Detection demo

Streaming detection during live communication

This demo uses Gemini 3.1 Flash TTS for real-time voice cloning and Google Meet as the communication platform to show how different streaming chunk sizes affect detection performance.

Selected detection sample · Gemini 3.1 Flash TTS · Google Meet

English · Female · 160 ms

Sample 01 of 16
Language
Voice
Streaming chunk size
Preparing selected sample…
Real Fake

05 / Generator demos

Commercial black-box voice-cloning APIs

We present streaming detection demos for the top five commercial black-box voice-cloning APIs ranked on the SuperCLUE-TTS benchmark.

Selected generation sample · Google Meet · 160 ms chunk size

Gemini 3.1 Flash TTS Preview · English · Female

Sample 01 of 20
Voice-cloning API
Language
Voice
Preparing selected sample…

06 / Platform demos

Commercial real-time communication platforms

We present detection demos on three real-world communication platforms, including Google Meet, Zoom, and Jitsi, using Gemini 3.1 Flash TTS and a default streaming chunk size of 160 ms.

Selected platform sample · Gemini 3.1 Flash TTS · 160 ms chunk size

Google Meet · English · Female

Sample 01 of 12
Platform
Language
Voice
Preparing selected sample…