StreamFake: Streaming Speech Deepfake Attack and Detection

An interactive companion to the paper, presenting real-time voice-cloning attacks and word-level deepfake detection as speech unfolds in live communication.

Teaser: Gemini · Mandarin · Google Meet · 160 ms streaming horizon.

01 / Paper overview

Abstract

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum. Integer posuere erat a ante venenatis dapibus posuere velit aliquet. Maecenas faucibus mollis interdum, et malesuada fames ac ante ipsum primis in faucibus.

02 / Method

Streaming word-level deepfake detection

StreamFake processes speech incrementally. Emission positions from the content decoder align incoming speech representations with word-level evidence, allowing the detection decoder to update its decision without waiting for a complete recording.

StreamFake architecture from slide 4 of the presentation, showing the streaming speech encoder, incremental content decoder, representation aggregation, and deepfake detection decoder.
Incremental decoding supplies content events; the detector aggregates the corresponding streaming speech representations and produces word-level authenticity decisions.

03 / Attack demo

Real-time voice-cloning attack

Compare the speech captured from the microphone with the complete output produced by the live cloning pipeline. Samples span Mandarin and English, male and female voices, and four streaming horizons.

Selected attack sample

Mandarin · Male · 160 ms

Sample 01 of 16
Language
Voice
Streaming horizon
Preparing selected sample…
A
Microphone inputOriginal captured speech
B
Cloned outputComplete generated speech

The screen recording retains the original speech captured by the microphone. Track B contains the complete output after real-time voice cloning.

04 / Detection demo

Streaming detection during live communication

With Gemini as the generation model and Google Meet as the communication platform, StreamFake captures system audio and exposes evolving word-level decisions along the active timeline.

Selected detection sample · Gemini · Google Meet

Mandarin · Male · 160 ms

Sample 01 of 16
Language
Voice
Streaming horizon
Preparing selected sample…
Real Fake

05 / Generator demos

Across five speech generation models

We present streaming detection examples for speech produced by Gemini, MiMo, MiniMax, Qwen, and Seed. Samples are organized by language, then generation model, then voice gender.

Selected generation sample · Google Meet · 160 ms horizon

Gemini 3.1 Flash TTS Preview · Mandarin · Male

Sample 01 of 20
Generation model
Language
Voice
Preparing selected sample…

06 / Platform demos

Across real-time communication platforms

Using Gemini and a 160 ms streaming horizon throughout, we show detector outputs for audio transmitted through Google Meet, Zoom, and Jitsi.

Selected platform sample · Gemini · 160 ms horizon

Google Meet · Mandarin · Male

Sample 01 of 12
Platform
Language
Voice
Preparing selected sample…