Selected attack sample
Mandarin · Male · 160 ms
The screen recording retains the original speech captured by the microphone. Track B contains the complete output after real-time voice cloning.
An interactive companion to the paper, presenting real-time voice-cloning attacks and word-level deepfake detection as speech unfolds in live communication.
01 / Paper overview
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.
Excepteur sint occaecat cupidatat non proident, sunt in culpa qui officia deserunt mollit anim id est laborum. Integer posuere erat a ante venenatis dapibus posuere velit aliquet. Maecenas faucibus mollis interdum, et malesuada fames ac ante ipsum primis in faucibus.
02 / Method
StreamFake processes speech incrementally. Emission positions from the content decoder align incoming speech representations with word-level evidence, allowing the detection decoder to update its decision without waiting for a complete recording.
03 / Attack demo
Compare the speech captured from the microphone with the complete output produced by the live cloning pipeline. Samples span Mandarin and English, male and female voices, and four streaming horizons.
Selected attack sample
The screen recording retains the original speech captured by the microphone. Track B contains the complete output after real-time voice cloning.
04 / Detection demo
With Gemini as the generation model and Google Meet as the communication platform, StreamFake captures system audio and exposes evolving word-level decisions along the active timeline.
Selected detection sample · Gemini · Google Meet
05 / Generator demos
We present streaming detection examples for speech produced by Gemini, MiMo, MiniMax, Qwen, and Seed. Samples are organized by language, then generation model, then voice gender.
Selected generation sample · Google Meet · 160 ms horizon
06 / Platform demos
Using Gemini and a 160 ms streaming horizon throughout, we show detector outputs for audio transmitted through Google Meet, Zoom, and Jitsi.
Selected platform sample · Gemini · 160 ms horizon