Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems

Атаки с использованием состязательных примеров на системы распознавания говорящих в реальном времени: универсальность и устойчивость
Yi Xie, Cong Shi, Zhuohang Li, Jian Liu, Yingying Chen, Bo Yuan
2020-04-09

audio-agnostic universal perturbationover-the-air robustnessroom impulse responsespeaker recognition systemsuniversal adversarial attack
As the popularity of voice user interface (VUI) exploded in recent years, speaker recognition system has emerged as an important medium of identifying a speaker in many security-required applications and services. In this paper, we propose the first real-time, universal, and robust adversarial attack against the state-of-the-art deep neural network (DNN) based speaker recognition system. Through adding an audio-agnostic universal perturbation on arbitrary enrolled speaker's voice input, the DNN-based speaker recognition system would identify the speaker as any target (i.e., adversary-desired) speaker label. In addition, we improve the robustness of our attack by modeling the sound distortions caused by the physical over-the-air propagation through estimating room impulse response (RIR). Experiment using a public dataset of 109 English speakers demonstrates the effectiveness and robustness of our proposed attack with a high attack success rate of over 90%. The attack launching time also achieves a 100× speedup over contemporary non-universal attacks.
1
An audio-agnostic universal perturbation can cause arbitrary enrolled speakers’ voice inputs to be classified as an attacker-selected target speaker.
2
Experiments on a public dataset of 109 English speakers achieve over 90% attack success rate.
3
Introduces the first real-time, universal, and robust adversarial attack against state-of-the-art DNN-based speaker recognition systems.
4
Modeling physical over-the-air distortions using estimated room impulse responses improves attack robustness.
5
The attack launches 100× faster than contemporary non-universal attacks.

DNN-based speaker recognition systems processing arbitrary enrolled speakers’ voice inputs

The effectiveness, universality, real-time performance, and over-the-air robustness of audio-agnostic universal adversarial perturbations for targeted speaker-label misidentification

Publication Details
Publication Date
2020-04-09
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Yi Xie
Cong Shi
Zhuohang Li
Jian Liu
Yingying Chen
Bo Yuan
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%