Abstract:As the research on audio adversarial attacks advances, improving the transferability of adversarial audio across different models and ensuring its imperceptibility (that is, highly similar to the original audio in auditory perception) at the same time have become a research hotspot. This study proposes a new method called speak information attack (SIAttack) that can simultaneously improve the imperceptibility and transferability of adversarial audio. Specifically, the core idea of this method is to decouple speaker information from content information in the audio, and then apply small perturbations only to the speaker information, thereby achieving efficient attacks on the speaker recognition system under the premise of keeping the content information unchanged. The experiments on four speaker recognition models and three mainstream commercial APIs show that the audio generated by SIAttack is almost indistinguishable from the original audio, and can mislead all test models with a high success rate. Additionally, the transfer success rate on speaker recognition models can reach up to 100%.