This is the second part of our input privacy series, focusing on the implications of established adversarial attacks on federated learning. It's worth noting that while extensive literature exists regarding adversarial attacks on centralized models, reassessing the attack scenario within federated environments is paramount. As explored in our previous post, the unique dynamics of federated training and the aggregator's role significantly alter the attack's repercussions.
This time we explore backdoor attacks and their effects on the global model. We assess the entire backdoor mechanism, closely examining its structural intricacies and impact. These new insights highlight the severity and implications of backdoor attacks on federated model training.
To get up to speed be sure to check out our previous blog posts:
The Backdoor Attack: A Stealthy Threat to Federated Learning
Among adversarial attacks, the backdoor attack stands out as a particularly insidious threat to federated model training. In a backdoor attack, a malicious actor introduces a hidden trigger or pattern into the training data, which the model learns to recognize. This trigger has no impact on normal data but can be exploited to manipulate the model's predictions when the trigger is present.
In the context of federated learning, a backdoor attack can have severe consequences. Since models are trained across distributed devices or servers, a compromised model with a backdoor can spread the vulnerability across the federated network. This not only compromises the model's accuracy and reliability but also raises concerns about the integrity of data across participating entities.
Experiment setup
For the experiments, we have used the MNIST handwritten dataset comprising a total of 70,000 images, divided by default into 60,000 images for training and 10,000 images for testing. These tests are then subdivided into 10 partitions, each assigned to an individual client. The data partitions are balanced using IID settings. The model was trained using a commonly known neural network classification PyTorch model architecture.
- MNIST Dataset: https://paperswithcode.com/dataset/mnist
- PyTorch Implementation: https://www.kaggle.com/code/faduregis/mnist-digit-classification-in-pytorch
The attack described is a targeted backdoor attack, in which images of the digit "6" are marked with a backdoor indicated by the symbol "+". These altered images belong only to the malicious clients.
For the experiments, we selected the following three settings:
- Baseline: This includes all honest clients and serves as the foundation for training the federated model and assessing its accuracy.
- 10% Malicious Clients: In this setting, 10% of the clients in the federation are malicious, and all the images of one specific digit “6” are marked with the backdoor “+”.
- 20% Malicious Clients: In this setting, 20% of the clients in the federation are malicious, and like in the previous setting, all images of one specific digit “6” are marked with the backdoor “+”.
We worked with a total of ten clients, of which 10% and 20% were malicious clients within the federation. The training dataset for malicious clients contains all images of the digit "6" with the backdoor, whereas normal clients have all clean images in the training dataset.
Results and discussion
The training and validation processes within federated settings show that backdoor attacks are not easily detectable through conventional metrics. Plots of training accuracy, training loss, validation accuracy, and validation loss showcased borderline differences between the honest baseline and the federated networks with 10% and 20% malicious clients.
However, the impact on the test dataset was severe. We evaluated five different scenarios:
- Test dataset with all clean images.
- Test dataset with all images having the exact defined structure of the backdoor "+".
- Variations in the structure of the backdoor:
- 3a. Extended horizontal and vertical lines.
- 3b. A single extended horizontal line.
- 3c. A single extended vertical line.
Key Findings:
- Honest Setup: The model is unaware of the backdoor; results across all datasets align with expectations, as the backdoor is essentially treated as noise.
- 10% Malicious Setup: The model behaves normally with clean images. However, when the backdoor is introduced, there is a drastic negative impact. Predictions are successfully pulled toward label “6”, and there is an increase in incorrect counts for label “4” due to visual similarity to “6”. This effect persists even with partial structures of the backdoor.
- 20% Malicious Setup: This resulted in an even more aggravated negative impact. Almost all test samples were classified as the targeted label (digit 6), even when structural variations were introduced.
Remedies of Backdoor attacks
The experiments underscore the threat posed even by a small number of malicious clients. Potential mitigation strategies from centralized training literature include:
- Identifying discernible traces in the latent or feature space.
- Utilizing neural attention distillation to eliminate triggers.
Scaleout Systems specializes in privacy-preserving machine learning and supports practitioners in adopting best practices to mitigate these technical challenges.
Summary
Backdoor attacks pose a considerable challenge because they are difficult to detect during traditional federated training. Both the exact structure and variations of the backdoor can significantly harm model performance while remaining concealed. While we used a basic backdoor for these experiments, more sophisticated techniques can create triggers that are visually indistinguishable from clean images, potentially increasing the impact.
Acknowledgments
The results presented in this post are derived from the project report "Evaluating Model Poisoning Attacks in Federated Machine Learning" prepared for the Computational Science project course at Uppsala University, Sweden.
References:
- [1] Doan, Khoa, Yingjie Lao, and Ping Li. "Backdoor attack with imperceptible input and latent modification." Advances in Neural Information Processing Systems 34 (2021): 18944-18957.
- [2] Li, Yige, et al. "Neural attention distillation: Erasing backdoor triggers from deep neural networks." arXiv preprint arXiv:2101.05930 (2021).
- [3] Wang, Bolun, et al. "Neural cleanse: Identifying and mitigating backdoor attacks in neural networks." 2019 IEEE Symposium on Security and Privacy (SP). IEEE, 2019.
- [4] Guo, Wei, Benedetta Tondi, and Mauro Barni. "An overview of backdoor attacks against deep neural networks and possible defences." IEEE Open Journal of Signal Processing 3 (2022): 261-287.