LDIAAD: A Lightweight Domain-Invariant Acoustic Anomaly Detector for Industrial Machine Condition Monitoring Under Operating Condition Shifts
Abstract
Acoustic monitoring offered a low-cost, non-invasive approach to industrial predictive maintenance, but anomalous sound detection (ASD) models trained under one operating condition were shown in prior work to lose substantial accuracy when operating conditions shifted, and lightweight models suitable for edge deployment were rarely evaluated jointly for cross-domain generalization and computational efficiency. This study proposed LDIAAD (Lightweight Domain Invariant Acoustic Anomaly Detector), a MobileNetV3-Small encoder trained on a multi-resolution log-Mel representation with a self-supervised auxiliary classification objective, an explicit domain-alignment loss, knowledge distillation from a larger EfficientNet-B0 teacher, and a normal-prototype compactness regularizer. Experiments followed the official DCASE 2022 Task 2 protocol on the MIMII DG domain-generalization benchmark. A three-way pilot selected MMD as the best-performing alignment loss (overall AUC 0.608, versus 0.576 for CORAL and 0.529 for DANN). Across ten independent random seeds, LDIAAD reached a mean overall AUC of 0.6001 (± 0.0234), significantly higher than the identical backbone trained without any domain-generalization component (0.5536 ± 0.0332; paired t-test p = 0.0029, Cohen's d = 1.28, surviving Holm-Bonferroni correction). LDIAAD was also more robust than the plain backbone under all nine tested audio degradations. A much smaller custom lightweight CNN baseline, however, significantly outperformed LDIAAD on raw AUC (0.6871 ± 0.0149; p < 0.001), a finding reported plainly rather than downplayed. LDIAAD is proposed as the preferable choice where consistent cross-domain behavior matters more than peak same-condition accuracy.
Keywords
Full Text:
PDFReferences
Abbasi, S., Famouri, M., Shafiee, M.J., & Wong, A. (2021). OutlierNets: Highly Compact Deep Autoencoder Network Architectures for On-Device Acoustic Anomaly Detection. https://arxiv.org/abs/2104.00528
Bai, J., Chen, J., Wang, M., Ayub, M.S., & Yan, Q. (2023). SSDPT: Self-Supervised Dual-Path Transformer for Anomalous Sound Detection in Machine Condition Monitoring. Digital Signal Processing, 135, 103939. https://doi.org/10.1016/j.dsp.2023.103939
Dohi, K., Nishida, T., Purohit, H., Tanabe, R., Endo, T., Yamamoto, M., Nikaido, Y., & Kawaguchi, Y. (2022). MIMII DG: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection for Domain Generalization Task. https://doi.org/10.5281/zenodo.6529888
Dohi, K., Imoto, K., Harada, N., Niizumi, D., Koizumi, Y., Nishida, T., Purohit, H., Endo, T., Yamamoto, M., & Kawaguchi, Y. (2022). Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques. https://arxiv.org/abs/2206.05876
Ganin, Y., Ustinova, E., Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., Marchand, M., & Lempitsky, V. (2016). Domain-Adversarial Training of Neural Networks. Journal of Machine Learning Research, 17(59), 1--35. https://jmlr.org/papers/v17/15-239.html
Gretton, A., Borgwardt, K.M., Rasch, M.J., Schölkopf, B., & Smola, A. (2012). A Kernel Two-Sample Test. Journal of Machine Learning Research, 13(25), 723--773. https://jmlr.org/papers/v13/gretton12a.html
Guan, J., Xiao, F., Liu, Y., Zhu, Q., & Wang, W. (2023). Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://arxiv.org/abs/2304.03588
Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. arXiv preprint arXiv:1503.02531. https://arxiv.org/abs/1503.02531
Howard, A., Sandler, M., Chu, G., Chen, L., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., Le, Q.V., & Adam, H. (2019). Searching for MobileNetV3. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1314--1324. https://openaccess.thecvf.com/content_ICCV_2019/html/Howard_Searching_for_MobileNetV3_ICCV_2019_paper.html
Jena, S., Pulkit, A., Singh, K., Banerjee, A., Joshi, S., Ganesh, A., Singh, D., & Bhavsar, A. (2024). Unified Anomaly Detection Methods on Edge Device Using Knowledge Distillation and Quantization. arXiv preprint arXiv:2407.02968. https://arxiv.org/abs/2407.02968 [Vision-domain (image) anomaly detection benchmark; cited here only as a general knowledge-distillation-plus-quantization edge-deployment precedent, not as an audio-domain result.]
Koizumi, Y., Kawaguchi, Y., Imoto, K., Nakamura, T., Nikaido, Y., Tanabe, R., Purohit, H., Suefusa, K., Endo, T., Yasuda, M., & Harada, N. (2020). Description and Discussion on DCASE2020 Challenge Task2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring. arXiv preprint arXiv:2006.05822. https://arxiv.org/abs/2006.05822
Nishida, T., Harada, N., Niizumi, D., Albertini, D., Sannino, R., Pradolini, S., Augusti, F., Imoto, K., Dohi, K., Purohit, H., Endo, T., & Kawaguchi, Y. (2023). Description and Discussion on DCASE 2023 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring. https://arxiv.org/abs/2305.07828
Nishida, T., Harada, N., Niizumi, D., Albertini, D., Sannino, R., Pradolini, S., Augusti, F., Imoto, K., Dohi, K., Purohit, H., Endo, T., & Kawaguchi, Y. (2025). Description and Discussion on DCASE 2025 Challenge Task 2: First-Shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring. https://arxiv.org/abs/2506.10097
Purohit, H., Tanabe, R., Ichige, K., Endo, T., Nikaido, Y., Suefusa, K., & Kawaguchi, Y. (2019). MIMII Dataset: Sound Dataset for Malfunctioning Industrial Machine Investigation and Inspection. Proceedings of the Detection and Classification of Acoustic Scenes and Events 2019 Workshop (DCASE2019). https://arxiv.org/abs/1909.09347
Saengthong, P., & Shinozaki, T. (2024). Deep Generic Representations for Domain-Generalized Anomalous Sound Detection. https://arxiv.org/abs/2409.05035
Sun, B., & Saenko, K. (2016). Deep CORAL: Correlation Alignment for Deep Domain Adaptation. Computer Vision -- ECCV 2016 Workshops. https://doi.org/10.1007/978-3-319-49409-8_35
Tan, M., & Le, Q. (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. Proceedings of the 36th International Conference on Machine Learning (ICML), 6105--6114. https://proceedings.mlr.press/v97/tan19a.html
Wilkinghoff, K. (2021). Sub-Cluster AdaCos: Learning Representations for Anomalous Sound Detection. Proceedings of the International Joint Conference on Neural Networks (IJCNN). https://dcase.community/documents/challenge2021/technical_reports/DCASE2021_Wilkinghoff_31_t2.pdf
Wilkinghoff, K. (2024). Self-Supervised Learning for Anomalous Sound Detection. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 276--280. https://arxiv.org/abs/2312.09578
Ye, T., Peng, T., & Yang, L. (2025). Review on Sound-Based Industrial Predictive Maintenance: From Feature Engineering to Deep Learning. Mathematics, 13(11), 1724. https://doi.org/10.3390/math13111724
Yeo, J.J.S., Tan, E., Bai, J., Pekski, S., & Gan, W. (2024). Data Efficient Acoustic Scene Classification Using Teacher-Informed Confusing Class Instruction. https://arxiv.org/abs/2409.11964 [DCASE 2024 Task 1 (acoustic scene classification, not anomaly detection); cited for its teacher-student knowledge-distillation design for lightweight audio models, which directly parallels LDIAAD's teacher-student distillation (Section 4.3).]
DOI: https://doi.org/10.30596/jcositte.v7i2.31674
Refbacks
- There are currently no refbacks.




.png)

