Innovation Series: Advanced Science (ISSN 2938-9933, CNKI Indexed)

Volume 3 · Issue 8 (2026)
122
views
DOI number:
10.66521/2938-9933-2026083101

Dual-Encoder ViT–SAM Heterogeneous Feature Fusion for Surface Defect Detection of Mechanical Parts

 

Jianqing Xin1,*, Yuhao Kong2, Minwen Guo3

1 Harbin Institute of Technology, Weihai 264200, China

2 Northwest A&F University, Xianyang 712100, China

3 Chengdu College of University of Electronic Science and Technology of China, Chengdu 611731, China

Corresponding Author: Jianqing Xin (1309515895@qq.com)

 

Abstract: Surface defect detection of mechanical parts is challenging because defects are often small, thin, low contrast, and visually similar to manufacturing textures or reflection artifacts [1–4]. This paper presents a dual-encoder framework that combines a lightweight ResNet-based local branch with the Vision Transformer image encoder of the Segment Anything Model (SAM) [10], [15]. A heterogeneous feature fusion module (HFM) aligns local texture features and global semantic features through cross-attention, channel recalibration, and residual preservation. Confidence-based sample filtering, uncertainty-aware multi-task weighting, and staged fine-tuning are used to reduce the effects of label noise and target-domain mismatch [22], [24], [25]. On a production-line dataset containing seven defect categories, the proposed model achieves 91.1% mAP, 88.3% mIoU, and 91.6% recall. It also reaches 58 FPS in a separate NVIDIA GTX 1070 throughput benchmark. These results indicate that the proposed fusion strategy improves detection and segmentation while retaining practical inference efficiency.

 

Keywords: Surface defect detection; Segment Anything Model; Heterogeneous feature fusion; multi-task learning; Industrial vision

 

References

[1]
J. Liu, G. Xie, J. Wang, S. Li, C. Wang, F. Zheng, and Y. Jin, “Deep industrial image anomaly detection: A survey,” Machine Intelligence Research, vol. 21, no. 1, pp. 104–135, 2024. https://doi.org/10.1007/ s11633-023-1459-z
[2]
Z. Li, Y. Yan, X. Wang, Y. Ge, and L. Meng, “A survey of deep learning for industrial visual anomaly detection,” Artificial Intelligence Review, vol. 58, no. 9, Art. no. 279, 2025. https://doi.org/10.1007/s10462-025-11287-7
[3]
X. Tao, X. Gong, X. Zhang, S. Yan, and C. Adak, “Deep learning for unsupervised anomaly localization in industrial images: A survey,” IEEE Transactions on Instrumentation and Measurement, vol. 71, pp. 1–21, 2022. https://doi.org/10.1109/TIM.2022.3196436
[4]
Z. Zhang, Z. Zhao, X. Zhang, C. Sun, and X. Chen, “Industrial anomaly detection with domain shift: A real-world dataset and masked multi-scale reconstruction,” Computers in Industry, vol. 151, Art. no. 103990, 2023. https://doi.org/10.1016/j.compind.2023.103990
[5]
C. Wang, W. Zhu, B.-B. Gao, Z. Gan, J. Zhang, Z. Gu, S. Qian, M. Chen, and L. Ma, “Real-IAD: A real-world multi-view dataset for benchmarking versatile industrial anomaly detection,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2024, pp. 22883–22892. https://doi.org/10.1109/CVPR52733.2024.02159
[6]
P. Bergmann, M. Fauser, D. Sattlegger, and C. Steger, “MVTec AD—A comprehensive real-world dataset for unsupervised anomaly detection,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2019, pp. 9584–9592. https://doi.org/10.1109/CVPR.2019.00982
[7]
S. Li, C. Wu, and N. Xiong, “Hybrid architecture based on CNN and Transformer for strip steel surface defect classification,” Electronics, vol. 11, no. 8, Art. no. 1200, 2022. https://doi.org/10.3390/ electronics11081200
[8]
L. Zhao, Y. Zheng, T. Peng, and E. Zheng, “Metal surface defect detection based on a Transformer with multi-scale mask feature fusion,” Sensors, vol. 23, no. 23, Art. no. 9381, 2023. https://doi.org/10.3390/s23239381
[9]
B. Tang, Z.-K. Song, W. Sun, and X.-D. Wang, “An end-to-end steel surface defect detection approach via Swin Transformer,” IET Image Processing, vol. 17, no. 5, pp. 1334–1345, 2023. https://doi.org/10.1049/ipr2.12715
[10]
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778. https://doi.org/10.1109/CVPR. 2016.90
[11]
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015, LNCS 9351, 2015, pp. 234–241. https://doi.org/10.1007/978-3-319-24574-4 28
[12]
T.-Y. Lin, P. Doll´ar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2117–2125. https://doi.org/10.1109/CVPR.2017.106
[13]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems 30, 2017, pp. 5998–6008. https://papers.nips.cc/paper/7181-attention-is-all-you-need
[14]
A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations (ICLR), 2021.
[15]
A. Kirillov, E. Mintun, N. Ravi, et al., “Segment Anything,” in Proc. IEEE/CVF International Conf. Computer Vision (ICCV), 2023, pp. 4015–4026.
[16]
C. Zhang, D. Han, Y. Qiao, J. U. Kim, S.-H. Bae, S. Lee, and C. S. Hong, “Faster Segment Anything: Towards lightweight SAM for mobile applications,”
[17]
W. Ji, J. Li, Q. Bi, T. Liu, W. Li, and L. Cheng, “Segment Anything is not always perfect: An investigation of SAM on different real-world applications,” Machine Intelligence Research, vol. 21, no. 4, pp. 617–630, 2024. https://doi.org/10.1007/s11633-023-1385-0
[18]
B. Hu, B. Gao, C. Tan, T. Wu, and S. Z. Li, “Segment Anything in defect detection,”
[19]
K. Song, W. Cui, H. Yu, X. Li, and Y. Yan, “SAM era: Can it segment any industrial surface defects?” Computers, Materials & Continua, vol. 78, no. 3, pp. 3953–3969, 2024. https://doi.org/10.32604/cmc.2024.048451
[20]
J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation networks,” in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7132–7141. https://doi.org/10.1109/CVPR.2018.00745
[21]
Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: Fully convolutional onestage object detection,” in Proc. IEEE/CVF International Conf. Computer Vision (ICCV), 2019, pp. 9626–9635. https://doi.org/10.1109/ICCV.2019.00972
[22]
A. Kendall, Y. Gal, and R. Cipolla, “Multi-task learning using uncertainty to weigh losses for scene geometry and semantics,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2018, pp. 7482–7491. https://doi.org/10.1109/CVPR.2018.00781
[23]
F. Milletari, N. Navab, and S.-A. Ahmadi, “V-Net: Fully convolutional neural networks for volumetric medical image segmentation,” in Proc. Fourth International Conf. 3D Vision (3DV), 2016, pp. 565–571. https://doi.org/10.1109/3DV.2016.79
[24]
B. Han, Q. Yao, X. Yu, et al., “Co-teaching: Robust training of deep neural networks with extremely noisy labels,” in Advances in Neural Information Processing Systems 31, 2018, pp. 8527–8537. NeurIPS paper page.
[25]
Z. Zhang and M. R. Sabuncu, “Generalized cross entropy loss for training deep neural networks with noisy labels,” in Advances in Neural Information Processing Systems 31, 2018, pp. 8778–8788. NeurIPS paper page.
[26]
L. van der Maaten and G. Hinton, “Visualizing data using t-SNE,” Journal of Machine Learning Research, vol. 9, pp. 2579–2605, 2008. https://www.jmlr.org/papers/v9/vandermaaten08a.html
[27]
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” in Proc. IEEE International Conf. Computer Vision (ICCV), 2017, pp. 618–626. https://doi.org/10.1109/ICCV.2017.74
Download PDF
Innovation Series

Innovation Series is an academic publisher publishing journals and books covering a wide range of academic disciplines.

Contact

Francesc Boix i Campo, 7

08038 Barcelona, Spain