Page 171 - 《软件学报》2026年第4期
P. 171
1612 软件学报 2026 年第 37 卷第 4 期
[30] Zhou BL, Sun YY, Bau D, Torralba A. Interpretable basis decomposition for visual explanation. In: Proc. of the 15th European Conf. on
Computer Vision. Munich: Springer, 2018. 122–138. [doi: 10.1007/978-3-030-01237-3_8]
[31] Goyal Y, Feder A, Shalit U, Kim B. Explaining classifiers with causal concept effect (CaCE). arXiv:1907.07165, 2020.
[32] Wu ZX, D’Oosterlinck K, Geiger A, Zur A, Potts C. Causal proxy models for concept-based model explanations. In: Proc. of the 40th Int’l
Conf. on Machine Learning. Honolulu: JMLR.org, 2023. 37313–37334.
[33] Zhou BL, Khosla A, Lapedriza A, Oliva A, Torralba A. Object detectors emerge in deep scene CNNs. arXiv:1412.6856, 2015.
[34] Bau D, Zhou BL, Khosla A, Oliva A, Torralba A. Network dissection: Quantifying interpretability of deep visual representations. In:
Proc. of the 2017 IEEE Conf. on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017. 3319–3327. [doi: 10.1109/CVPR.
2017.354]
[35] Fong R, Vedaldi A. Net2Vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks. In: Proc. of the
2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Salt Lake City: IEEE, 2018. 8730–8738. [doi: 10.1109/CVPR.2018.
00910]
[36] Ghorbani A, Wexler J, Zou J, Kim B. Towards automatic concept-based explanations. In: Proc. of the 33rd Int’l Conf. on Neural
Information Processing Systems. Vancouver: Curran Associates Inc., 2019. 9277–9286.
[37] Yeh CK, Kim B, Arik SÖ, Li CL, Pfister T, Ravikumar P. On completeness-aware concept-based explanations in deep neural networks.
In: Proc. of the 34th Int’l Conf. on Neural Information Processing Systems. Vancouver: Curran Associates Inc., 2020. 20554–20565.
[38] Zhang RH, Madumal P, Miller T, Ehinger KA, Rubinstein BIP. Invertible concept-based explanations for CNN models with non-negative
concept activation vectors. In: Proc. of the 35th AAAI Conf. on Artificial Intelligence. AAAI, 2021. 11682–11690. [doi: 10.1609/aaai.
v35i13.17389]
[39] Vielhaben J, Blücher S, Strodthoff N. Multi-dimensional concept discovery (MCD): A unifying framework with completeness guarantees.
arXiv:2301.11911, 2023.
[40] Fel T, Picard A, Bethune L, Boissin T, Vigouroux D, Colin J, Cadéne R, Serre T. CRAFT: Concept recursive activation factorization for
explainability. In: Proc. of the 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 2711–2721.
[doi: 10.1109/CVPR52729.2023.00266]
[41] Yang LJ, Wang JQ, Jing LP, Yu J. Semantic representation learning of convolutional neural network based on tensor computation.
Chinese Journal of Computers, 2023, 46(3): 568–578 (in Chinese with English abstract). [doi: 10.11897/SP.J.1016.2023.00568]
[42] Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature
Machine Intelligence, 2019, 1(5): 206–215. [doi: 10.1038/s42256-019-0048-x]
[43] Zhai JH, Zhang SF, Chen JF, He Q. Autoencoder and its various variants. In: Proc. of the 2018 IEEE Int’l Conf. on Systems, Man, and
Cybernetics (SMC). Miyazaki: IEEE, 2018. 415–419. [doi: 10.1109/SMC.2018.00080]
[44] Kingma DP, Welling M. Auto-encoding variational Bayes. arXiv:1312.6114, 2022.
[45] Koh PW, Nguyen T, Tang YS, Mussmann S, Pierson E, Kim B, Liang P. Concept bottleneck models. In: Proc. of the 37th Int’l Conf. on
Machine Learning. JMLR.org, 2020. 5338–5348.
[46] Kim E, Jung D, Park S, Kim S, Yoon S. Probabilistic concept bottleneck models. In: Proc. of the 40th Int’l Conf. on Machine Learning.
Honolulu: JMLR.org, 2023. 16521–16540.
[47] Chen Z, Bei YJ, Rudin C. Concept whitening for interpretable image recognition. Nature Machine Intelligence, 2020, 2(12): 772–782.
[doi: 10.1038/s42256-020-00265-z]
[48] Zarlenga ME, Barbiero P, Ciravegna G, Marra G, Giannini F, Diligenti M, Shams Z, Precioso F, Melacci S, Weller A, Lio P, Jamnik M.
Concept embedding models: Beyond the accuracy-explainability trade-off. In: Proc. of the 36th Int’l Conf. on Neural Information
Processing Systems. New Orleans: Curran Associates Inc., 2022. 21400–21413.
[49] Yuksekgonul M, Wang M, Zou J. Post-hoc concept bottleneck models. arXiv:2205.15480, 2023.
[50] Rigotti M, Miksovic C, Giurgiu I, Gschwind T, Scotton P. Attention-based interpretability with concept Transformers. In: Proc. of the
10th Int’l Conf. on Learning Representations. OpenReview.net, 2021.
[51] Zhang QS, Wu YN, Zhu SC. Interpretable convolutional neural networks. In: Proc. of the 2018 IEEE/CVF Conf. on Computer Vision and
Pattern Recognition. Salt Lake City: IEEE, 2018. 8827–8836. [doi: 10.1109/CVPR.2018.00920]
[52] Alvarez-Melis D, Jaakkola TS. Towards robust interpretability with self-explaining neural networks. In: Proc. of the 32nd Int’l Conf. on
Neural Information Processing Systems. Montréal: Curran Associates Inc., 2018. 7786–7795.
[53] Wang BW, Li LZ, Nakashima Y, Nagahara H. Learning bottleneck concepts in image classification. In: Proc. of the 2023 IEEE/CVF
Conf. on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023. 10962–10971. [doi: 10.1109/CVPR52729.2023.01055]
[54] Rajagopal D, Balachandran V, Hovy EH, Tsvetkov Y. SelfExplain: A self-explaining architecture for neural text classifiers. In: Proc. of

