Page 318 - 《软件学报》2026年第2期
P. 318
强敏杰 等: 结合多模态信息的定制化评论生成 797
Ruan Jian Xue Bao/Journal of Software, 2017, 28(3): 708–720 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/5163.
htm [doi: 10.13328/j.cnki.jos.005163]
[26] Chen BY, Huang L, Wang CD, Jing LP. Explicit and implicit feedback based collaborative filtering algorithm. Ruan Jian Xue
Bao/Journal of Software, 2020, 31(3): 794–805 (in Chinese with English abstract). http://www.jos.org.cn/1000-9825/5897.htm [doi: 10.
13328/j.cnki.jos.005897]
[27] Choi J, Hong S, Park N, Cho SB. Blurring-sharpening process models for collaborative filtering. In: Proc. of the 46th Int’l ACM SIGIR
Conf. on Research and Development in Information Retrieval. Taipei: ACM, 2023. 1096–1106. [doi: 10.1145/3539618.3591645]
[28] Sun PJ, Wu L, Zhang K, Su Y, Wang M. An unsupervised aspect-aware recommendation model with explanation text generation. ACM
Trans. on Information Systems (TOIS), 2021, 40(3): 63. [doi: 10.1145/3483611]
[29] Shuai J, Zhang K, Wu L, Sun PJ, Hong RC, Wang M, Li Y. A review-aware graph contrastive learning framework for recommendation.
In: Proc. of the 45th Int’l ACM SIGIR Conf. on Research and Development in Information Retrieval. Madrid: ACM, 2022. 1283–1293.
[doi: 10.1145/3477495.3531927]
[30] Ni JM, McAuley J. Personalized review generation by expanding phrases and attending on aspect-aware representations. In: Proc. of the
56th Annual Meeting of the Association for Computational Linguistics,Vol. 2 (Short Papers). Melbourne: Association for Computational
Linguistics, 2018. 706–711. [doi: 10.18653/v1/P18-2112]
[31] Li P, Tuzhilin A. Towards controllable and personalized review generation. arXiv:1910.03506, 2020.
[32] Sun PJ, Wu L, Zhang K, Fu YJ, Hong RC, Wang M. Dual learning for explainable recommendation: Towards unifying user preference
prediction and review generation. In: Proc. of the 2020 Web Conf. Taipei: ACM, 2020. 837–847. [doi: 10.1145/3366423.3380164]
[33] Xi WD, Huang L, Wang CD, Zheng YY, Lai JH. Deep rating and review neural network for item recommendation. IEEE Trans. on
Neural Networks and Learning Systems, 2022, 33(11): 6726–6736. [doi: 10.1109/TNNLS.2021.3083264]
[34] Ni JM, Lipton ZC, Vikram S, McAuley J. Estimating reactions and recommending products with generative models of reviews. In: Proc.
of the 8th Int’l Joint Conf. on Natural Language Processing, Vol. 1 (Long Papers). Taipei: Asian Federation of Natural Language
Processing, 2017. 783–791.
[35] Tang J, Yang YF, Carton S, Zhang M, Mei QZ. Context-aware natural language generation with recurrent neural networks.
arXiv:1611.09900, 2016.
[36] Peng QY, Liu HT, Xu HY, Yang Q, Shao ML, Wang WJ. Review-LLM: Harnessing large language models for personalized review
generation. arXiv:2407.07487, 2024.
[37] Atrey PK, Hossain MA, El Saddik A, Kankanhalli MS. Multimodal fusion for multimedia analysis: A survey. Multimedia Systems, 2010,
16(6): 345–379. [doi: 10.1007/s00530-010-0182-0]
[38] Bramon R, Boada I, Bardera A, Rodriguez J, Feixas M, Puig J, Sbert M. Multimodal data fusion based on mutual information. IEEE
Trans. on Visualization and Computer Graphics, 2012, 18(9): 1574–1587. [doi: 10.1109/TVCG.2011.280]
[39] Zadeh A, Chen MH, Poria S, Cambria E, Morency LP. Tensor fusion network for multimodal sentiment analysis. arXiv:1707.07250, 2017.
[40] Tsai YHH, Bai SJ, Liang PP, Kolter JZ, Morency LP, Salakhutdinov R. Multimodal Transformer for unaligned multimodal language
sequences. In: Proc. of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for
Computational Linguistics, 2019. 6558–6569. [doi: 10.18653/v1/P19-1656]
[41] Huang J, Tao JH, Liu B, Lian Z, Niu MY. Multimodal Transformer fusion for continuous emotion recognition. In: Proc. of the 2020 IEEE
Int’l Conf. on Acoustics, Speech and Signal Processing (ICASSP 2020). Barcelona: IEEE, 2020. 3507–3511. [doi: 10.1109/
ICASSP40776.2020.9053762]
[42] Yang HZ, Gao XQ, Wu JL, Gan T, Ding N, Jiang FJ, Nie LQ. Self-adaptive context and modal-interaction modeling for multimodal
emotion recognition. In: Findings of the Association for Computational Linguistics: ACL 2023. Toronto: Association for Computational
Linguistics, 2023. 6267–6281. [doi: 10.18653/v1/2023.findings-acl.390]
[43] Zhu JN, Li HR, Liu TS, Zhou Y, Zhang JJ, Zong CQ. MSMO: Multimodal summarization with multimodal output. In: Proc. of the 2018
Conf. on Empirical Methods in Natural Language Processing. Brussels: Association for Computational Linguistics, 2018. 4154–4164.
[doi: 10.18653/v1/D18-1448]
[44] Chung HW, Hou L, Longpre S, et al. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 2024, 25(1):
3381–3433.
[45] Lin TY, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL. Microsoft COCO: Common objects in context. In:
Proc. of the 13th European Conf. on Computer Vision (ECCV 2014). Zurich: Springer Int’l Publishing, 2014. 740–755. [doi: 10.1007/978-
3-319-10602-1_48]
[46] Yan A, He ZK, Li JC, Zhang TY, McAuley J. Personalized showcases: Generating multi-modal explanations for recommendations. In:
Proc. of the 46th Int’l ACM SIGIR Conf. on Research and Development in Information Retrieval. Taipei: ACM, 2023. 2251–2255. [doi:

