Page 21 - 《软件学报》2026年第5期
P. 21
1900 软件学报 2026 年第 37 卷第 5 期
在序列语境中动态融合知识特征, 从而同时兼顾多层次结构与长距离依赖, 有效缓解了纯序列模型忽略音乐层级
规律且缺乏乐理知识的问题. 在 6 个公开符号音乐理解数据集 Pop1K7、ASAP、POP909、Pianist8、EMOPIA 和
ADL 上的实验结果表明, 该模型在旋律识别、力度预测、作曲家分类、情感分类和流派分类这 5 项符号音乐理
解基准任务上均优于基线模型, 尤其在作曲家分类、情感分类和流派分类任务上分别取得 7.14、4.59 和 3.97 个
百分点的显著性能提升. 消融实验进一步验证了音乐知识特征对深层音乐结构与语义理解的关键作用. 本文提出
的方法为符号音乐理解领域提供了新的研究视角, 也为音乐生成、实时情感识别和多模态音乐分析提供了技术支撑.
References
[1] Salamon J, Gómez E, Ellis DPW, Richard G. Melody extraction from polyphonic music signals: Approaches, applications, and
challenges. IEEE Signal Processing Magazine, 2014, 31(2): 118–134. [doi: 10.1109/MSP.2013.2271648]
[2] Simonetta F, Cancino-Chacón CE, Ntalampiras S, Widmer G. A convolutional approach to melody line identification in symbolic scores.
In: Proc. of the 20th Int’l Society for Music Information Retrieval Conf. Delft: ISMIR, 2019. 924–931. [doi: 10.5281/zenodo.3527965]
[3] Jeong D, Kwon T, Kim Y, Lee K, Nam J. VirtuosoNet: A hierarchical RNN-based system for modeling expressive piano performance. In:
Proc. of the 20th Int’l Society for Music Information Retrieval Conf. Delft: ISMIR, 2019. 908–915. [doi: 10.5281/zenodo.3527962]
[4] Jeong D, Kwon T, Kim Y, Nam J. Graph neural network for music score data and modeling expressive piano performance. In: Proc. of
the 36th Int’l Conf. on Machine Learning. Long Beach: PMLR, 2019. 3060–3070.
[5] Tsai T, Ji K. Composer style classification of piano sheet music images using language model pretraining. In: Proc. of the 21st Int’l
Society for Music Information Retrieval Conf. Montreal: ISMIR, 2020. 176–183. [doi: 10.5281/zenodo.4245398]
[6] Kim S, Lee H, Park S, Lee J, Choi K. Deep composer classification using symbolic representation. arXiv:2010.00823, 2020.
[7] Grekow J, Raś ZW. Detecting emotions in classical music from MIDI files. In: Proc. of the 18th Int’l Symp. on Foundations of Intelligent
Systems. Prague: Springer, 2009. 261–270. [doi: 10.1007/978-3-642-04125-9_29]
[8] Panda R, Malheiro R, Paiva RP. Musical texture and expressivity features for music emotion recognition. In: Proc. of the 19th Int’l
Society for Music Information Retrieval Conf. Paris: ISMIR, 2018. 383–391. [doi: 10.5281/zenodo.1492431]
[9] Good M. MusicXML: An Internet-friendly format for sheet music. 2001. http://michaelgood.info/publications/music/musicxml-an-
internet-friendly-format-for-sheet-music/
[10] Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional Transformers for language understanding. In: Proc.
of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1
(Long and Short Papers). Minneapolis: ACL, 2019. 4171–4186. [doi: 10.18653/v1/N19-1423]
[11] Floridi L, Chiriatti M. GPT-3: Its nature, scope, limits, and consequences. Minds and Machines, 2020, 30(4): 681–694. [doi: 10.1007/
s11023-020-09548-1]
[12] Zeng ML, Tan X, Wang R, Ju ZQ, Qin T, Liu TY. MusicBERT: Symbolic music understanding with large-scale pre-training. In: Proc. of
the 2021 Findings of the Association for Computational Linguistics. ACL, 2021. 791–800. [doi: 10.18653/v1/2021.findings-acl.70]
[13] Chou YH, Chen IC, Ching J, Chang CJ, Yang YH. MidiBERT-Piano: Large-scale pre-training for symbolic music classification tasks.
Journal of Creative Music Systems, 2024, 8(1): 1–19. [doi: 10.5920/jcms.1064]
[14] de Berardinis J, Vamvakaris M, Cangelosi A, Coutinho E. Unveiling the hierarchical structure of music by multi-resolution community
detection. Trans. of the Int’l Society for Music Information Retrieval, 2020, 3(1): 82–97. [doi: 10.5334/tismir.41]
[15] Conklin D, Witten IH. Multiple viewpoint systems for music prediction. Journal of New Music Research, 1995, 24(1): 51–73. [doi: 10.
1080/09298219508570672]
[16] Yang LC, Lerch A. On the evaluation of generative models in music. Neural Computing and Applications, 2020, 32(9): 4773–4784. [doi:
10.1007/s00521-018-3849-7]
[17] Le DVT, Bigo L, Herremans D, Keller M. Natural language processing methods for symbolic music generation and information retrieval:
A survey. ACM Computing Surveys, 2025, 57(7): 175. [doi: 10.1145/3714457]
[18] Mikolov T, Chen K, Corrado G, Dean J. Efficient estimation of word representations in vector space. arXiv:1301.3781, 2013.
[19] Mikolov T, Sutskever I, Chen K, Corrado G, Dean J. Distributed representations of words and phrases and their compositionality. In:
Proc. of the 27th Int’l Conf. on Neural Information Processing Systems. Lake Tahoe: Curran Associates Inc., 2013. 3111–3119.
[20] Li ZC, Gong RH, Chen YN, Su KH. Fine-grained position helps memorizing more, a novel music compound Transformer model with
feature interaction fusion. In: Proc. of the 37th AAAI Conf. on Artificial Intelligence. Washington: AAAI, 2023. 5203–5212. [doi: 10.
1609/aaai.v37i4.25650]
[21] Su JL, Ahmed M, Lu Y, Pan SF, Bo W, Liu YF. RoFormer: Enhanced Transformer with rotary position embedding. Neurocomputing,

