Page 22 - 《软件学报》2026年第5期
P. 22

黄恒焱 等: 融合音乐知识结构化表征的高精度符号音乐理解                                                    1901


                     2024, 568: 127063. [doi: 10.1016/j.neucom.2023.127063]
                 [22]   Qin Y, Xie HM, Ding SX, Li YJ, Tan BY, Ye MC. Score images as a modality: Enhancing symbolic music understanding through large-
                     scale multimodal pre-training. Sensors, 2024, 24(15): 5017. [doi: 10.3390/s24155017]
                 [23]   Clark K, Khandelwal U, Levy O, Manning CD. What does BERT look at? An analysis of BERT’s attention. In: Proc. of the 2019 ACL
                     Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. Florence: ACL, 2019. 276–286. [doi: 10.18653/v1/W19-
                     4828]
                 [24]   Hsiao WY, Liu JY, Yeh YC, Yang YH. Compound word Transformer: Learning to compose full-song music over dynamic directed
                     hypergraphs. In: Proc. of the 35th AAAI Conf. on Artificial Intelligence. AAAI, 2021, 35(1): 178–186. [doi: 10.1609/aaai.v35i1.16091]
                 [25]   Foscarin F, McLeod A, Rigaux P, Jacquemard F, Sakai M. ASAP: A dataset of aligned scores and performances for piano transcription.
                     In:  Proc.  of  the  21st  Int’l  Society  for  Music  Information  Retrieval  Conf.  Montreal:  ISMIR,  2020.  534–541.  [doi:  10.5281/zenodo.
                     4245490]
                 [26]   Wang ZY, Chen K, Jiang JY, Zhang YY, Xu MR, Dai SQ, Xia G. POP909: A pop-song dataset for music arrangement generation. In:
                     Proc. of the 21st Int’l Society for Music Information Retrieval Conf. Montreal: ISMIR, 2020. 38–45. [doi: 10.5281/zenodo.4245366]
                 [27]   Kong QQ, Li BC, Song XC, Wan Y, Wang YX. High-resolution piano transcription with pedals by regressing onset and offset times.
                     IEEE/ACM Trans. on Audio, Speech, and Language Processing, 2021, 29: 3707–3717. [doi: 10.1109/TASLP.2021.3121991]
                 [28]   Hung HT, Ching J, Doh S, Kim N, Nam J, Yang YH. EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-
                     based music generation. In: Proc. of the 22nd Int’l Society for Music Information Retrieval Conf. ISMIR, 2021. 318–325. [doi: 10.5281/
                     zenodo.5624518]
                 [29]   Ferreira L, Lelis LHS, Whitehead J. Computer-generated music for tabletop role-playing games. In: Proc. of the 16th AAAI Conf. on
                     Artificial Intelligence and Interactive Digital Entertainment. AAAI, 2020. 59–65. [doi: 10.1609/aiide.v16i1.7408]
                 [30]   Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, Davison J, Shleifer S, von Platen P,
                     Ma C, Jernite Y, Plu J, Xu CW, Le Scao T, Gugger S, Drame M, Lhoest Q, Rush A. Transformers: State-of-the-art natural language
                     processing. In: Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations. ACL, 2020.
                     38–45. [doi: 10.18653/v1/2020.emnlp-demos.6]
                 [31]   Loshchilov I, Hutter F. Decoupled weight decay regularization. arXiv:1711.05101, 2019.
                 [32]   Qiu JB, Chen CLP, Zhang T. A novel multi-task learning method for symbolic music emotion recognition. arXiv:2201.05782, 2022.
                 [33]   Liu SK, Xu HG, Xu K. An optimized method for large-scale pre-training in symbolic music. In: Proc. of the 16th IEEE Int’l Conf. on
                     Anti-counterfeiting, Security, and Identification. Xiamen: IEEE, 2022. 105–109. [doi: 10.1109/ASID56930.2022.9995766]
                 [34]   Yang D, Tsai TJ. Composer classification with cross-modal transfer learning and musically-informed augmentation. In: Proc. of the 22nd
                     Int’l Society for Music Information Retrieval Conf. ISMIR, 2021. 802–809.
                 [35]   Fernández-Sotos A, Fernández-Caballero A, Latorre JM. Influence of tempo and rhythmic unit in musical emotion regulation. Frontiers in
                     Computational Neuroscience, 2016, 10: 80. [doi: 10.3389/fncom.2016.00080]

                  附录  A

                  A.1   音符密度偏差

                    在情感分类任务的典型失败案例中             (见图  A1(a)), CNN-Midiformer 出现了系统性的误判现象. 该样本在音高
                 分布、时值均值及和弦复杂度等核心特征上均符合“高兴低激                     (HALV)”类别的统计特征, 但模型却将其错误归类
                 为“低兴低激    (LALV)”. 深入分析发现, 关键问题在于音符密度的异常偏离: 该片段的音符密度达到                       3.2 notes/beat,
                 相比同类   HALV  样本的均值高出约       15%, 并与  LALV  样本的密度分布区间产生显著重叠.
                    这一现象的根本原因在于互补音乐特征提取模块中音长转移矩阵                         (NLTM) 对节奏律动变化的高度敏感性.
                 当  NLTM  检测到异常的高密度节奏结构时, CNN           层能够提炼出与      LALV  训练样本相似的密集节奏特征模式. 在
                 特征融合阶段, 门控机制会在中间层自动调整各通道的权重分配, 显著提升节奏特征通道的影响力, 导致模型对情
                 绪“激发度”维度的判断向        LALV  类别偏移. 消融实验结果进一步证实了这一机制: 在移除互补特征提取模块                       (w/o
                 CNN) 的配置下, 模型性能虽然回落至          RoAR  基线水平, 但不再对密度异常样本表现出过度敏感, 这表明该模块在
                 提升整体性能的同时, 也成为此类误判的直接触发因素.
   17   18   19   20   21   22   23   24   25   26   27