针对现有语音识别方法在安多藏语语音识别中存在多尺度全局依赖特征及局部细节特征提取不充分,导致无法有效区分同音字和近音字的问题,提出一种基于聚合Conformer多尺度特征的安多藏语语音识别方法.首先,采用速度扰动和频谱增强,增加语音的多样性;其次,编码器采用多层Conformer提取多尺度的语音特征,将多尺度的全局依赖特征和局部细节特征进行聚合,通过注意力统计池化层获取语音特征的加权均值和加权标准差,对局部细节特征以及全局依赖特征分别赋予不同的权重;最后,采用CTC/Attention进行联合解码,实现语音识别.在安多藏语XBMU-AMDO31语音数据集上的实验结果显示,该方法较基准模型Conformer,性能提升36.68%,能够对安多藏语语音中的同音字和近音字进行有效识别.
To address the problem that existing speech recognition methods have insufficient multi-scale global dependent features and local detail feature extraction in Amdo Tibetan speech recognition, leading to the inability to effectively distinguish homophones and near-tone characters, this paper proposes a method based on aggregated Conformer multi-scale features for Amdo Tibetan speech recognition. Firstly, speed perturbation and spectral enhancement are employed to enhance data diversity. Secondly, a multi-layer Conformer encoder is used to extract multi-scale speech features, aggregating global dependency and local detailed representations. An attention-based statistical pooling layer is adopted to obtain weighted means and standard deviations of these features, allowing differential weighting of global and local characteristics. Finally, a joint decoding approach using CTC, and an Attention mechanism is applied to perform speech recognition. Experimental results on the XBMU-AMDO31 Amdo Tibetan speech dataset show that the proposed model outperforms the baseline Conformer model by 36.68% in terms of performance, effectively recognizing homophones and near-tone characters in Amdo Tibetan speech.
[1] 羊忠旦增.藏语三大方言比较研究[D].北京:中央民族大学,2014:58-69.
[2] Prabhavalkar R,Hori T,Sainath T N,et al.End-to-end speech recognition:a survey[J].IEEE/ACM Transactions on Audio,Speech,and Language Processing,2024,32:325-351.
[3] Qin S,Wang L,Li S,et al.Improving low-resource Tibetan end-to-end ASR by multilingual and multilevel unit modeling[J].EURASIP Journal on Audio,Speech,and Music Processing,2022,2022:2.
[4] 王嘉文,高定国,索朗曲珍.藏语语声识别声学模型建模单元研究[J].应用声学,2025,44(2):405-412.
[5] 边巴旺堆,王希,王君堡.藏语语音识别研究进展综述[J].高原科学研究,2022,6(4):76-84.
[6] Li Q,Mai Q,Wang M,et al.Chinese dialect speech recognition:a comprehensive survey[J].Artificial Intelligence Review,2024,57(2):25.
[7] Zhao Y,Yue J,Xu X,et al.End-to-end-based Tibetan multitask speech recognition[J].IEEE Access,2019,7:162519-162529.
[8] 贡保加.基于MRDCNN_CTC&Transformer的安多藏语语音识别技术研究[D].西宁:青海师范大学,2022:36-42.
[9] 王之杰.基于迁移学习的跨语言藏语拉萨话语音识别研究[D].北京:中央民族大学,2023:27-38.
[10] 宁洁雯.低资源条件下的端到端半监督语音识别研究[D].兰州:西北民族大学,2023:24-30.
[11] 更藏措毛,黄鹤鸣.双向循环神经网络在语音识别中的应用[J].计算机与现代化,2019(10):1-6.
[12] Dong L,Xu S,Xu B.Speech-transformer:a No-recurrence sequence-to-sequence model for speech recognition[C]//2018 IEEE International Conference on Acoustics,Speech and Signal Processing (ICASSP).Calagry,AB:IEEE,2018:5884-5888.
[13] Gulati A,Qin J,Chiu C C,et al.Conformer:convolution-augmented transformer for speech recognition[J/OL].[2024-07-10].http://arxiv.org/abs/2005.08100?context=cs.LG.
[14] 吕坤儒,吴春国,梁艳春,等.融合语言模型的端到端中文语音识别算法[J].电子学报,2021,49(11):2177-2185.
[15] 陈艳,李图雅,马志强,等.基于端到端的蒙古语异形同音词声学建模方法[J].中文信息学报,2022,36(3):27-35.
[16] 付强,徐振平,盛文星,等.结合字节级别字节对编码的端到端中文语音识别方法[J].计算机应用,2025,45(1):318-324.
[17] 赵宏,岳鲁鹏,常兆斌,等.基于多特征I-Vector的说话人识别算法[J].兰州理工大学学报,2021,47(5):93-98.
[18] 姜囡,庞永恒,高爽.基于注意力机制语谱图特征提取的语音识别[J].吉林大学学报 (理学版),2024,62(2):320-330.
[19] 洪青阳,李琳.语音识别:原理与应用[M].北京:电子工业出版社,2020:154-162.
[20] Watanabe S,Hori T,Kim S,et al.Hybrid CTC/Attention architecture for end-to-end speech recognition[J].IEEE Journal of Selected Topics in Signal Processing,2017,11(8):1240-1253.
[21] Yao Z,Wu D,Wang X,et al.Wenet:production oriented str-eaming and non-streaming end-to-end speech recognition toolkit[J/OL].[2024-07-10].http://arxiv.org/abs/2102.01547.
[22] 李运鹏.基于迁移学习的藏语语音识别研究[D].兰州:西北民族大学,2024:32-41.