兰州理工大学学报 ›› 2026, Vol. 52 ›› Issue (4): 103-110.doi: 10.13295/j.cnki.issn1673-5196.2026.04.012

• 自动化技术与计算机技术 • 上一篇    下一篇

基于聚合Conformer多尺度特征的安多藏语语音识别方法

赵宏*, 扎西草   

  1. 兰州理工大学 计算机与人工智能学院, 甘肃 兰州 730050
  • 收稿日期:2024-09-20 出版日期:2026-08-28 发布日期:2026-09-03
  • 通讯作者: 赵宏(1971-),男,甘肃西和人,博士,教授,博导.Email:zhaoh@lut.edu.cn
  • 基金资助:
    国家自然科学基金(62166025)

Research on Amdo Tibetan speech recognition method based on aggregation conformer multi-scale features

ZHAO Hong, ZHA Xi-cao   

  1. School of Computer Science and Artificial Intelligence, Lanzhou University of Technology, Lanzhou 730050, China
  • Received:2024-09-20 Online:2026-08-28 Published:2026-09-03

摘要: 针对现有语音识别方法在安多藏语语音识别中存在多尺度全局依赖特征及局部细节特征提取不充分,导致无法有效区分同音字和近音字的问题,提出一种基于聚合Conformer多尺度特征的安多藏语语音识别方法.首先,采用速度扰动和频谱增强,增加语音的多样性;其次,编码器采用多层Conformer提取多尺度的语音特征,将多尺度的全局依赖特征和局部细节特征进行聚合,通过注意力统计池化层获取语音特征的加权均值和加权标准差,对局部细节特征以及全局依赖特征分别赋予不同的权重;最后,采用CTC/Attention进行联合解码,实现语音识别.在安多藏语XBMU-AMDO31语音数据集上的实验结果显示,该方法较基准模型Conformer,性能提升36.68%,能够对安多藏语语音中的同音字和近音字进行有效识别.

关键词: Conformer, 安多藏语, 语音识别, 多尺度特征, 特征聚合

Abstract: To address the problem that existing speech recognition methods have insufficient multi-scale global dependent features and local detail feature extraction in Amdo Tibetan speech recognition, leading to the inability to effectively distinguish homophones and near-tone characters, this paper proposes a method based on aggregated Conformer multi-scale features for Amdo Tibetan speech recognition. Firstly, speed perturbation and spectral enhancement are employed to enhance data diversity. Secondly, a multi-layer Conformer encoder is used to extract multi-scale speech features, aggregating global dependency and local detailed representations. An attention-based statistical pooling layer is adopted to obtain weighted means and standard deviations of these features, allowing differential weighting of global and local characteristics. Finally, a joint decoding approach using CTC, and an Attention mechanism is applied to perform speech recognition. Experimental results on the XBMU-AMDO31 Amdo Tibetan speech dataset show that the proposed model outperforms the baseline Conformer model by 36.68% in terms of performance, effectively recognizing homophones and near-tone characters in Amdo Tibetan speech.

Key words: Conformer, Amdo Tibetan, speech recognition, multi-scale features, feature aggregation

中图分类号: