Journal of Lanzhou University of Technology ›› 2026, Vol. 52 ›› Issue (4): 103-110.doi: 10.13295/j.cnki.issn1673-5196.2026.04.012

• Automation Technique and Computer Technology • Previous Articles     Next Articles

Research on Amdo Tibetan speech recognition method based on aggregation conformer multi-scale features

ZHAO Hong, ZHA Xi-cao   

  1. School of Computer Science and Artificial Intelligence, Lanzhou University of Technology, Lanzhou 730050, China
  • Received:2024-09-20 Online:2026-08-28 Published:2026-09-03

Abstract: To address the problem that existing speech recognition methods have insufficient multi-scale global dependent features and local detail feature extraction in Amdo Tibetan speech recognition, leading to the inability to effectively distinguish homophones and near-tone characters, this paper proposes a method based on aggregated Conformer multi-scale features for Amdo Tibetan speech recognition. Firstly, speed perturbation and spectral enhancement are employed to enhance data diversity. Secondly, a multi-layer Conformer encoder is used to extract multi-scale speech features, aggregating global dependency and local detailed representations. An attention-based statistical pooling layer is adopted to obtain weighted means and standard deviations of these features, allowing differential weighting of global and local characteristics. Finally, a joint decoding approach using CTC, and an Attention mechanism is applied to perform speech recognition. Experimental results on the XBMU-AMDO31 Amdo Tibetan speech dataset show that the proposed model outperforms the baseline Conformer model by 36.68% in terms of performance, effectively recognizing homophones and near-tone characters in Amdo Tibetan speech.

Key words: Conformer, Amdo Tibetan, speech recognition, multi-scale features, feature aggregation

CLC Number: