<?xml version="1.1" encoding="utf-8"?>
<article xsi:noNamespaceSchemaLocation="http://jats.nlm.nih.gov/publishing/1.1/xsd/JATS-journalpublishing1-mathml3.xsd" dtd-version="1.1" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"><front><journal-meta><journal-id journal-id-type="publisher-id">JERA</journal-id><journal-title-group><journal-title>Journal of Electronic Research and Application</journal-title></journal-title-group><issn>2208-3502</issn><eissn>2208-3510</eissn><publisher><publisher-name>Bio-Byword Scientific Publishing Pty. Ltd.</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.26689/jera.v10i3.14660</article-id><article-categories><subj-group subj-group-type="heading"><subject>Article</subject></subj-group></article-categories><title>A Tibetan Speaker Verification Method Based on the Improved MFA-NConformer Model</title><url>https://artdesignp.com/journal/JERA/10/3/10.26689/jera.v10i3.14660</url><author>GongYitong,ChenYuting</author><pub-date pub-type="publication-year"><year>2026</year></pub-date><volume>10</volume><issue>3</issue><history><date date-type="pub"><published-time>2026-04-24</published-time></date></history><abstract>MFA-conformer methods are widely used in English and Chinese speaker recognition. Theoretically language-independent but practically language-related, Tibetan speaker recognition currently relies on traditional models with poor performance. To address this, we adopt MFA-conformer as the basic framework and propose improvements: integrating 1D depth-wise separable convolution and channel attention into the conformer feed-forward network, fusing multi-block features, and adding an intra-class correlation regularizer to GE2E loss. Experiments show the improved model reduces the equal error rate (EER) compared with the conformer baseline.</abstract><keywords/></article-meta></front><body/><back><ref-list><ref id="B1" content-type="article"><label>1</label><element-citation publication-type="journal"><p>Boles A, Rad P, 2017, Voice Biometrics: Deep Learning-based Voiceprint Authentication System, In: 2017 12th System of Systems Engineering Conference (SoSE), 1–6.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B2" content-type="article"><label>2</label><element-citation publication-type="journal"><p>Chowdhury F, Wang Q, Moreno I, 2018, Attention-based Models for Text-Dependent Speaker Verification, In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 5359–5363.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B3" content-type="article"><label>3</label><element-citation publication-type="journal"><p>Peddinti V, Povey D, Khudanpur S, 2015, A Time Delay Neural Network Architecture for Efficient Modeling of Long Temporal Contexts., in Interspeech, 3214–3218.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B4" content-type="article"><label>4</label><element-citation publication-type="journal"><p>Sang M, Zhao Y, Liu G, et al., 2023, Improving Transformer-based Networks with Locality for Automatic Speaker Verification, In: ICASSP 2023- 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B5" content-type="article"><label>5</label><element-citation publication-type="journal"><p>Cai D, Li M, 2023, Leveraging ASR Pretrained Conformers for Speaker Verification through Transfer Learning and Knowledge Distillation, arXiv, https://doi.org/10.1109/TASLP.2024.3419426</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B6" content-type="article"><label>6</label><element-citation publication-type="journal"><p>Gulati A, Qin J, Chiu C, et al., 2020, Conformer: Convolution-Augmented Transformer for Speech Recognition, arXiv, https://doi.org/10.48550/arXiv.2005.08100</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B7" content-type="article"><label>7</label><element-citation publication-type="journal"><p>Zhang Y, Lv Z, Wu H, et al., 2022, Mfa-Conformer: Multi-Scale Feature Aggregation Conformer for Automatic Speaker Verification, arXiv, https://doi.org/10.48550/arXiv.2203.15249</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B8" content-type="article"><label>8</label><element-citation publication-type="journal"><p>Li L, Wang D, Rozi A, et al., 2017, Cross-Lingual Speaker Verification with Deep Feature Learning, In: 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 1040–1044.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B9" content-type="article"><label>9</label><element-citation publication-type="journal"><p>Wu Y, Liao W, 2021, Toward Text-Independent Cross-Lingual Speaker Recognition using English-Mandarin-Taiwanese Dataset, In: 2020 25th International Conference on Pattern Recognition (ICPR), 8515–8522.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B10" content-type="article"><label>10</label><element-citation publication-type="journal"><p>Thienpondt J, Desplanques B, Demuynck K, 2022, Tackling the Score Shift in Cross-Lingual Speaker Verification by Exploiting Language Information, In: ICASSP 2022–2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7187–7191.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B11" content-type="article"><label>11</label><element-citation publication-type="journal"><p>Schuessler A, 2024, Sino-Tibetan in Tibetan and Old Chinese. Language and Linguistics, 80–122.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B12" content-type="article"><label>12</label><element-citation publication-type="journal"><p>Mokgonyane T, Sefara T, Manamela M, et al., 2019, The Effects of Data Size on Text-Independent Automatic Speaker Identification System, In: 2019 International Conference on Advances in Big Data, Computing and Data Communication Systems (icABCD), 1–6.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B13" content-type="article"><label>13</label><element-citation publication-type="journal"><p>Wan L, Wang Q, Papir A, et al., 2018, Generalized End-to-End Loss for Speaker Verification, In: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4879–4883.</p><pub-id pub-id-type="doi"/></element-citation></ref><ref id="B14" content-type="article"><label>14</label><element-citation publication-type="journal"><p>Lyu H, Sha N, Qin S, et al., 2019, Advances in Neural Information Processing Systems. Advances in Neural Information Processing Systems, 32(2019).</p><pub-id pub-id-type="doi"/></element-citation></ref></ref-list></back></article>
