Dialect Identification from Automatically Transcribed Speech
DOI:
https://doi.org/10.15475/calcip.2026.2.3Keywords:
ASR, automatic speech recognition, dialect identificationAbstract
This study investigates whether automatic speech recognition (ASR) output can support rapid dialect identification in Khiamniungan, an under-documented Tibeto-Burman language spoken across Northeast India and Myanmar. Using transcripts produced by an existing wav2vec 2.0 ASR model, we train two text-based dialect identification systems on 83.93 minutes of speech from six speakers representing three dialects: Thang, Nokhu, and Peshu. The systems use either regularized logistic regression over syllable, syllable-pair, and phoneme n-gram features, or character n-gram language models scored by perplexity. Under file-grouped cross-validation, both models correctly classify 37 of 38 held-out files, achieving 0.974 accuracy and 0.975 macro F1. The models also recover the same dialect relationships, grouping Thang and Nokhu against Peshu. On two previously unseen broadcasts, including one by a speaker whose home variety differs from the variety spoken on air, both systems correctly identify Thang from approximately half a minute of automatic transcription. These results show that ASR transcripts can provide sufficient phonological and lexical information for efficient dialect identification without additional neural modeling, offering a lightweight tool for organizing and comparing documentary language archives.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Copyright remains with the author.

This work is licensed under a Creative Commons Attribution 4.0 International License.
As a general rule, all articles in this journal are published with CC-BY Attribution 4.0 License.