Dialect Identification from Automatically Transcribed Speech

Authors

  • Kellen Parker van Dam

DOI:

https://doi.org/10.15475/calcip.2026.2.3

Keywords:

ASR, automatic speech recognition, dialect identification

Abstract

This study investigates whether automatic speech recognition (ASR) output can support rapid dialect identification in Khiamniungan, an under-documented Tibeto-Burman language spoken across Northeast India and Myanmar. Using transcripts produced by an existing wav2vec 2.0 ASR model, we train two text-based dialect identification systems on 83.93 minutes of speech from six speakers representing three dialects: Thang, Nokhu, and Peshu. The systems use either regularized logistic regression over syllable, syllable-pair, and phoneme n-gram features, or character n-gram language models scored by perplexity. Under file-grouped cross-validation, both models correctly classify 37 of 38 held-out files, achieving 0.974 accuracy and 0.975 macro F1. The models also recover the same dialect relationships, grouping Thang and Nokhu against Peshu. On two previously unseen broadcasts, including one by a speaker whose home variety differs from the variety spoken on air, both systems correctly identify Thang from approximately half a minute of automatic transcription. These results show that ASR transcripts can provide sufficient phonological and lexical information for efficient dialect identification without additional neural modeling, offering a lightweight tool for organizing and comparing documentary language archives.

Downloads

Published

2026-09-23

How to Cite

Dialect Identification from Automatically Transcribed Speech. (2026). Computer-Assisted Language Comparison in Practice: Tutorials on Computational Approaches to the History and Diversity of Languages, 9(2), 95-110. https://doi.org/10.15475/calcip.2026.2.3