Mường Phonetic Data: From Linguistic Research To An Open Digital Resource

The Mường language is not only a means of communication among Mường communities but also an especially important source of evidence for studying the history and diversity of the Việt–Mường languages. Mường, however, is not entirely uniform across regions. Each area may preserve distinctive pronunciations, tones, and vocabulary. Collecting, systematizing, and comparing these differences is therefore highly significant for both linguistic research and the preservation of cultural heritage.
To make this valuable material more accessible to researchers and the public, the Digitizing Việt Nam team has developed the online dataset “Mường Phonetic Data,” based on linguist Nguyễn Văn Tài’s Ngữ âm tiếng Mường qua các phương ngôn (Mường Phonetics across Dialects). Published in Hanoi by the Encyclopedic Dictionary Publishing House in 2005, the book comprises 350 pages and includes maps. It examines the phonetics of numerous Mường dialects while also discussing the development of an appropriate transcription system for the language.
A Substantial Comparative Vocabulary Collection
At the heart of the dataset is a comparative vocabulary table containing 941 Vietnamese headwords and their corresponding forms in 30 Mường dialects. The varieties represented include Mường Thải, Tân Phong, Huy Thượng, Giáp Lai, Yến Mao, Mường Bi, Mường Khói, Mường Thàng, and Mường Động, as well as Mường Ống, Mường Rặc, Nghĩa Mai, Sông Con, Lâm La, and Cổ Liêm (Nguồn).
For each headword—such as “ác,” “ai,” “anh,” “ao,” “áo,” or “ăn”—users can examine its corresponding pronunciation in each locality. The lexical forms are presented in phonetic notation, with numerals indicating tones. By placing data from all 30 dialects side by side, the interface allows users to quickly identify similarities and differences in consonants, vowels, syllable structures, and tonal systems.
For example, the same Vietnamese headword may have similar pronunciations in some regions but markedly different forms in others. These variations are more than differences in ways of speaking: they also reflect the historical development of the language, patterns of contact between communities, and the geographical distribution of different population groups.
From Printed Pages to Searchable Data
One of the project’s most notable achievements is the transformation of a substantial body of specialized material from a printed book into structured digital data. Instead of manually searching through hundreds of pages, users can enter a Vietnamese word into the search box and compare its pronunciation across all 30 dialects.
This structure considerably broadens the potential uses of Nguyễn Văn Tài’s work. Linguists can draw on the data to study dialectology, phonetics, historical linguistics, and the relationship between Vietnamese and Mường. Teachers and students can use the comparative table as a visual resource in lessons on language variation. Mường people and other interested readers can likewise explore the linguistic richness found across different localities.
The dataset is classified as a digital resource in Mường and Vietnamese, covering the fields of linguistics and Vietnam’s ethnic communities. It is openly accessible for educational and research purposes, while commercial use is prohibited.
Preserving Mường Linguistic Diversity in the Digital Environment
At a time when many minority languages and dialects are under pressure from urbanization, migration, and the widespread use of languages with larger speaker populations, digitizing Mường-language materials has significance extending well beyond the boundaries of an academic project. Every recorded pronunciation preserves a trace of communal memory and of the cultural environment in which that linguistic form emerged and was passed down.
The dataset cannot replace the living speech of native speakers, nor does it encompass the full linguistic life of Mường communities. Nevertheless, it provides an important reference point. When combined with audio materials, fieldwork, and community participation, the data can support dialect mapping, the development of teaching materials, research into language change, and the preservation of features that may gradually be disappearing.
Building on Nguyễn Văn Tài’s meticulous scholarship, Digitizing Việt Nam has created a digital platform through which users can access Mường phonetic and lexical resources more easily. The dataset both extends the value of a foundational linguistic study and demonstrates the potential of digital technology to preserve, organize, and share Vietnam’s linguistic heritage.
Explore the dataset on the Digitizing Việt Nam platform:
https://www.digitizingvietnam.com/en/our-collections/dan-toc-viet-nam-ngon-ngu-van-hoa/ngu-am-tieng-muong