Proceedings of the sixth International Conference on Asian Geolinguistics
Abstract
This volume contains papers presented at the sixth International Conference on Asian Geolinguistics (ICAG) held at Universiti Kebangsaan Malaysia, Kuala Lumpur, Malaysia, on the 5th of May, 2025.
Full text
Proceedings of the sixth International Conference on Asian Geolinguistics EDITED BY NOR HASHIMAH JALALUDDIN, HIROYUKI SUZUKI, AND MITSUAKI ENDO Geolinguistic Society of Japan
ISSN 2436-6471 Proceedings of the sixth International Conference on Asian Geolinguistics EDITED BY N OR H ASHIMAH J ALALUDDIN , H IROYUKI S UZUKI , AND M ITSUAKI E NDO Geolinguistic Society of Japan 2025
Studies in Geolinguistics, Monograph series, No. 10 ISSN 2436-6471 Proceedings of the sixth International Conference on Asian Geolinguistics, edited by Nor Hashimah Jalaluddin, Hiroyuki Suzuki, and Mitsuaki Endo, 2025 Cover photo: Institute of the Malay world and civilisation © 2025 Mika Fukazawa First published 2025 doi: https://doi.org/10.5281/zenodo.17204665 © 2025 by Authors. All rights reserved. Published by: Geolinguistic Society of Japan Website: https://geolinguistics.sakura.ne.jp/ Office address: Room G1305, Aoyama Gakuin Univeristy, 4-4-25 Shibuya, Shibuya-ku, Tokyo
Proceedings of the sixth International Conference on Asian Geolinguistics i Preface This volume contains papers presented at the sixth International Conference on Asian Geolinguistics (ICAG) held at Universiti Kebangsaan Malaysia, Kuala Lumpur, Malaysia, on the 5th of May, 2025. The previous ICAGs were held as follows: 1st, at Aoyama Gakuin University, Tokyo, 2012; 2nd, at Chulalongkorn University, Bangkok, 2014; 3rd, at Royal University of Phnom Penh, Phnom Penh, 2016; 4th, at Fakultas Ilmu Pengetahuan Budaya Universitas Indonesia, Jakarta, 2018; and 5th, at the University of Social Sciences and Humanities, VNU, Ha Noi, Vietnam, 2023. The proceedings of each ICAG (PICAG) have been edited and published as open access documents; see below for bibliographical information. Before PICAG-5, the titles are denoted as Papers from...; as of PICAG-5, Proceedings of... is used. PICAG-1: Endo, Mitsuaki (ed.) (2012) Papers from the First International Conference on Asian Geolinguistics. https://doi.org/10.5281/zenodo.6423581 PICAG-2: Endo, Mitsuaki (ed.) (2014) Papers from the Second International Conference on Asian Geolinguistics. https://doi.org/10.5281/zenodo.6423601 PICAG-3: Endo, Mitsuaki (ed.) (2016) Papers from the Third International Conference on Asian Geolinguistics. Fuchu: Research Institute for Languages and Cultures of Asian and Africa. https://publication.aa-ken.jp/papers_3IC_Asian_geolinguistics_2016.pdf PICAG-4: Suzuki, Hiroyuki and Mitsuaki Endo (eds.) (2018) Papers from the Fourth International Conference of Asian Geolinguistics. Fuchu: Research Institute for Languages and Cultures of Asia and Africa. https://publication.aaken.jp/papers_4IC_Asian_geolinguistics_2018.pdf PICAG-5: Trịnh, Cẩm Lan, Trần Thị Hồng Hạnh, Hiroyuki Suzuki, and Mitsuaki Endo (eds.) (2023) Proceedings of the fifth International Conference on Asian Geolinguistics. Tokyo: Geolinguistic Society of Japan. https://doi.org/10.5281/zenodo.8382130 The editors
ii
Proceedings of the sixth International Conference on Asian Geolinguistics iii TABLE OF CONTENTS Preface i Conference Program v Chitsuko Fukushima Revisiting ‘sun’ in the Linguistic Atlas of Asia: Focusing on semantic extensions and differentiation . . . . . . . . . . . . . . . . . . . . . . . . . . 1 Mitsuaki Endo Homelands, migration, and dispersion of the Kra-Dai . . . . . . . . . . . . . . . . . . . . . 18 TRẦN Thị Hồng Hạnh and TRƯƠNG Nhật Vinh A geographic distribution of the word form for ‘frog/toad’ among Mon-Khmer languages in Vietnam and its implication . . . . . . . . . . . . . . . . 33 Hiroyuki Suzuki Geolinguistic patterns of the word form for ‘butterfly’ in Tibetic languages of the eastern Tibetosphere . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 50 Mika Fukazawa Extraction of regularities and geographical patterns from the basic vocabulary of the Ainu language . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62 Trịnh Cẩm Lan The distribution of /l/ and /n/ variants in the Red River Delta, Vietnam . . . . . . . 81 Atsuko Utsumi Diversity in noun markers and grammatical voice systems in the languages of the Philippines and Indonesia . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 Khairul Ashraaf Saari, Nor Hashimah Jalaluddin, and Harishon Radzi Dialect transition along the Perak River . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 118 ❦
iv
Proceedings of the sixth International Conference on Asian Geolinguistics v CONFERENCE PROGRAM The Sixth International Conference on Asian Geolinguistics Venue: Sudut Wacana, Level 3, Institut Alam dan Tamadun Melayu, Universiti Kebangsaan Malaysia Kuala Lumpur, Malaysia Online: Zoom Meeting 5th May 2025 Opening ceremony 8:55 – 9:00 Malaysia time Nor Hashimah Jalaluddin (Universiti Kebangsaan Malaysia) Welcoming speech First session 9:00 – 10:30 Chair: Nor Hashimah Jalaluddin (Universiti Kebangsaan Malaysia) 1. Chitsuko Fukushima (University of Niigata Prefecture), Revisiting ‘Sun’ in the Linguistic Atlas of Asia: Focusing on semantic extensions and differentiation 2. Ian Joo (Otaru University of Commerce) and Yu-Yin Hsu (The Hong Kong Polytechnic University), Correlation between longitude and maximal lengths of onset and coda in Eurasia 3. Mitsuaki Endo (Aoyama Gakuin University), Homelands, migration, and dispersion of the Kra-Dai Second session 10:45 - 11:45 Chair: Trịnh Cẩm Lan (University of Social Sciences and Humanities, VNU) 4. [online] Li Jinhua and Wang Lei (Nanjing University), Geolinguistic Analysis of the Word Form for ‘Cat’ in Chinese Korean 5. Trần Thị Hồng Hạnh and TRUONG Nhat Vinh (University of Social Sciences and Humanities, VNU), A geolinguistic analysis of the word form for ‘frog and toad’ among Mon Khmer languages in Vietnam 6. [online] Hiroyuki SUZUKI (Kyoto University), Geolinguistic analysis of the word form for ‘butterfly’ in Tibetic languages of the eastern Tibetosphere
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 6 show the distributions of etymological motivations. Again, most distributions are found nearby that of ‘day’, specifically in Southeast and South Asia, except that in Arabic. Figure 6: Words meaning ‘day’ in Asia F igure 7: Words meaning ‘sunshine, light’ in Asia
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 7 Figure 8: Words meaning ‘daytime’ in Asia Figure 9 Words meaning ‘sky, heaven’ Figure 9: Words meaning ‘sky, heaven’ in Asia Figure 10: Words meaning ‘time’ in Asia
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 8 3. Semantic extensions and differentiation of ‘sun’ in individual languages/language groups in Asia In this section, maps are introduced to show the variation in each language/language group. The comparison of maps shows the changes that happened. Figure 11: Words meaning ‘hot’ in Asia Figure 12: Words meaning 'mountain' in Asia
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 9 3.1. Japanese and Sinitic See Figure 13. Japanese shows a small distribution of hi which means ‘day’ since Chinese words (e.g. taiyo) and honorific words (e.g. ohisama) which mean only ‘sun’ are in everyday use. Sinitic shows the distribution of ri only in the South while disyllabic words and words of personification are used in many locations. This map shows the change from polysemy to monosemy in Japanese and Sinitic. 3.2. Austronesian Figure 14 shows the distribution of words which mean ‘day’ and ‘light’. Figure 15 shows the distribution of mata ‘eye’ and its compounds which mean only ‘sun’. The distribution of Figure 15 seems newer compared with that of Figure 14. These maps also show the change from polysemy to monosemy. 3.3. Austroasiatic Figure 16 shows the semantic variation and Figure 17 shows the distribution of the compound type. The maps also show the change from polysemy to monosemy. 3.4. Kra-Dai Figure 18 shows the semantic variation of ‘day’ and ‘time’. Figure 19 shows the distribution of the word meaning ‘eye’ and its compounds. The maps again show the change from polysemy to monosemy. Figure 13: Words meaning ‘day’ in Asia: Focus on Japanese and Sinitic
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 10 Figure 14: Semantic variation of ‘sun’ in Austronesian Figure 15: Distribution of mata ‘eye’ and its compounds in Austronesian Figure 16: Semantic variation of ‘sun’ in Austroasiatic
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 11 Figure 17: Distribution of the compound type in Austroasiatic Figure 18: Semantic variation of ‘sun’ in Kra-Dai Figure 19: Distribution of ‘eye’ and its compounds in Kra-Dai
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 12 3.5. Tibeto-Burman Figure 20 shows the following semantic variation: ‘day, dwell’, ‘sunshine, bright’, ‘sky, heaven, clouds’, and ‘set (of the sun)’. Figure 21 shows the distribution of various types of compounds in the south-eastern part of the Tibeto-Burman area. The distribution of the compound type in Tibeto-Burman (Figure 21) is in succession with those of the compound type in Kra-Dai (figure 17) and in Austroasiatic (Figure 19). Figure 20: Semantic variation of ‘sun’ in Tibeto-Burman Figure 21: Distribution of the compound type in Tibeto-Burman <enlarged>
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 13 3.6. Turkic Figures 22 shows the semantic variation in Turkic. The change from ‘sun’ and ’day’ to ‘sun’ is attested. According to Saito (2021: 48): The form kün, which originally meant ‘sun,’ has extended to refer also to ‘day.’ … Chulym developed a compound word kün kara ɣ ɪ with the word karak ‘eye.’ … The type B word (küneš type) is used (almost?) exclusively for ‘sun’ in Turkish, Crimean Tatar, and Chuvash, while the type A word (kün type) is used for both ’sun’ and ‘day’ in the other languages. The ‘eye of the day’ compound is used in two distant locations (Chulym and Nogai 1 ). This information is valuable since Urban (2010) identified the linguistic pattern only in the language families of Southeast Asia and Oceania. 3.7. Uralic Figure 23 shows the semantic variation in Uralic. The languages with the meaning of ‘day’ are located in the north. This is a peripheral distribution, so the polysemy is older. 1 kün közi (köz ‘eye’). This information of Nogai has been recently provided by Prof. Yoshio Saitô. Figure 22: Semantic variation of ‘sun’ in Turkic Figure 23: Semantic variation in Uralic
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 14 3.8. South Asia Figures 24 to 27 show the semantic variation in each language group in South Asia. Figure 24 shows the distribution of ‘sunshine’ and ‘hot’ in Iranian, Figure 25 that of ‘sunshine’, ‘daytime’, ‘mountain’, and ‘hot’ in Aryan, Figure 26 that of ‘time’ in Dravidian, and Figure 27 that of ‘hot’ in Andaman. Figure 24: Semantic variation of ‘sun’ in Iranian Figure 25: Semantic variation in Aryan
F UKUSHIMA : R EVISITING ‘ SUN ’ IN THE L INGUISTIC A TLAS OF A SIA 15 Figure 26: Semantic variation in Dravidian Figure 27: Semantic variation in Andaman
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 22 Figure 5: Initial consonants of the word for “six” in Kra-Dai Figure 6: Initial consonants of the word for “two” in Kra-Dai
E NDO : H OMELANDS , MIGRATION AND DISPERSION OF THE K RA -D AI 23 Figure 7: The word for “iron” in Kra-Dai 2.4. Evidence from genetics Changmai et al. (2023), who analyzed the genome data of the Kra-Dai-speaking population in Thailand, found geographical differences in DNA proportion (Figure 8). The Lao element was generally predominant, implying that the mainstream Thai population came through Laos. In the northeastern part, the Cambodian element is observed in some places, indicating that mixture occurred between Kra-Dai-speaking people and indigenous Austroasiatic-speaking people. In fact, half of the population around the Surin area near the Cambodian border are Austronesian-speaking people. Thailand’s northern district shows a considerably high ratio of Zhuang from China and Thai Vietnam, suggesting the homeland and stopover point of these human groups. The Dai element is also recognized in this area. Dai people dwelled in Yunnan, China. The Zhuang element occupies the main part of Hmong Daw in the Golden Triangle, which
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 24 means they experienced a language shift from Kra-Dai to Hmong. Thus, genetics provides not only evidence of population movements but also migration dating. Figure 8: Sources of recent ancestry in Thai groups 3. Existence of the earlier Kra-Dai population in Guangdong province Figure 1 shows how the Kra-Dai is exceptionally distributed in the Guangdong province today. However, because Kra-Dai origin toponyms are generally found in this area, Kra-Dai people can be said to have lived there in ancient times. Xu (1939) enumerated toponyms with Kra-Dai elements in Guangdong province as shown in Figures 9–16, where the maps present different distribution patterns. Some toponyms meant “village” (Figures 11, 12, and 15) because of dialectal differences. Some of them are generally distributed in Hainan Island, covering the area where the Hlai people lived (Figures 9 and 16) or partially encompassing the area of the Lingao people (Figures 12 and 14). Ostapirat (1998) reported a “Chinese” dialect in the Leizhou Peninsula in Guangdong, in which a small amount of the basic vocabulary is of Kra-Dai origin; otherwise, the whole linguistic system resembles a Chinese dialect and is a vestige of a mainland Lingao language, which is essentially the same as the nature of Cantonese dialects.
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 25 Figure 9: Toponyms with na (rice field) Figure 10: Toponyms with du (10 villages) Figure 11: Toponyms with si (village) Figure 12: Toponyms with gu (village) Figure 13: Toponyms wit lu (mountain) Figure 14: Toponyms with luo (various meanings) Figure 15: Toponyms with yun (baan, village) Figure 16: Toponyms with li
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 26 4. The Kra-Dai population in southeastern China as a whole 4.1. The earlier Kra-Dai population in Chu state during the Warring States period Lin et al. (2004) examined , meaning “one,” in bamboo slip documents in the Chu state around the Hunan and Hubei areas dating back to the Warring States period (BC 5–3) and further compared it with [nɯŋ] “one” in modern Tai. They considered that some Tai-speaking people adopted the Tai sound 能, meaning “one,” as a phonetic element. 4.2. Kra-Dai toponyms in southeastern China Zhou and You (2019) expanded the scope of the Kra-Dai origin toponyms to the entire southeastern China. Figure 17 illustrates the geographical distribution of the Kra-Dai toponyms lai “stream” (◯), luo “mountain” (▲), and tan (no meaning provided). They are vestiges of the Kra-Dai in this area. Li (2001) reported that a dialect in the Daic substratum was used near Shanghai. Figure 17: Kra-Dai toponyms lai, luo, and tan in southern China (Zhou and You 2019)
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 27 5. Theories on the Kra-Dai migration process 5.1. Chamberlain (2016) Chamberlain (2016: 67-70) provided a general picture of the Kra-Dai migration process as follows: Figure 18: The Kra-Dai migration process (Chamberlain 2016) Phase 1 – 1100-800 BCE Shang ends, Zhou begins, Chǔ is born. Kra separates from the Kra-Dai mainstream. Phase 2 – Early Spring and Autumn Period, 771-685 BCE Political upheavals in Zhou causing capital to relocate Li splits off and moves south/southeast. Phase 3 – Chǔ Hegemony – Late Spring and Autumn, 613-453 BCE King Zhuang expands Chǔ which becomes the most powerful state Yue (Be-Tai) separates from Kam-Sui moves east to the coast, conquers Wu. Phase 4 – Warring States Period, 475-221 BCE Chǔ annexes Yue, 333 BCE; Qin conquers Chǔ 223 BCE. Yue royal families begin to move south, forming the Bai Yue. Luo Yue (Central Southwestern Tai) overruns the lands of the Rei and the Kra in southern Lingnan and Annam.
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 28 Xi Ou (Northern Tai) follows and comes to dominate northern Lingnan, including some areas formerly held by Luo Yue in Jiuzhen. Phase 5 – Qin Dynasty 223-206 BCE; Han Dynasty, 206 BCE–220 CE Qin-Han begins colonization of the south, establishing commanderies at Canton, Jiaozhi, and Jiuzhen. Recorded history of the south begins. Mobile Yue Central Southwestern Tai polities continue to establish chiefdoms, dominating the original Kra and populations in western Annam, and Rei in Jiuzhen. Ou Yue and Yi Northern Tais push west from Nan Yue. Be and Sek separate and move west and south from Nan Yue, respectively. 5.2. Blench (2018) Blench (2018) discussed the genetic relation between Daic and Austronesian based on basic vocabulary and provided a scheme (Figure 19). Figure 19: Genesis of Daic languages and their association with Austronesian (Blench 2018; https://en.wikipedia.org/wiki/Kra%E2%80%93Dai_languages#cite_note-Blench2018-42) 5.3. Tao et al. (2023) Tao et al. (2023) inferred migration routes and chronology by applying Bayesian phylogenetic method to 600 vocabulary items of 100 Kra-Dai languages. However, the language dispersion process and migration routes do not necessarily always coincide. In addition, the chronology in Figure 20 shows mean dating, while the minimal datings in Figure 21 seem more plausible.
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 29 Figure 20: Inferred dispersal routes of Kra-Dai speakers and their languages in prehistory (Tao et al. 2023) Figure 21: Maximum clade credibility tree with divergence time of the Kra-Dai languages (Tao et al. 2023)
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 30 6. Conclusion In sum, Kra-Dai, as a branch of Austronesian, once spread in almost all areas in southeastern China. Following Sagart’s Sino–Tibetan–Austronesian theory, it has further northwestern origins from the Gansu–Qinghai–Shaanxi area, migrating through Henan, Shandong, and Jiangsu and then to Fujian province. Endo (forthcoming) provides an in-depth discussion of the maritime migration routes of the Hlai and Kra branches of Kra-Dai. As observed in section 2.4 above, language shifts also occurred rather frequently. Li et al. (2010) showed that Mien (Hmong-Mien) is clustered with Sui, Hlai, Bouyei, and Thai (all are Kra-Dai) in terms of Y chromosomes, reflecting paternal origins, which suggests that Mien people are of mainly Kra-Dai origin and accepted Hmong through language contact. Such a language shift is rather attributed to language dispersion and not migration. Even today, southern China is home to Austronesian, Austroasiatic, Hmong-Mien, and Sino-Tibetan people besides the Kra-Dai. After the Qin-Han period from 221 BC to 220 AD, the Chinese-speaking area significantly expanded, with human groups migrating, mixing, and changing their languages before and after. Their macro and micro histories can be successfully explained because of the collaboration between linguistics, archaeology, genetics, and other related disciplines. References Baker, Chris (2002) From Yue to Tai. Journal of Siam Society 90(1/2): 1–26. https://thesiamsociety.org/wpcontent/uploads/2002/03/JSS_090_0b_Baker_YueToThai.pdf Baker, Chris and Pasuk Phongpaichit (2017) A History of Ayutthaya. Cambridge: Cambridge University Press. Blench, Roger (2018) Tai-Kadai and Austronesian are Related at Multiple Levels and their Archaeological Interpretation (draft). https://en.wikipedia.org/wiki/Kra%E2%80%93Dai_languages Chamberlain, James R. (2016) Kra-Dai and proto history of South China and Vietnam, Journal of Siam Society 104: 27-77. https://thesiamsociety.org/wp-content/uploads/2016/04/ JSS_104_0c_Chamberlain_KraDaiAndProtoHistoryofSouthChinaAndVietnam.pdf Changmai, P., Phongbunchoo, Y., Kočí, J. et al. (2023) Reanalyzing the genetic history of KraDai speakers from Thailand and new insights into their genetic interactions beyond Mainland Southeast Asia. Scientific Reports 13, 8371. https://doi.org/10.1038/s41598-02335507-8
ENDO: HOMELANDS, MIGRATION AND DISPERSION OF THE KRA-DAI 31 Endo, Mitsuaki (2016) Geographical distribution of the /r/ type sounds in Zhuang. Papers from the Third International Conference on Asian Geolinguistics. 46–71. URI: https://publication.aa-ken.jp/papers_3IC_Asian_geolinguistics_2016.pdf Endo, Mitsuaki (2017) Iron: Tai-Kadai, Studies in Asian Geolinguistics 5: 15–16. https://publication.aa-ken.jp/sag5_iron_2017.pdf Endo, Mitsuaki (2022) Subgrouping of Kra-Dai. In Hiroyuki Suzuki, Mika Fukazawa, Akiko Yokoyama, and Mitsuaki Endo (eds.) Linguistic Atlas of Asia and Africa I, 9. Tokyo: Geolinguistic Society of Japan. https://doi.org/10.5281/zenodo.7118188 Endo, Mitsuaki (forthcoming) Words for the numbers one to ten reflect the migration patterns of the Hlai and Kra people from Taiwan. Li, Dongna et al. (2010) Genetic origin of Kadai-speaking Gelong people on Hainan island viewed from Y chromosomes. Journal of Human Genetics 55: 462–468. https://doi.org/10.1038/jhg.2010.50 Li, Fang Kuei (1977) A handbook of comparative Tai. Honolulu: University of Hawai‘i Press. Li, Hui [李辉] (2001) Shanghai Maqiao hua de Taiyu diceng cihui 上海马桥话的台语底层词 汇 [Daic Background Vocabulary in Shanghai Maqiao Dialect]. Qiong-Tai shaoshu minzu xueshu wenhua jiaoliu yantaohui lunwenji《琼台少数民族学术文化交流研讨会论文集》 [Proceedings for Conference of Minority Cultures in Hainan and Taiwan]. https://web.archive.org/web/20180327144856/http://loca.fudan.edu.cn/lh/Doc/D02.pdf Liang, Min [梁敏] and Junru Zhang [张均如] (1996) Dongtai Yuzu Gailun《侗台语族概论》 [Synopsis of Kra-Dai]. Beijing : Zhongguo Shehui Kexue Chubanshe. Lin, Hongying [林虹瑛], Murase Nozomi [村瀬望], Furuya Akihiro [古屋昭弘] (2004) Sengoku moji 𬚆 ni tsuite [On the character 𬚆 during Warring States period]. Kaipian『開 篇』23: 71–75. Pittayawat, Pittayaporn (2009) The phonology of proto-Tai. Ph. D. Dissertation, Cornell University. https://ecommons.cornell.edu/server/api/core/bitstreams/6af02aa7-c444-481c8d1b-ac0c25346f20/content Ostapirat, Weera (1998) A mainland Be language?. Journal of Chinese Linguistics 26(2): 338– 344. Sagart, Laurent (1999) The roots of Old Chinese. Amsterdam : John Benjamins. https://starlingdb.org/Texts/Students/Sagart%2C%20Laurent/The%20Roots%20of%20Ol d%20Chinese%20%281999%29.pdf Tao, Y., Wei, Y., Ge, J. et al. (2023) Phylogenetic evidence reveals early Kra-Dai divergence and dispersal in the late Holocene. Nature Communications 14, 6924. https://www.nature.com/articles/s41467-023-42761-x Yuan, Jiahua [袁家骅] (1963) Zhuangyu / r / de Fangyin Duiying 壮语/ r /的方音对应 [Sound correspondences of / r / among the Zhuang dialects]. Yuyanxue Luncong《语言学论丛》5: 187–218. Xu, Songshi [徐松石] (1939) Yuejiang liuyu renmin shi《粤江流域人民史》[History of people in Yuejiang river basin].
T RẦN AND T RƯƠNG : A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘ FROG / TOAD ’ 38 Table 1: Types of Mon Khmer terms for ‘frog’ and ‘toad’ Table 1 indeed reveals a significant phonological structural consistency across the 'frog' and 'toad' terms in Mon-Khmer languages, despite their classification into monosyllabic, sesquisyllabic, and disyllabic forms. Regarding the geographical distribution of these variants, Map 1 provides a visual representation. Map 1: Distribution map of “toad” in Mon-Khmer languages in Vietnam. Type Subtype Presyllable Main Syllable Tone Consonant Vowel Final consonant A Monosyllable CVC A1 /ʔ/ /ɛ/ /eː/ /u//u:/ /o/ /ɔ/ /ɑ/, /ɒ/ /k/ /t/ sắc (rising tone), nặng (low dropping tone) A2 /k/, /kʰ/ A3 /ɣ/ A4 /r/, /ʑ/ A5 /tʰ/ B Sesquisyllable CvCVC B /a/ /u/ /k/ /t/ /r/, /ʑ/ C Disyllabic/ Reduplicative form C Reduplicative form /k/ /r/ /o/ /ɔ/ /ɑ/, /ɒ/ /e/ /k/ /t/ //∅
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 39 Type A's distribution largely coincides with the primary habitation areas of Vietic language speakers. Type B is predominantly found in regions inhabited by conservative subgroups of Vietic, Khmuic, and Katuic languages. Type C, conversely, is more widely distributed across regions inhabited by Vietic, Bahnaric, Khmeric, and Katuic language communities. Despite these observed distributional differences among Type A, B, and C variants, their overall spatial arrangement exhibits considerable interspersion. This pattern may be indicative of historical internal evolution of subgroups within the Mon-Khmer branch under different contexts. However, our fieldwork also revealed phonological forms similar to Type A variants in several Tai languages: Tày language in Vietnam and Zhuang language in Guangxi, China. Table 2: ‘Toad’ in some Tai languages in Vietnam and Southern China Word for ‘toad’ Language/Dialect /kɔk/ Zhuang (Liujiang, Liuzhou) /kok/ Zhuang (Du'an, Hezhi) /kok/ Zhuang (Longzhou, Chongzuo) /kok/ Zhuang (Nanning) cáy gộc /kaj ɣɔk/ Tày Bắc Giang cáy rộc /kaj rɔk/ Tày Trùng Khánh cấy cộc /kə i kok/ Tày Lạng Sơn ca cộc /ka kok/ Nùng Bắc Cạn This is quite compatible with the word ‘frog’ from the Zhuang-Chinese-English Dictionary, a resource primarily based on Nong Zhuang dialect from Zhetu District in Guangnan County, Yunnan Province. In this dictionary, the entry for 'frog' is annotated as: frog: [n] goap (gaep gaep); [n] gvej2; [n] gaeuh 2 frog, large edible [n]: goep frog, large male; bullfrog [n]: goepcwz, goepndai frog, male of a small variety of edible [n]: gvejndai frog, mountain [n]: goep-ndoeng
T RẦN AND T RƯƠNG : A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘ FROG / TOAD ’ 40 However, the dictionary's entry for 'toad' displays a phonological form that bears little clear connection, that is gungqsou. Map 2: Distribution map of “toad” in Mon-Khmer and Taic languages in Vietnam and Southern China It is worth noting that, in a previous study, Endo (2001) Endo mentioned the word for ‘toad’ in several dialects in southern China and stated that they are related to Tai languages, including Zhuang. In discussing the word '蛤' (frog/toad) in several Chinese dialects, he noted the word for toads *kop in southern dialects of Chinese, Yue (Cantonese), Min (Fukkienese), and Xiang (Hunanese). It also appeared in the Fangyan dating back to the Han dynasty with locality of the Guangxi Zhuang Autonomous Region now. According to Hashimoto (1976, as cited in Endo 2001, p. 146), it came from Tai languages. In Guangzhou dialect, '蛤' is pronounced [kɐp 下陰入/xià yīn rù/a low entering tone] and refers to a type of frog or toad. According to Hashimoto (1976, as cited in Endo 2001, p. 146), the word "frog" in Cantonese is found in Guangzhou [kap], Taishan [ka:p], Yangchun [kap], and Huaxian [kuop] (all with yin), which correspond to the Tai languages Boai [kop], Zhuang [kop], and Siamese [kop] (all with 7 tones), and is a very clear layer of Thai language. Although the Chinese Dialect Vocabulary lists words like '蜂蜂' and '田鶏' as equivalents for ‘frog’ in Cantonese, other dictionaries do indeed include terms containing '蛤'. Based on the characters, it would appear that Northern dialects also have ' 蛤蟆' (háma) for frog/toad, but phonologically, the Beijing Mandarin háma is better understood as deriving from '蝦' (pronounced with the “hu + jia” reading in Guangyun), rather than from '蛤' (with the 'gu + xiang' reading in Guangyun). In the Chinese Dialect Vocabulary entry for
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 41 'frog' (青蛙), dialects such as Yangjiang, Xiamen, Chaozhou, Fuzhou, and Changsha — primarily Yue and Hakka dialects — include terms that contain '蛤'. Thus, a key question arises from this close geographic distribution presented on the Map 2: does it represent mere spatial coincidence, or does it suggest the occurrence of a specific historical linguistic event? To address this question and to investigate the etyma for ‘frog’ and ‘toad’ in the Mon-Khmer languages, the next section of this paper will continue the discussion, supported by evidence from historical linguistics, archaeology, and genetics. 4. Discussion 4.1. Historical linguistic data Regarding the Vietnamese word ếch ‘frog’, the online World Loanword Database (WOLD) indicates no evidence of borrowing. Data for ‘toad’ is absent from the database. As Vietnamese is the only Mon-Khmer language represented in the WOLD corpus, its value for broader comparative analysis within the Mon-Khmer language family is consequently limited in the present discussion. This section therefore examines the reconstructed forms for ‘frog’ and ‘toad’ in Mon-Khmer languages, drawing primarily on data from the Mon-Khmer Etymological Dictionary (MKED). Although certain reconstruction results in the MKED remain open to scholarly debate (cf. Sidwell & Alves 2023), the database is nonetheless considered a highly reliable source for the historical linguistic study of Mon-Khmer languages. Table 3: Reconstructed forms of ‘frog’ and ‘toad’ in Mon-Khmer languages Word Branch/group Reconstructed forms frog Proto-Mon Khmer *[ ]r[ɔ]k, *kit, *kiit, *ku[ə]t, *kət, *kəət Proto-Bahnaric *ki(ː)t Proto-Katuic *ʔaguut Proto-Khmuic *ɟʔɛʔ Proto-Vietic *gɔːt, *ʔeːk toad Proto-Mon Khmer *[ ]r[ɔ]k Proto-Bahnaric *-rɔk Proto-Katuic *ʔaguut traak, *ʔagɔk, *ʔarɔk, *ʔaʔɔk
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 42 Proto-Khmuic *hro(ː)k Proto-Vietic *-rɔk, *raːk, *roːk, *-duːt It appears that Proto-Mon-Khmer likely possessed a single term *[ ]r[ɔ]k encompassing both 'frog' and 'toad'. This form represents a robust and well-supported reconstruction for 'toad' across several branches—including Bahnaric, Katuic, Khmuic, and Vietic—and for 'frog' at the deepest Proto-Mon Khmer level. As such, it stands out as a strong candidate for the common etymon of these amphibians. According to Alves (2009, p. 4), a form recognized in Proto-Mon-Khmer is not considered a loanword even when it appears in other language families, such as Tai-Kadai or Sino-Tibetan. Thus, based on this evidence, it can be asserted that the etymon for both ‘frog’ and ‘toad’ in Mon-Khmer languages is *[ ]r[ɔ]k. This analysis leads us to revisit the question, introduced in Section 3, regarding any potential connection or implication between the phonetic forms of ‘frog/toad’ in MonKhmer and some Tai languages. The interaction and relationship between Mon-Khmer and Tai languages have been a significant and extensively researched topic in the historical linguistics of Southeast Asia. Although they belong to two distinct language families, their long history of geographical and cultural contact has led to considerable influences. Alves (2009, p. 4) also stated that, based on somewhat tentative data—essentially a few dozen loanwords at most (Nguyễn Tài Cẩn 1995, p. 322, as cited in Alves 2009, p. 4)—Proto-Tai peoples may have come into contact with the ancestors of the Vietnamese before the Han Dynasty (200 BCE to 200 CE). At that time, they presumably shared with the Vietnamese some technology related to agriculture and animal husbandry and the associated vocabulary (e.g., vịt “duck” and mương “canal”). Pittayaporn (2019) determines Proto-Tai is the ancestor of the Tai languages of Mainland Southeast Asia. In his Proto-Tai reconstructions, the form for ‘frog’ is: *kɤpD, and for ‘small frog’ is: *krwe:C. This reconstructed form, however, does not provide significant insight into the question under consideration. Therefore, additional evidence is needed to resolve the issue. 4.2. Additional archaeological and genetic data Drawing upon historical linguistic analysis and archaeological-genetic research findings, scholars such as Alves (2019, 2022b) and Trần Trí Dõi (2022) agree that speakers of Austroasiatic languages have been present in the Red River Delta since approximately 2000 BCE. During the Phùng Nguyên culture, the region's population
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 43 was predominantly Austroasiatic-speaking. By the Đông Sơn culture, this Austroasiatic-speaking populace gradually diversified into groups speaking languages of the Vietic branch, who likely engaged in varying degrees of contact with neighboring Tai-speaking populations. This period of initial linguistic contact between Vietic and Tai peoples likely commenced in the Iron Age, around 500 BCE, coinciding with the onset of the Đông Sơn culture and significant socio-cultural developments leading to the Cổ Loa site in northern Vietnam (Alves 2022a, p. 18). Trần Trí Dõi (2022) affirms that among various hypotheses regarding the language of Đông Sơn culture residents, the proposition that this community primarily spoke Vietic languages and engaged in contact with neighboring Tai-speaking populations seems to be strongly supported by linguistic, archaeological evidence. Consequently, this view is largely accepted by many researchers of Vietnamese linguistic, cultural, and societal history during the Đông Sơn period. This historical and linguistic understanding of the region's population dynamics is strongly corroborated by archaeological and recent genetic evidence. Indeed, specific genetic analyses provide concrete support for this narrative. For instance, research by Vietnamese and international scholars (Dang Liu et al. 2019) indicates that the genetic make-up of the Vietnamese people is directly inherited from the Đông Sơn culture, showing only approximately 10-15% genetic divergence compared to ancient samples from the Đông Sơn culture site of Núi Nấp. Furthermore, evidence from craniometric measurements suggests that the Phùng Nguyên and Đông Sơn cultures were characterized by Austroasiatic-speaking populations (Matsumura et al., 2019). Such findings reinforce the long-standing presence and predominant influence of Austroasiatic-speaking groups in the Red River Delta during these formative periods. When examining Đông Sơn culture, the bronze drum emerges as a profoundly significant archaeological artifact. Scholarly discourse concerning its origins primarily revolves around two ancient formative centers: Vietnam and Yunnan (China). Both are recognized as the foremost production hubs during the early development of bronze drum culture. As Chiou-Peng (2009) notes, "The data accumulated for more than half a century have now attested that comparable artifacts from the Dong Son region and Yunnan belonged to two cognate cultures in the Yue-based cultural sphere of southwest China and the Indochinese peninsula" (Chiou-Peng 2009, p. 34). Regarding the classification of bronze drums, a slight divergence exists between Chinese and Vietnamese scholarly systems. Chinese scholars have categorized drums discovered within China and those held in Chinese collections into eight distinct types. This differs from the system proposed by Fr. Heger, currently adopted by Vietnamese
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 44 researchers, which includes one fewer type. Despite these classification differences, there is a broad consensus among researchers regarding the paramount importance of Heger Type I bronze drums, widely considered the earliest form. Notably, Vietnam possesses the largest quantity of these crucial artifacts, with their number nearly double that found in China and equivalent to the combined total from all other nations (Li & Huang 2016). In Vietnam, Heger I bronze drums, also designated as Đông Sơn bronze drums, are further categorized by archaeologists into five chronologically ordered groups: A, B, C, D, and Đ. A detailed presentation of this system is omitted here, as bronze drum research is not the primary focus of this paper. Nevertheless, this classification is pertinent, as toad/frog statues, a key subject of interest, began to appear on the drum faces of those belonging to Group C (Phạm 2024, pg.48). A notable point of contention among scholars is the ongoing lack of consensus regarding the precise identification and nomenclature of these symbols—specifically, whether they represent frogs or toads (Phạm 2024, pg.44). This inconsistency in identification is entirely explainable. Firstly, from a visual standpoint, the shape of these symbols is not sufficiently distinctive, thereby hindering clear differentiation between toads and frogs. Secondly, within Vietnamese culture, both frogs and toads hold profound cultural symbolic significance, being intrinsically linked to the characteristics of wet-rice agriculture and the associated rain-praying beliefs. This also provides a foundational basis for understanding why only a single Proto Mon-Khmer reconstructed form exists for both ‘frog’ and ‘toad’. Numerous archaeological studies indicate that the discovery of Đông Sơn bronze drums (particularly Heger Type I) in areas of Southern China provides significant evidence for extensive trade and cultural exchange during and after the Đông Sơn culture. The widespread discovery of Heger Type I bronze drums and other artifacts across both Vietnam and the South China region has significant implications regarding cultural contact and the potential migrations of populations and language families in history. Trịnh Sinh (2024) posits the extensive spread of Đông Sơn culture throughout Southeast Asia and Southern China and suggests insightful assumptions regarding the ancient routes of cultural exchange based on the discovery of Đông Sơn significant artifacts. Among the routes of Đông Sơn culture diffusion he proposes, there is a northeast coastal route to Guangdong and Zhejiang that facilitated trade and cultural interaction and a route along the Red River extending southwest, fostering exchanges with the Dien culture in Yunnan.
T RẦN AND T RƯƠNG : A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘ FROG / TOAD ’ 45 Figure 1: The coastal exchange route of Dong Son culture extended up to Zhejiang, evidenced by: 1. Nine bronze situlae in the Nanyue King's Tomb; 2. Mingqi drums (funerary drums) in the Shangmashan Tomb; 3. Bronze situlae and axes in Zhaoqing; 4. Daggers with human-figure hilts from Shumuling (Trịnh Sinh 2024, pg. 41) Figure 2: The exchange routes of Dong Son culture with cultures in South China, evidenced by excavated and discovered Dong Son drums and situlae, include locations such as: 1. Shizhaishan; 2. Kaihua; 3. Azhangzhai; 4. Lijiashan; 5. Putuo; 6. Diandong; 7. Guixian; 8. Luobowan; 9. Xuzhuang; 10. Huili (Trịnh Sinh 2024, pg. 47) The diffusion of Dong Son culture, as demonstrated by archaeologists, suggests a close contact between speakers of Taic and Vietic languages. This interaction likely stemmed not only from Taic southward migration but also from Dong Son inhabitants' trade and exchange with people in the North. A hypothesis can be suggested: the exchange of artifacts, especially bronze drums featuring toad/frog symbols, facilitated the diffusion of Vietic phonological forms /kok/ and /kɔk/ into certain Taic languages. This diffusion subsequently led to the retention of these forms in some Taic dialects in Southern China today, as evidenced by earlier data.
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 46 Thus, building upon these archaeological and linguistic observations, a plausible scenario for language contact during the Dong Son period, particularly concerning the terms for 'toad' and 'frog', can be proposed: The Proto Mon-Khmer etymon for both 'toad' and 'frog' was a shared form *[ ]r[ɔ]k. As Mon-Khmer languages evolved, a distinction between 'toad' and 'frog' emerged in certain groups, though this differentiation wasn't always clear. By the Dong Son period, Mon-Khmer languages, primarily represented by the Vietic group, began linguistic contact with southwardmigrating Taic-speaking populations. Simultaneously, with the diffusion of Dong Son culture, some Vietic speakers also interacted with Taic speakers to the north, allowing the proto-Vietic form for 'toad' *roːk to spread to some Taic languages in that region. The shared initial consonants and syllable-final structures between Proto-Tai and ProtoVietic phonological forms likely led to their interchangeable use across various Taic dialects. However, this hypothesis remains tentative due to the limited Taic language data presented in this paper. Further in-depth research with a more comprehensive dataset is required to construct more detailed linguistic maps, which would enable a more definitive affirmation or refutation of this hypothesis. Moreover, Alves (2015) highlights imitative onomatopoeia as a notable phenomenon in the coinage of terms for various bird species in Asia, both wild and domesticated. For instance, 'crow', 'pigeon/dove', and 'owl' have names clearly related to their calls. This mechanism is presented as a reasonably intuitive process, also seen in the term for 'cat', represented as MAO or a similar sounding word across Chinese, Tai, Vietnamese, and other regional languages (Alves 2014, pg. 40). Given that imitative onomatopoeia represents a significant universal lexical innovation, it is thus proposed that the phonological similarity between the Vietic word form for 'toad' and 'frog' (e.g., /kok/, /kɔk/) and its counterparts in some Taic dialects could also stem from this phenomenon. This suggests that these shared forms might have originated from the imitation of the amphibian's call, rather than exclusively through direct language contact and diffusion. “Such sound-symbolism adds uncertainty to claims of borrowing in a certain direction, especially at time depths of thousands of years, unless there is sufficient clarifying linguistic and extralinguistic evidence” (Alves 2015, pg. 41). 5. Conclusion This paper investigates the possibility of linking geographical distribution and word form similarity to specific language interaction events, focusing on 'frog' and 'toad' terms. A preliminarily supported hypothesis proposed here suggests that the
TRẦN AND TRƯƠNG: A GEOGRAPHIC DISTRIBUTION OF THE WORD FORM FOR ‘FROG/TOAD’ 47 geographical distribution of 'frog' and 'toad' variants in Taic languages, which are closely aligned with or even interspersed among Mon-Khmer languages, especially Vietic variants, resulted from language and ethnic contact during the Đông Sơn period. However, due to the limitation of current data, this remains a tentative hypothesis requiring further evidence. Additionally, the potential for the phonological similarity between Taic and Vietic variants to stem from imitative onomatopoeia must also be considered. Future research, therefore, must aim to thoroughly map these linguistic variations to definitively clarify the complex interplay between population movements, cultural exchange, and the evolution of shared lexical forms in ancient Southeast Asia. References Alves, Mark (2009) Loanwords in Vietnamese. In Martin Haspelmath and Uri Tadmor (eds). Loanwords in the World's Languages: A Comparative Handbook: 617-637. De Gruyter Mouton. Alves, Mark (2015) Etyma for ‘Chicken’, ‘Duck’, & ‘Goose’ among Language Phyla in China & Southeast Asia. Journal of the Southeast Asian Linguistics Society 8:39-55. http://hdl.handle.net/1885/16086 Alves, Mark (2019) Dữ liệu liên ngành: chi Vietic kết nối với văn hóa Đông Sơn [Data from Multiple Disciplines Connecting Vietic with the Dong Son Culture]. Presentation at University of Social Science and Humanities, Vietnam National University of Hanoi. doi: https://doi.org/10.13140/RG.2.2.26690.22720 Alves, Mark (2022a) Lexical evidence of the Vietic household before and after language contact with Sinitic. Trang Phan, John Phan, and Mark Alves (eds) Vietnamese Linguistics: State of the field, JSEALS Special Publication No9 Vol.15 (4): 15-58. University of Hawai‘i Press. http://hdl.handle.net/10524/52500 Alves, Mark (2022b). The Ðông Sơn Speech Community: Evidence for Vietic. Crossroads. doi: https://doi.org/10.1163/26662523-bja10002 Alves, Mark (2023a) Proto-Austroasiatic Etymologies of Words Related to Household Structures. In Paul Sidwell & Mark Alves (eds) Proceedings of the 9th and 10th International Conferences of Austroasiatic Linguistics, JSEALS Special Publication No. 12: 1-18. University of Hawai‘i Press. http://hdl.handle.net/10524/52517 Alves, Mark (2023b) Preliminary Etymological Notes on Vietnamese Words for Pottery. In Alves, M. J., Lâm, Q. Đ., Trịnh, C. L., Trần, T. H. H., & Dương, X. Q. (eds) Researching and Applying Linguistics and Vietnamese Language Studies 1-22. Tokyo: Geolinguistic Society of Japan. Alves, Mark (2024) An Etymological Study of Vietnamese Words for Weaving and Woven Objects. In Phan, T., Nguyen, TC., Shimizu, M. (eds) Studies in Vietnamese Historical Linguistics. Global Vietnam: Across Time, Space and Community. Singapore: Springer. doi: https://doi.org/10.1007/978-981-97-4314-8_2
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 54 This quadrisyllabic form can be analysed as ‘open-or-labtse’ describing the butterfly’s wings with labtse-like antennae, as depicted in Figure 1. Although LT does not include a word form for ‘butterfly containing lab rtse or la btsas’, each morpheme of the quadrisyllabic form is independently attested in LT. 2.2. Other forms Two additional types are identified: one with reduplication syllable, sometimes linked to LT equivalents; the other is miscellaneous. The reduplicated type includes /ta pheː pheː/ (Lhagang), /mbẽ bo/ (Sangdam), and /pi bi tha ro/ (dGonpa).3 These examples contain the bilabial sounds /ph/ and /b/. This sound may reflect the iconicity of the action of flying, or to an onomatopoetic expression. Additionally, the form of sNyingthong, /ɕhә mә phɔʔ/, appears to be a hybrid of LT phye ma and the onomatopoetic bilabial sound. The miscellaneous type includes /ȵi mᴀ le le/ (Zhollam), /ʔa ku pɛ leʔ/ (sKobsteng), /ʔɐ kho peː mɐ/ (nKhorlo), /ndza mo la tsa/ (sTaglo), /xhɔ ɦba hta pa/ (dGonpa), /ʔa ɕhә n the ma ka/ (Thopa), /ɦgo xhɯ tɕiʔ leʔ/ (Gadnagshod), /mbɯ xu ra˞/ (Ongsum), inter alia. The forms of Zhollam, sKobsteng, and nKhorlo are loosely related. They belong to the Melung subgroup of Sems-kyi-nyila Tibetan: the syllable /le(ʔ)/, which may be related to the LT morpheme leb ‘flat’, occurs in the region between Zhollam and sKobsteng. The disyllabic morphemes /ʔa ku/ and /ʔɐ kho/, likely related to the word for ‘uncle’ (LT a khu), are found between sKobsteng and nKhorlo. 3. Mapping and discussions Based on the morphological analysis presented in Section 2, cartographic representations are examined according to the classification outlined below. A: phye ma leb - phye ma ka leb type (including LT equivalents of phye and leb) A1a: trisyllabic form without syllable coalescence e.g., ɕhe ma leʔ, ɕhә ma leː A1b: trisyllabic form with some modifications e.g., ptshaː bɯ leʔ, ɕhõ ɦa leʔ, ɕhә wã la A1c: coalescent form from A2a e.g., ɕhɑː ka ljɔʔ, ɕhɑː ka ljәʔ, ɕhaː kә lɔʔ A2a: quadrisyllabic form 3 The suprasegmental description is uniformly omitted when citing word forms.
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 55 e.g., ɕhә mã kә ljɔʔ, ɕha ma ka lɑʔ, ɕhe mɑː ɦa lɤʔ, sho mɔ ka leʔ A2b: quadrisyllabic form with some modifications e.g., ɕhe mɑː ɦa lɤʔ, ɕhɤ wã kә ljɤʔ, ʔa ma kә ljɯ B: phye ma lab rtse - phye ma la btsas type (connected to labtse) e.g., ɕhi ma low htse, xha mo la tse, ho mo la dza C: reduplicative forms including labial initials (see Section 2) D: miscellaneous (see Section 2) 3.1. Mapping the morpholofical structure Figure 4 shows a map reflecting the aforementioned classifications. Type A is represented by the black symbols. Type B is sky blue, whereas Types C and D have different shapes in red. Notably, Type D includes forms not classified in the other types; hence, the distribution of Type D does not represent any geolinguistic significance. Figure 4 indicates the red line dividing the target region into two parts. The northeastern region (upper part of the red line) is traditionally referred to as Amdo, and the south-eastern region (lower part of the red line) is named Khams. 4 An overall observation of the distribution of the three types (A, B, and C) is as follows: Type A is mainly found in Khams, and Type B is dominant in Amdo. Type C is concentrated in the central area of the map, that is, the easternmost area (Suzuki and Sonam Wangmo 2015) of the Khams region. It is noteworthy that Type B also forms a minority in the distribution area of Type A. Based on current reports (Suzuki and Sonam Wangmo 2016, 2017, 2019; Suzuki 2018, 2022), these are exclaves of Amdo Tibetan with oral histories which can relate the speakers’ ancestors to their homeland in Amdo. The distribution of Type B in Khams denotes the retention of the lexical item for ‘butterfly’, showing neither influence from, nor exerting influence on, neighbouring Khams varieties. Next, we focus on the distribution of Type A subclasses in the Khams region. Types A1a and A1b (trisyllabic forms) are predominant in northern Khams, whereas the others are mainly observed in southern Khams. The phenomenon of Type A1c belongs to a phonological change (coalescence of a LT suffix po, pa, mo, or ma and its preceding syllable), and hence does not appear only in this lexical item. Varieties with this phonological change are distributed in a given area (sPomborgang, Chaphreng, and Sems-kyi-nyila). The map indicates that Types A1a, A1c, A2a, and A2b coexisted in the central Khams area. 4 For a more detailed information of the traditional description and language classification, refer to Ryavec (2015) and Tournadre and Suzuki (2023).
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 56 Figure 4: Map based on the morphological structure
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 57 Figure 5: Map based on the syllable number
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 58 3.2. Mapping the number of syllables Figure 5 indicates the number of syllables in word forms for ‘butterfly’, regardless of lexical type, ranging from two to five. Disyllabic forms appear centrally within a limited area, bordered by trisyllabic and quadrisyllabic forms, with some exceptions in Southern Khams. Trisyllabic forms span a broad central area, overlapping with quadrisyllabic forms, which dominate the eastern Tibetosphere. Quadrisyllabic forms are also found in both northern and southern regions, separated by disyllabic and trisyllabic zones. A cinquisyllabic form is found in isolation, in Thopa (/ʔa ɕhә n the ma ka/). Considering, language classification, quadrisyllabic forms, along with disyllabic and trisyllabic forms in the central areas, are found in Amdo Tibetan varieties (cf. the distribution of Type B in Figure 4). However, trisyllabic and quadrisyllabic forms also co-occur in several Khams Tibetan subgroups. Hence, it is essential to interpret the distribution in light of language affiliations and internal developments. 3.3. Discussions Based on Figures 4 and 5, the following observations are made: language-level tendency, the ABA (concentric) distribution of triand quadrisyllabic forms is observed in Khams, reflecting a relationship between the trisyllabic and quadrisyllabic forms of Type A. The language-level tendencies are conceptualised as follows: Amdo Tibetan varieties predominantly exhibit Type B forms. Even in in Khams speaking areas (the central region of the maps), Type B is retained. In contrast, Khams Tibetan varieties display disyllabic, trisyllabic, or quadrisyllabic forms of Type A, with no attestation of Type B. In relation to syllable number, the ABA distribution of triand quadrisyllabic forms is attested in Khams. As indicated in Figure 5, the quadrisyllabic form appeared later in Khams. This suggests that Amdo-speaking people later came to the Khams area, which is historically correct (Suzuki and Sonam Wangmo 2019). Based on the observations in Section 3.1, the relationship between the trisyllabic and quadrisyllabic forms of Type A, that is, A1c and A2a, is recognised. By definition, the trisyllabic form A1c is derived from the coalescence of two syllables within the quadrisyllabic form A2a. The condition for the coalescence of two syllables is that the second syllable is equivalent to the LT suffixes pa, po, ma, mo, etc. For the lexical form for ‘butterfly’, the first two syllables are phye ma, which fulfils a condition of the coalescence. This phenomenon has been attested in many Tibetic varieties, from Southern Khams to mNga’-ris (Jiang 2002:70–76). Particularly, the Southern Khams
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 59 varied across varieties. Coalescence does not always appear in each variety in a regular manner; some have many words with coalescence, whereas others do not. This observation suggests that coalescent forms for ‘butterfly’ display notable geographical continuity, as shown in Figure 5. This highlights the need to geolinguistically analyse each word from exhibiting syllabic coalescence. 3.4. Notes from the viewpoint of motivation Motivation is a method of cross-linguistically effective geolinguistic analysis (Del Giudice & Brun-Trigaud 2024). From this standpoint, a brief analysis of Tibetic, Sinitic, and Japanese forms for ‘butterfly’ is provided. Motivation is analysed by comparing morpheme-level meanings of the word forms. As provided in Section 2, Tibetic word forms includes the following semantic units: ‘open’ (phye), ‘flat’ (leb), ‘lid’ (kha leb), and ‘labtse’ (la btsas). Of them, LT leb ‘flat’ is cognate with Sinitic 蝶 die ‘butterfly’. Its Old Chinese form is *lʕep (Baxter and Sagart 2014), and its Middle Chinese form is depD, which was actually borrowed into Old Japanese てふ tehu. Its reduplication form (てふてふ tehutehu) is the modern spoken form ちょうちょう tyootyoo. Therefore, from a viewpoint of motivation, the forms for ‘butterfly’ in Tibetic, Sinitic, and Japanese include a common semantic feature; however, its original meaning has been lost. Another noteworthy feature is the use of labial sounds, such as /p/ and /(m)b/, which may reflect onomatopoeia iconicity of wing movement. Indeed, searching for word forms for ‘butterfly’ in Tibeto-Burman languages with the STEDT database (Matisoff 2015), we find many forms including /p/ with a Proto-Tibeto-Burman form pur ⪤ pwar (#355; BUTTERFLY). A detailed discussion is beyond the scope of this study, but future research could explore sound iconicity as part of the motivation in constructing lexical forms for ‘butterfly’. 5. Concluding remarks This article examined the word form for ‘butterfly’ in Tibetic languages of the eastern Tibetosphere by analysing variation in lexical forms and syllable count. It then presented linguistic maps for both features and offers an interpretive analysis of their distributions. Three principal lexical forms are recognised: phye ma leb - phye ma ka leb type, phye ma lab rtse - phye ma la btsas type, and reduplicative forms, including labial initials. The first two types are further characterised by the number of syllables. However, the number of syllables is not interpreted as reflecting ABA distribution for
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 60 the temporal differentiation, but rather as resulting from inherited sound change tendencies within a each variety. This study broadens the perspective on lexical motivation by showing that Tibetic forms for ‘butterfly’ often include the LT morpheme leb ‘flat’, which may be cognate with a Chinese word for ‘butterfly’ and possibly related to a Japanese equivalent. Additionally, the frequent use of the labial initials suggests onomatopoetic iconicity linked to wing movement. References Baxter, William H. & Laurent Sagart (2014) Baxter-Sagart Old Chinese reconstruction, version 1.1. Unpublished manuscript, retrieved from https://sites.lsa.umich.edu/ocbaxtersagart/ Del Giudice, Philippe & Guylaine Brun-Trigaud (2024) Linguistic motives common to Japanese and Romance dialects: Two examples with maps. Studies in Geolinguistics 4: 20–39. https://doi.org/10.5281/zenodo.13948553 Geng, Xianzong, Junying Li, and Lhun-grub rDo-rje (2007) Anduo Zangyu kouyu cidian《安多 藏语口语词典》[Dictionary of spoken Amdo Tibetan]. Lanzhou: Gansu Minzu Chubanshe. Hua, Kan and Klu-’bum-rgyal (1993) Anduo Zangyu kouyu cidian《安多藏语口语词典》 [Dictionary of spoken Amdo Tibetan]. Lanzhou: Gansu Minzu Chubanshe. Jiang, Di [江荻] (1990) Zangyu yuyinshi yanjiu《藏语语音史研究》[Study on Tibetan sound history]. Beijing: Minzu Chubanshe. Matisoff, James A. (2015) The Sino-Tibetan etymological dictionary and thesaurus. Berkeley: The Regents of the University of California. Data are also available as online STEDT Database. http://stedt.berkeley.edu/~stedt-cgi/rootcanal.pl Ryavec, Karl E. (2015) A historical atlas of Tibet. Chicago: University of Chicago Press. https://doi.org/10.7208/chicago/9780226243948.001.0001 Suzuki, Hiroyuki [铃木博之] (2018) Litangxian ji qi zhoubian Zangzu yuyan xianzhuang diaocha yu fenxi 理塘县及其周边藏族语言现状调查与分析 [Research and analysis of the current statud of Tibetans’ languages in Lithang County and its surroundings]. Minzu Xuekan 2: 35–44+106–109. https://doi.org/10.3969/j.issn.1674-9391.2018.02.005 Suzuki, Hiroyuki (2022) Amdo Tibetan-speaking Khampas in Lithang: Their language, identity, and migration history. Paper presented at 16th Seminar of the International Association for Tibetan Studies (Praha). https://doi.org/10.13140/RG.2.2.24354.50887 Suzuki, Hiroyuki (2024) Evolution of dorsal fricatives in rGyalthangic varieties of Khams Tibetan. Journal of the Phonetic Society of Japan 28(2): 107–118. https://doi.org/10.24467/onseikenkyu.28.2_107
SUZUKI: GEOLINGUISTIC PATTERNS OF THE WORD FORM FOR ‘BUTTERFLY’ IN TIBETIC LANGUAGES 61 Suzuki, Hiroyuki & Sonam Wangmo (2015) Challenge to discover endangered Tibetic varieties in the easternmost Tibetosphere: a case study on Dartsendo Tibetan. Linguistics of the Tibeto-Burman Area 38(2): 256–270. https://doi.org/10.1075/ltba.38.2.07suz Suzuki, Hiroyuki & Sonam Wangmo (2016) Vocabulary of Shingnyag Tibetan: A dialect of Amdo Tibetan spoken in Lhagang, Khams Minyag. Asian and African Languages and Linguistics (AALL) 11: 101–127. https://doi.org/10.15026/89211 Suzuki, Hiroyuki & Sonam Wangmo (2017) Language evolution and vitality of Lhagang Tibetan: a Tibetic language as a minority in Minyag Rabgang. International Journal of the Sociology of Language 245: 63–90. https://doi.org/10.1515/ijsl-2017-0003 Suzuki, Hiroyuki & Sonam Wangmo (2019) Migration history of Amdo-speaking pastoralists in Lhagang, Khams Minyag, based on narratives and linguistic evidence. Archiv Orientální Supplementa XI: 203–222. Tournadre, Nicolas & Hiroyuki Suzuki (2023) The Tibetic languages: An introduction to the family of languages derived from Old Tibetan. LACITO Publications. https://doi.org/10.5281/zenodo.10026628
Proceedings of the sixth International Conference of Asian Geolinguistics 62 Extraction of regularities and geographical patterns from the basic vocabulary of the Ainu language Mika Fukazawa (National Ainu Museum) Abstract: This paper demonstrates how to extract “regularities,” such as phonological correspondences and morphological and grammatical rules, from basic vocabulary items and create relational maps in the Ainu language. Studies on basic Ainu vocabulary (Asai 1974), are based on Hattori & Chiri’s (1960) list of 200 items taken from Swadesh’s (1954, 1955) personal letters to Hattori (cf. Hattori 1954). This basic vocabulary contains information about lexical forms as well as various lexical correspondences and rules. For instance, the regularity of the semivowel (glide) type, - iw-:-uy, can be extracted from the items ciw:cuy for ‘pierce (stab)’ and nociw, noociw:nocuy for ‘star’ (Fukazawa & Ono 2025, Ono & Fukazawa 2025). When drawing a regularity map such as a lexical form map, it may be suggested that several lexical forms are integrated into one symbol with the corresponding regularity. However, as integrating symbols may exclude dialectal details, this study proposes drawing relational maps between different basic lexical maps. Key words: basic vocabulary, Ainu, phonological correspondences, lexical forms 1. Ainu language and the dialects The Ainu language was (or still is, as a second language) spoken in Hokkaido, Sakhalin, the Kuril islands, and the northern part of mainland Japan. It is a language isolate that is typologically different from Japanese and other Asian languages (cf. Tamura 2000; Bugaeva 2022). Recently, it has been designated as an endangered language in Japan. It is now impossible to conduct a large-scale survey of the Ainu vocabulary. Even in 1960, a linguist of the Ainu language, Hattori Shirō, reported that some of his informants were among the last native speaker(s) of the Ainu dialect, and that they were very old. He also mentioned that some could speak the Ainu language fluently, while others knew only several words (Hattori & Chiri 1960). FUKAZAWA, Mika. 2025. Extraction of regularities and geographical patterns from the basic vocabulary of the Ainu language. In Nor Hashimah Jalaluddin, Hiroyuki Suzuki and Mitsuaki Endo (eds.) Proceedings of the sixth International Conference of Asian Geolinguistics, 62–80. doi: https://doi.org/10.5281/zenodo.17204595 This work was supported by JSPS KAKENHI Grant Numbers JP25K04114, JP23K25322. We would like to thank Editage (www.editage.jp) for English language editing.
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 63 When studying the Ainu dialects in the present, we have no choice but to use materials from the Ainu language of the past. The most reliable source of data for considering the classification of the Ainu dialects is the 200 (Asai’s 202)1 basic vocabulary items of the Ainu language, left by Hattori & Chiri (1960) and Asai (1974). Hattori & Chiri’s (1960) data was mostly corrected by Hattori Shirō, Chiri Mashiho, and some collaborators from 1955 to 1956, and the data of Asai (1974) was based on the Hattori & Chiri (1960), his original survey of Asahikawa, Obihiro, and Kushiro dialects, and other written materials of Torii (1903), Murayama (1971), and Pinart (1872). Asai (1974) modified the data of Hattori & Chiri (1960) and added two dialects, Chitose and northern Kuril. See Table 1 and Figure 1 for the Ainu dialects used by Hattori & Chiri (1960) and Asai (1974). Table 1: Ainu dialects Hokkaido dialects 1. Yakumo, 2. Oshamambe, 3. Horobetsu, 4. Biratori, 5. Nukibetsu, 6. Niikappu, 7. Samani, 8. Obihiro, 9. Kushiro, 10. Bihoro, 11. Asahikawa, 12. Nayoro, 13. Soya, 21. Chitose Sakhalin dialects 14. Ochiho, 15. Tarantomari, 16. Maoka, 17. Shiraura, 18. Raichishika, 19. Nairo Northern Kuril dialect 20. Shumushu Figure 1: Map of the Ainu dialects2 1 Asai (1974) divided the two basic vocabulary items of Hattori & Chiri (1960) into two respectively, increasing the total number of items from 200 to 202. 2 The numbers correspond to those in Table 1.
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 70 3. Ways to draw regularity maps from basic vocabulary items In this section, we will attempt to consider a method for drawing a regularity map of basic vocabulary items with the same patterns and characteristics. Basic vocabulary items contain information about lexical forms as well as various lexical correspondences and rules. The following examples use the basic vocabulary data from Hattori & Chiri (1960) and Asai (1974). 3.1. Regularity of the semivowel (glide) type The regularity of the semi-vowel (glide) type -iw:-uy can be extracted from the items of (4). (4) Regularity of semi-vowel (glide) types: -iw > -uy (metathesis form of -iw) (a) Lexical forms of ‘star’ (No. 74): nociw~noociw:nocuy (b) Lexical forms of ‘pierce (stab)’ (No. 179): ciw:cuy Figure 7 shows the lexical forms of ‘star’ (No. 74). The lexical form rikop can be analyzed as rik, the noun ‘sky,’ o, the transitive verb ‘be in,’ and -p, the nominalizer of ‘thing,’ mening ‘thing in the sky.’ This form is probably newer than keta or nociw, although there is some controversy regarding which is older, keta or nociw (Fukazawa 2017: 100). The semivowel (glide) type -iw:-uy can be found in the lexical forms nociw, noociw, and nocuy. The -uy form of nocuy exists in the Kushiro dialect. Figure 8 shows the lexical forms of ‘pierce (stab)’ (No. 179). Hattori (1964:146) recorded the lexical form eciekara in Asai (1974) as the -iw form eciwkara. The form of ciri in Asai (1974) appears to be a miswriting of civi, which can be assumed to be the -iw form of /ciwi/ (cf. Murayama 1971: 152). Thus, two distinct lexical formtypes are observed: {otke} and {ciw}. However, they may suggest a historical convergence from distinct semantic origins since {ciw} can also express the meaning of ‘to ripple.’ The semivowel (glide) type -iw:-uy can be found in the lexical forms ciw, ciwi, and cuy. The -uy form of cuy exists in the Bihoro dialect. The regularity map in Figure 9 is derived from Figures 7 and 8. The -uy form is found only in eastern Hokkaido dialects, such as Kushiro and Bihoro.
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 71 Figure 7: Maps of the lexical forms of ‘star’ (No. 74) Figure 8: Maps of the lexical forms of ‘pierce (stab)’ (No. 179) nociw~noociw~nocuy keta rikop ciw~ciwi~cuy eciwkara otke~otki~okke nociw~noociw nocuy ciw~ciwi~ eciwkara cuy
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 72 Figure 9: Regularity map of the semivowel (glide) type, -iw:-uy 3.2. Regularity of the consonant (C1C2:C2C2:hC2) type The lexical forms of ‘pierce (stab)’ (No. 179) also contain the consonant (C1C2:C2C2:hC2) regularity, which is related to the other basic items, such as those for ‘rain’ (No. 76), as in (5). (5) The regularity of the consonant (C1C2:C2C2:hC2) type (a) Lexical forms of ‘rain’ (No. 76): apto:atto:ahto (b) Lexical forms of ‘pierce (stab)’ (No. 179): otke~otki:okke (c) Lexical forms of ‘lie’ (No. 67): hotke:hokke Figure 10 shows the maps of the lexical forms of ‘rain’ (No. 76). The lexical form weni in the westernmost Hokkaido dialects, Yakumo and Oshamambe, seems to be derived from the word wen for ‘bad’ (cf. Fukazawa 2021: 289). The lexical form sirun in Shumushu corresponds to sirwen ‘bad weather’ in the Hokkaido dialects. In the eastern Hokkaido dialects, the lexical forms ruyanpe and ruwanpe are used for ‘rain,’ while in the western Hokkaido dialects, they mean ‘storm,’ and the lexical form {apto} is used for ‘rain.’ The lexical form of C1C2 or C2C2, apto or atto, is also expected to exist in the eastern Hokkaido dialects but cannot be found. -iw -uy
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 73 Figure 10: Maps of the lexical forms of ‘rain’ (No. 76) Figure 11: Maps of the lexical forms of ‘pierce (stab)’ (No. 179) apto~atto~ahto ruyanpe~ruwanpe weni sirun ciw~ciwi~cuy eciwkara otke~otki~okke apto atto ahto otke~otki okke
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 74 Figure 11 shows the maps of the lexical forms of ‘pierce (stab)’ (No. 179), one of which was shown in Figure 8. The C1C2 forms, otke and otki, are found in western Hokkaido and northern Kuril Islands. The C2C2 form okke is found in the southernmost Hokkaido dialect, Samani. Figure 12 shows the maps of the lexical forms of ‘lie’ (No. 67), which also belongs to the consonant regularity, C1C2:C2C2:hC2. The C1C2 and C2C2 forms of {hotke}, hotke and hokke, are widely distributed in the Hokkaido dialects. The hokke form can be found in northeastern Hokkaido dialects, including Asahikawa, Obihiro, Kushiro, and Bihoro. Many dialects use both the forms of various actions. The lexical forms of {hotke} mean ‘to lie down,’ while the forms of {situri} are used to mean ‘to stretch out.’ According to Hattori (1964), in a Sakhalin dialect, Raichishika, the hC2 form of {hotke}, hohke, means ‘to go to bed and lie down.’ Figure 12: Maps of the lexical forms of ‘lie’ (No. 67) Figure 13 shows the regularity map derived from the lexical maps shown in Figures 10, 11, and 12. The C1C2 form is found in the Hokkaido and northern Kuril dialects, C2C2 form is found in the northeastern Hokkaido and a few Sakhalin dialects, and the hC2 form is found only in Sakhalin dialects. These dialectal features originate from at least three basic vocabulary items. They compensate for each other. hotke~hokke situri nokataki hotke hokke
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 75 Figure 13: Map of the the consonant regularity (C1C2:C2C2:hC2) type 4. Relational map among basic vocabulary items This section will attempt to draw a relational map by combining the different basic lexical maps with a regularity map. Regularity maps such as those in Figures 9 and 13 exclude lexical information. However, they are quantitatively layered maps, also known as “phonological maps.” Lexical maps trace the histories of individual vocabulary items. This study proposes that the necessary information can be easily obtained from each type of map, although this sometimes requires the exclusion of unnecessary details. Relational maps, or map relations, fill in the missing information from both regularity and lexical maps. For example, when creating a relational map with Figures 7, 8, and 9, the regularity map of the semi-vowel (glide) type, -iw:-uy (Figure 9), is placed in the center and combined with the lexical maps of Figures 7 and 8. The relational map is shown in Figure 14. Similarly, a relationship map (Figure 15) can be created by placing the C1C2:C2C2:hC2 consonant regularity map (Figure 13) in the center and combining it with the lexical maps shown in Figures 10, 11, and 12. As Figures 8 and 11 show, the C 1C2 C2C2 hC2
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 76 lexical maps of ‘pierce (stab)’ are related to the two regularity maps. One lexical map can be associated with several regularity maps, thus constituting another relational map. As noted above, a relational map can show the lexical form that would be inferred if it had not been recorded. For example, the lexical form hohke for ‘lie’ was not recorded in Hattori & Chiri (1960) and Asai (1974), but Hattori (1964) mentioned it in the Raichishika dialect of Sakhalin. Given the relational map between consonant (C1C2:C2C2:hC2) regularity and lexical maps, the lack of data would be compensated for. Since the Hokkaido dialects have the hotke (C1C2) and hokke (C2C2) forms, one can infer that the Sakhalin dialects have the hC2 form, hohke. Considering which regularity map each lexical form is associated with is useful to trace the history of a lexical item. Figure 14: Relational Map of the semivowel (glide) type, -iw:-uy Figure 9 -iw -uy Figure 7 ‘star’ Figure 8 ‘pierce (stab)’
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 77 Figure 15: Relational Map of the consonant (C1C2:C2C2:hC2) type Figure 10 ‘rain’ Figure 11 ‘pierce (stab)’ Figure 13 C 1C2 C2C2 hC2 Figure 12 ‘lie’
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 78 5. Concluding remarks This study proposes the creation of relational maps of basic vocabulary items in Ainu. Fukazawa’s (2018) preliminary study started with the first 100 basic vocabulary maps but found that they often had a monotonous distribution. One solution to this issue, as proposed in this study, is to distinguish between regularity and lexical maps and create relational maps. The regularity map used in this study is similar to the “phonological map” in structural geolinguistics. Future studies will also aim to create grammatical and morphological maps as regularity maps, such as those showing how to combine firstand second-person affixes with transitive verbs. This study also proposes drawing a connection between regularity and lexical maps. A lack of information on a map is acceptable because if the symbols are appropriately assigned and the relations between the maps are distinctly created, they can be more easily interpreted. Relational maps demonstrate the relationships between regularity and lexical maps, and can provide additional information if word recordings are lacking in the linguistic survey. The histories of vocabulary items in the regularity and lexical maps differ significantly. Lexical maps frequently show more complex distributions because words trace their own histories, including borrowings. Regularity maps may show dialectal categories extracted from quantitative vocabulary items. Relational maps possess the advantages of both types, and are useful for tracing word histories and identifying trends in linguistic regularities. Abbreviations 1: 1st person, A: transitive subject, C: consonant, EXCL: exclusive (personal affix). PL: plural, V: vowel. References Asai, Toru (1974) Classification of dialects: Cluster analysis of Ainu dialects. Bulletin of the Institute for the Study of North Eurasian Cultures, Hokkaido University 8: 45–136. Bugaeva, Anna (2022) Introduction. In: Anna Bugaeva (ed.) Handbook of the Ainu language, 1–19. Berlin: Mouton De Gruyter. doi: https://doi.org/10.1515/9781501502859-001 Chiri, Mashiho (1976 [1962]) Bunrui Ainugo jiten, Dōbutsuhen『分類アイヌ語辞典 動物 篇』[Classified Ainu dictionary]. Reprinted: Chiri Mashiho chosakushū bekkan 『知里真
FUKAZAWA: EXTRACTION OF REGULARITIES AND GEOGRAPHICAL PATTERNS FROM THE BASIC VOCABULARY 79 志保著作集 別巻1』[Collected works of Chiri Mashiho Separate volume 1]. Tokyo: Heibonsha. Fukazawa, Mika [深澤美香] (2017) Kagake monjo ni okeru Ainugo no bunkengakuteki kenkyū 『加賀家文書におけるアイヌ語の文献学的研究』[Philological study of the Ainu in Kaga family’s archives]. Doctoral dissertation. Chiba: Chiba University. doi: https://doi.org/10.20776/103627 ; URL: https://opac.ll.chiba-u.jp/da/curator/103627/ Fukazawa, Mika (2018) Geographical distribution patterns of basic Ainu vocabulary in Hattori and Chiri (1960). In. Suzuki, Hiroyuki & Mitsuaki, Endo (eds.) Studies in Asian Geolinguistics, Monograph Series No. 4: Papers from the Fourth International Conference on Asian Geolinguistics, 91–103. URL: https://publication.aaken.jp/papers_4IC_Asian_geolinguistics_2018.pdf Fukazawa, Mika (2021) ‘It rains’ in Ainu. In. Endo, Mitsuaki, Makoto Minegishi, Satoko Shirai, Hiroshi Suzuki, and Keita Kurabe (eds.) Linguistic Atlas of Asia. 289–290. Tokyo: Hituzi Syobo. Fukazawa, Mika [深澤美香] (2025) Kokuritsu Ainuminzoku hakubutsukan shozō Chiri Mashiho kinyu no Ainugo kisogoi chosahyō「〔資料紹介〕国立アイヌ民族博物館所蔵 知里真志保記入のアイヌ語基礎語彙調査表」[Preliminary survey of the basic Ainu vocabulary: Chiri Mashiho’s materials from the National Ainu Museum]. National Ainu Museum Journal 3: 136–158. doi: https://doi.org/10.57545/namjournal.2024-06 Fukazawa, Mika [深澤美香] & Yōhei Ono [小野洋平] (2025) Kizogoi niyoru hōgen bunrui no shomondai: Asai (1974) no Ainugo hōgen dēta wo saidaigen katsuyōsuru tameni「基礎語 彙による方言分類の諸問題:Asai (1974) のアイヌ語方言データを最大限活用する ために」[Challenges of dialect classification in extracting maximum information from the basic vocabulary of Ainu in Asai (1974)]. Northern Language Studies 15: 141–163. doi: https://doi.org/10.14943/112975 Hattori, Shirō [服部四郎] (1954) “Gengonendaigaku” sunawachi “goitōkeigaku” no hōhō nitsuite「「言語年代學」即ち「語彙統計學」の方法について:日本祖語の年代」[On the method of Glottochronology and the time-depth of Proto-Japanese]. Gengo Kenkyu [Journal of the Linguistic Society of Japan] 26, 27: 29–37. doi: https://doi.org/10.11435/gengo1939.1954.26-27_29 Hattori, Shirō [服部四郎] and Mashiho Chiri [知里真志保] (1960) Ainugo shohōgen no kisogoi tōkeigakuteki kenkyū 「アイヌ語諸方言の基礎語彙統計学的研究」 [Alexicostatistic study on the Ainu dialects]. Japanese Journal of Ethnology 24(4): 307– 342. doi: https://doi.org/10.14890/minkennewseries.24.4_307 Murayama, Shichirō [村山七郎] (1971) Kitachishima Ainugo『北千島アイヌ語』[Northern Kuril Ainu]. Tokyo: Yoshikawa Kobunkan. Nakagawa, Hiroshi & Mika Fukazawa (2022) Hokkaido Ainu dialects: Towards aclassification of Ainu dialects. In. Anna Bugaeva (ed.) Handbook of the Ainu language, 253–328. Berlin: Mouton De Gruyter. doi: https://doi.org/10.1515/9781501502859-009 Ono, Yōhei [小野洋平] and Mika Fukazawa [深澤美香] (2023) Ainugo shohōgen no gokei no ruiji ni kansuru kiso dētano fukugen「アイヌ語諸方言の語形の類似に関する基礎デー タの復元:論文に書ききれなかった研究者の判断・思考に迫る」[Reconstruction of
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 86 to surrounding areas in a uniform and continuous manner. The degree of influence diminishes with increasing distance from the center. On dialect maps, the Wave Model is represented by isoglosses — lines that demarcate linguistic boundaries — shaped like concentric circles or curves radiating from a point. Areas close to the center exhibit linguistic changes more clearly and strongly, while more distant areas show weaker influence. This model effectively explains transitional dialect regions, where linguistic features change gradually over space rather than along abrupt boundaries. Later, in 1973, the model was further developed by Charler Bailey to better accommodate findings in modern sociolinguistics (Bailey 1973). In Bailey’s version, wave models no longer appear strictly as ideal concentric circles. Instead, they account for more complex, directional diffusion patterns, especially when natural or social barriers (e.g., mountain ranges, rivers, political upheavals, or borders) interrupt the spread at certain points. These obstacles can cause the wave to propagate in a single direction. This evolution made Bailey’s wave model more dynamic and explanatory (Trịnh Cẩm Lan 2008). 2.2.2. Social Network Model Developed by Lesley Milroy in her study of Belfast, the Social Network Model shifts the focus from macro-level factors such as population and geographic distance to micro-level interactions among individuals within a community (Milroy 2002). The key principle of this model is that the spread and maintenance of linguistic variants depend on the social network structure of speakers. There are two main types of networks: A closed network is one with high density, where members all know each other and have few external ties. Such networks tend to reinforce internal linguistic norms and resist external influence. An open network, by contrast, has lower density and more external connections. These networks serve as effective channels for the introduction and spread of new linguistic variants (Milroy 2002). On dialect maps, this model helps explain the presence of dialectal "islands" or highly conservative linguistic areas even when they are located near innovative urban centers. For example, a rural community with a tightly-knit social structure (based on dense kinship and neighborhood relations) may preserve traditional dialect features despite being near a major city. Conversely, a social group with an open network — such as university students from various regions — forms an ideal environment for the emergence and diffusion of new linguistic variants, such as internet language or student slang.
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 87 3. Data and Research Methods The subjects of this survey are the phonetic variants of the initials /l/ and /n/ across 11 provinces in the Red River Delta. The investigators, who also served as data providers, were students from the University of Social Sciences and Humanities (majoring in Linguistics) and the University of Education (majoring in Literature Education), both part of Vietnam National University, Hanoi. These students were trained as field investigators to observe the pronunciation habits of residents in the localities where they were born and raised, and to record the findings in a questionnaire. The geographical scope of the survey includes provinces and cities in the Red River Delta, as well as a few neighboring provinces for comparative purposes. In each province/city, data was collected — where possible — from each district, town, or city under provincial administration, with an average of one sampling point per administrative unit. The total number of surveyed points was 170, covering 14 provinces (including 11 Red River Delta provinces and 3 neighboring provinces — Hòa Bình, Phú Thọ, and Bắc Giang — for comparative analysis). Map 1: The Red River Delta (Source: Vietnam Law Library, https://thuvienphapluat.vn/)
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 88 In addition to the data provided by student investigators, the author also referred to the results of the Vietnamese Comprehensive Linguistic Survey Program 1998–2000, Audio Field Recordings, currently archived at the Institute of Linguistics (Vietnam), and held discussions with the original investigators of the program to verify and refine the research findings. For certain localities where the collected data was incomplete or potentially biased, the researcher personally visited the area to conduct direct observation and data collection. Data Analysis Methods: The collected data — including linguistic data (phonetic variants of /l/ and /n/ in each locality) and geographical data (longitude and latitude of each survey point) — was encoded and entered into Excel. The dialect maps were then generated using ArcGIS Online, a web-based geographic information system developed by Esri, following the guidelines of Mitsuaki Endo (2016) and Mika Fukazawa (2017). 4. Research Findings 4.1. Patterns of /l/ and /n/ Variation in the Red River Delta In the Red River Delta, the initial /n/ exhibits two variants. The first is [n], an alveolar nasal stop, which is the standard, officially recognized pronunciation of /n/. The second variant is [l], an alveolar lateral approximant. Similarly, the consonant /l/ also has two variants: the standard [l], and a variant pronounced as [n]. Survey results show that the patterns of /l/–/n/ confusion are quite complex. The realizations of /l/ and /n/ across the surveyed regions present several key features: * In terms of frequency: For /l/, data indicates that in 84 out of 170 localities, /l/ has two variants, while in 86 out of 170 localities, /l/ has only one variant. For /n/, 61 out of 170 localities show two variants of /n/, whereas 109 localities have only one variant. Thus, overall, the shift /l/ > [n] appears to be more prominent than /n/ > [l]. * In terms of regional variation patterns, the data identifies several models of variation: - Model A: In some localities, both /l/ and /n/ have only one unique variant each, maintaining the phonological contrast between them. There is no confusion. - Model B: /n/ has two variants ([n] and [l]), while /l/ remains unchanged with only the [l] variant. This is a unidirectional confusion model: /n/ > [l], with /l/ remaining stable.
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 89 - Model C: /l/ has two variants ([l] and [n]), while /n/ has only one variant ([n]). This is also a unidirectional confusion model, but in the opposite direction: /l/ > [n], with /n/ remaining stable. - Model D: Both /l/ and /n/ exhibit two variants. This represents bidirectional confusion: /n/ > [l] and /l/ > [n]. Table 2. Regional Models of /l/–/n/ Confusion Model /n/ Variant(s) /l/ Variant(s) Description No. of Localities A [n] [l] No confusion 57 B [n], /n/ > [l] [l] Unidirectional confusion: /n/ > [l] 32 C [n] [l], /l/ > [n] Unidirectional confusion: /l/ > [n] 52 D [n], /n/ > [l] [l], /l/ > [n] Bidirectional confusion: /n/ > [l], /l/ > [n] 29 In terms of degree of confusion, the data shows that there is no case of complete confusion (i.e., “everyone in the village mispronounces”) for both phonemes in any locality. In all surveyed location, the confusion appears in a patchy or leopard-skin pattern — that is, a portion of the population maintains clear /l/–/n/ distinction, while others exhibit confusion in either one direction (unidirectional) or both (bidirectional). At the individual level, there are: (1) Speakers who clearly distinguish /l/ and /n/ (no confusion); (2) Speakers who fully confuse them in one direction (e.g., consistently pronouncing both /l/ and /n/ as either [l] or [n]); (3) Speakers who show partial or inconsistent confusion, where some words are affected while others are not. Preliminary observations suggest that confusion tends to concentrate among specific social groups: the elderly, young children, farmers, individuals with lower levels of education, or those with limited social interaction. Similar patterns have also been observed in dialect communities of Southern China (Ng, Choi Lee Charlie, 2017; Yuyan Zhou, Ying Wu, 2019).
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 90 4.2. Distribution of /l/ and /n/ Variants in the Red River Delta 4.2.1. Distribution of /l/ Variants Map 2. Distribution of /l/ Variants Regions with a single variant [l]: 86 localities; Regions with two variants [l] and /l/ > [n]: 84 localities) Map 2 shows that regions with only one variant of /l/ tend to be concentrated in the northwestern part of the area, while regions with two variants are mostly found in the southeastern area, particularly in the northern coastal provinces. Survey data reveals several regionally or provincially distinctive centers: * Regions mainly exhibiting two variants (i.e., regions with confusion): southern districts of Hanoi; provinces of Hải Dương, Hưng Yên, Hải Phòng, Hà Nam, Nam Định, Thái Bình, and Ninh Bình. * Regions mainly exhibiting a single variant (i.e., no or minimal confusion): urban districts of Hanoi, northern and northwestern districts of Hanoi, Bắc Ninh, and parts of Vĩnh Phúc. * Bordering provinces of the Red River Delta such as Hòa Bình, Phú Thọ, and Bắc Giang exhibit only one variant, meaning the phenomenon /l/ > [n] does not occur there.
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 91 4.2.2. Distribution of /n/ Variants Map 3. Distribution of /n/ Variants Regions with a single variant [n]: 109 localities; Regions with two variants [n] and /n/ > [l]: 61 localities. According to Map 3, regions with a single variant of /n/ tend to cluster in the southeastern part of the area, while regions with two variants are more concentrated in the northwest, particularly in provinces bordering hilly terrain, such as Bắc Giang, Hòa Bình, and Phú Thọ. The survey data reveals several regional/provincial centers for /n/ variation: * Regions with a tendency toward two variants (i.e., with confusion): northwestern districts of Hanoi, Bắc Ninh, and parts of Vĩnh Phúc. * Regions mainly exhibiting a single variant (i.e., no or minimal confusion): urban districts of Hanoi; southern districts of Hanoi; provinces of Hải Dương, Hưng Yên, Hải Phòng, Nam Định, Thái Bình, and Ninh Bình. * Bordering provinces such as Hòa Bình, Phú Thọ, and Bắc Giang only exhibit a single variant, indicating no occurrence of the /n/ > [l] phenomenon.
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 92 4.2.3. Overall Distribution of /l/ and /n/ Variants Map 4. Distribution of Variation Patterns by Locality Map 4 provides an overall picture of /l/–/n/ confusion models across localities: ● Regions where both /l/ and /n/ are pronounced accurately and distinctly (i.e., no confusion) are concentrated in the urban districts of Hanoi, with scattered occurrences elsewhere in the delta. Peripheral areas not technically part of the delta, such as Hòa Bình, Phú Thọ, and Bắc Giang, show virtually no confusion. ● Regions with unidirectional confusion /n/ > [l] are concentrated in the northwest. ● Regions with unidirectional confusion /l/ > [n] are concentrated in the southeast. ● Regions with bidirectional confusion /n/ > [l] and /l/ > [n] are distributed in peripheral areas around Hanoi and scattered throughout the delta. 4.3. Interpretation of the Distribution of /l/ and /n/ Variants 4.3.1. According to the Wave Model The wave model suggests that linguistic changes spread from a central point outward to surrounding areas, like ripples on water. The farther from the center, the weaker the influence of these ‘waves’. Based on the results shown in Maps 2, 3, and 4, several centers and waves of variation can be identified: Non-variation center (areas without confusion): Map 4 shows that the nonconfused areas (Model A, red triangle) are mainly concentrated in the western and
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 93 northwestern parts of the delta (Hòa Bình, Phú Thọ, Bắc Giang). These are regions outside the Red River Delta, included in the study for comparative purposes. They maintain a clear phonemic distinction between /l/ and /n/, suggesting that the wave of change has not reached or has been blocked from these areas. In fact, the hilly and mountainous terrain in these regions can be seen as natural barriers that prevent the propagation of linguistic waves from the delta, in accordance with Bailey’s directional wave model (Bailey, 1973). Wave of /l/ > [n] transformation: This is the strongest and most widespread wave. In Map 3, the regions showing unidirectional confusion /l/ > [n] (Model C, blue dots) cover almost the entire eastern and southeastern coastal area of the delta (Hải Phòng, Hải Dương, Hưng Yên, Thái Bình, Nam Định, Ninh Bình). This wave seems to have originated in coastal centers and spread inland. The spread is not abrupt but creates transitional zones, where mixed variants appear—consistent with the idea of waves gradually weakening as they move inward. Wave of /n/ > [l] transformation: This is a smaller-scale wave represented by unidirectional confusion /n/ > [l] (Model B, green diamonds), forming a cluster in Vĩnh Phúc, Bắc Ninh, and the northern outskirts of Hanoi. It may have originated from a local center in the north of Hanoi (the Kinh Bắc region) and spread outward. Transition zone (Hanoi and surrounding areas): The area in and around Hanoi presents a complex picture. Maps 2 and 3 show a concentration of standard pronunciation with clear /l/–/n/ distinction. However, Map 4 shows the presence of all variation types, especially bidirectional confusion (Model D, yellow stars). According to the wave model, Hanoi serves as both a center for standardization and a convergence point where transformation waves from different directions collide (i.e., /l/ > [n] from the southeast and /n/ > [l] from the northwest), resulting in a non-homogeneous dialectal area, especially in suburban villages. Based on previous research results not only in Vietnamese, it can be seen that the above change of /l/ and /n/ is not an isolated phenomenon in the Red River Delta and in Vietnamese. Near Vietnam, the /l/ - /n/ confusion is noted to be very common in dialects at the south of the Yangtze River (China), both in the river valley and in the entire vast South China Plain (Tonghe 2023: 33; Linguistics Atlas of Chinese Dialects, Phonetics, 057). Hoang Thi Chau believes that the /l/ - /n/ confusion may be a linguistic wave spreading across a large area from southern China to northern Vietnam (Hoang Thi Chau 2004: 137). We agree with her that this phenomenon may be related in some way to changes in Chinese in the South China region, but more specific explanations are needed regarding the direction of the spread of linguistic waves. This is related to
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 94 direct contact between Vietnamese and Chinese, which must have been a living Chinese language in Giao Chau (Nguyen Tai Can 1979: 38), after Vietnamese separated from Muong language. The fact that no confusion between /l/ and /n/ was found in Muong language shows that the confusion that exists in Vietnamese must have occurred after Vietnamese separated from Muong language and entered the Old Vietnamese period (13th to 16th century). This is the time when history witnessed many waves of Chinese migration into the Red River Delta. The famous commercial centers in the Red River Delta where the Chinese played a central role in commercial activities were the commercial urban area of Van Don (Quang Ninh), the urban area of Pho Hien (Hung Yen), the port urban area of Hai Phong and the urban area of Thang Long - Hanoi... Along with these commercial urban centers are chains of trading routes along the major rivers spreading throughout the provinces in the lower reaches of the Red River such as Hai Duong, Hung Yen, Thai Binh, Nam Dinh... (Chau Thi Hai 2001). This reality completely coincides with the waves that we have explained above. The wave of change /l/ > [n] originated from the coastal centers (Quang Ninh, Hai Phong) and spread deep inland (the entire lower reaches of the Red River) in the Northwest direction. The wave of change /n/ > [l] originated from the north of the urban area of Thang Long - Hanoi and spread in the Southeast direction. The two waves met to create an intersection in the center of Hanoi and the surrounding areas. 4.3.2. According to the Social Network Model The Social Network Model posits that the maintenance or change of language depends on the structure and density of social relationships within a community. Closed networks (high density) tend to reinforce local pronunciation norms and resist external influences, while open networks (low density) facilitate the adoption of new linguistic variants from outside. Closed networks and the maintenance of ‘local norms’: In the region exhibiting /l/ > [n] confusion (Model C) along the southeastern coast, an interesting phenomenon is observed. Instead of resisting the change, communities here appear to have adopted the non-standard /l/ > [n] variant as their own local norm. The consistency across this large region suggests that the rural village-based closed social networks are strong in reinforcing and spreading this variant. The pronunciation of /l/ as [n] has become a community marker, maintained through repeated social interactions. Similarly, the /n/ > [l] cluster (Model B) in the northwest may also represent a community with a strong enough network to develop and preserve a distinct second local norm, different from both the standardized and coastal patterns.
TRỊNH: THE DISTRIBUTION OF /l/ AND /n/ VARIANTS IN THE RED RIVER DELTA, VIETNAM 95 In fact, Vietnamese dialectology has long noted that the Red River Delta hosts more sub-dialects than any other region (Hoàng Thị Châu 2004: 219). Archaeological evidence suggests this region is the cradle of the Vietnamese people, and where such cradles exist, multiple dialects tend to emerge. In these areas, subsistence-based economies in wet rice agricultural societies fostered village-based settlements, where people lived in near isolation, creating closed social networks with high internal density and few external links. Such networks reinforce and preserve long-standing local norms that diverge from national standards (Hoàng Thị Châu 2004: 221–222). Open networks and linguistic diversity: This is evident in central Hanoi and its periphery, where bidirectional confusion (Model D) occurs. This is characteristic of open social networks. As a political, economic, and cultural center, Hanoi attracts residents from many regions with diverse linguistic habits. Social relationships here tend to be looser and less dense compared to rural communities, creating a linguistically diverse environment where no single standard prevails. As a result, multiple variants coexist, and bidirectional confusion is observed especially in suburban villages. This reflects linguistic instability and variability among individuals in loose, expansive social networks. Thus, the combination of the two explanatory models above shows that the phenomenon of confusion of /l/ and /n/ in the Red River Delta is a complex process, reflecting both the history of migration, economic and cultural exchange (Wave Model) and social structure and community identity (Social Network Model). Conclusion This study provides a comprehensive overview of the distribution of /l/ and /n/ variants in the Red River Delta, Northern Vietnam. It identifies the phenomenon not only as a widespread pronunciation pattern, but more importantly, as a distinctive dialectal feature of the region. The key findings are as follows: 1. Solid phonological foundation: The confusion between /l/ and /n/ stems from their articulatory similarities — same place of articulation (alveolar), shared sonority, and voicing — which facilitate assimilation and can lead to eventual phonemic merger over time. 2. Diverse transformation models: The study identifies four main local variation models: i. No confusion (Model A); ii. Unidirectional confusion /n/ > [l] (Model B); iii.
UTSUMI: DIVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 102 (1) Indonesian a. anwar me-mukul anjing Anwar AV-beat dog b. anjing itu di-pukul Anwar dog that UV-beat Anwar (PATIENT as the subject) “Anwar beat a dog” (2) Indonesian a. Ibu Puji mem-bawa buku ini mother Puji AV-bring book this b. buku ini di-bawa Ibu Puji book this UV-bring mother Puji (CONVEYED THEME as the subject) “Puji’s mother brought this book” Indonesian-type languages are not equipped with complete (productive?) nounmarking systems nor pronominal case systems, as shown in examples (1) and (2), but instead have productive applicative affixes, as observed in (3), in which the Indonesian applicative affixes -kan and -i are exemplified (present?). Voice-marking affixes typically exclusively indicate voice, as in examples (1-3), unlike Philippine-type languages which have Tense-Aspect-Mood and Voice complex affixes. (3) a. pelayan meng-ambil-kan tamu segelas air waiter AV-take-APPL guest one.glass water “The waiter took a glass of water to the guest”(Adapted from Kyokasho Indonesia-go, p 121) b. dia me-naik-i gunung Merapi 3sg AV-go.up-APPL mountain Merapi. “He climbed Mount Merapi”(Adapted from Kyokasho Indonesia-go, p 125) Philippine-type languages, which exhibit two or more Undergoer voices, often display a noun-marking system which plays an important role in a construction. Additionally, a regular alternation between noun-markers and voices is found, as shown in the Tagalog example (4) below. The three pronominal cases are nominative, genitive, and dative/locative, and common nouns are marked by corresponding markers; ang for nominative, ng for genitive, and sa for dative/locative. The pivot noun, which is underscored, is marked by a noun marker ang in (4b), (4c), and (4d), or exhibits
UTSUMI: DIVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 103 nominal case if the pivot is a pronoun, as in 4a. The verb in each sentence includes different voice-marking affixes; the infix -umindicates Actor voice, the suffix -in Patient voice, the suffix -an Locative voice, and the prefix iBeneficiary voice. In addition, Tagalog exhibits two aspects and two moods for each voice form, resulting in four different forms: non-realis/perfective, non-realis/imperfective, realis/perfective, and realis/imperfective (Himmelmann 2005b). (4) Tagalog a. Actor voice h<um>iram =ka ng libro sa aklatan para sa anak=ko <AV>borrow= 2s.NOM GEN book DAT library for DAT child=1sg.GEN b. Patient voice hiram-in=mo ang libro sa aklatan para sa anak=ko borrow-PV=2sg.GEN NOM book DAT library for DAT child=1sg.GEN c. Locative voice hiram-an=mo ng libro ang aklatan para sa anak=ko borrow-LV=2sg.GEN GEN book NOM library for DAT child=1sg.GEN d. Beneficiary voice i-hiram=mo ng libro sa aklatan ang anak=ko BV-borrow=2sg.GEN GEN book DAT library NOM child=1sg.GEN Transitional languages that have mixed features of both Philippine-type and Indonesian-type languages exhibit fewer voice alternations and an imperfect nounmarking system, as in the Talaud examples shown in (5) and the Bantik examples shown in (6). Talaud and Bantik, two languages spoken in north Sulawesi province, Indonesia, have one Actor voice and two Undergoer voices, which are indicated by TAM and voice complex affixes. Their noun-marking system, too, is incomplete when compared with Tagalog. Pronouns, proper names, and singular human nouns exhibit three cases: nominative, genitive, and locative. Other common nouns, on the other hand, do not take a noun marker when in the nominative case, so daho udde (these guests) in (5b) lacks any noun marking. In Talaud, many verbs exhibit a three-way voice alternation, but semantically intransitive verb bases show derivational meaning in Undergoer voice constructions. For example, although the same verb base saɾaiŋ (dance) appears in Example (5a),(5b), and (5c), in the Actor voice (5a) it means “to dance”, in the Goal voice (5b) “to be entertained by watching dance”, and in the Conveyance voice (5c) “use something for dancing”. Bantik, unlike typical Philippinetype languages, exhibits applicative constructions, and a three-way voice alternation is only found with applicative verbs, as shown in example (6), as well as with causative
UTSUMI: DIVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 104 verbs. Conveyance voice is indicated by a zero morpheme, as shown in (6c). All the verbs in examples (5) and (6) are in the non-past tense. (5) Talaud a. i-maŋitou ma-saɾaiŋŋ u-lamaʔa su-daho udde NOM-3pl AV-dance GEN-dish LOC-guests that “They will dance a dish dance to these guests” Actor voice, non-past b. daho udde saɾaiŋ-an ni-maŋitou riŋan nu-saɾaiŋŋa lamaʔa guest that dance-GV GEN-3pl with GEN-dance dish “Those guests will be entertained by them with ‘dish dance’” Goal voice, nonpast c. lamaʔa i-saɾaiŋŋ i-maŋitou dish CV-dance GEN-3pl “Dishes will be used by them in a dance” Conveyance voice, non-past (6)Bantik a.i-remi ma-pa-mandaŋ nu-pisou=ne su-pun nu-teɾiŋ NOM-Remi AV-APPL-test GEN-knife=GEN-3sg LOC-tree GEN-bamboo (AV) b. pun nu-teɾiŋ pa-mandam-en ni-remi nu-pisou=ne tree GEN-bamboo APPL-test-GV GEN-Remi GEN-knife=GEN-3sg (GV) c. pisou=ne Ø-pa-mandaŋ ni-remi su-pun nu-teɾiŋ knife=GEN-3sg Ø-APPL-test GEN-Remi LOC-tree GEN-bamboo (CV) “Remi will test his knife on a bamboo trunk.” To summarize, Philippine-type languages exhibit one Actor voice and two or more Undergoer voices, altogether more than three. Typical Philippine-type languages, such as Tagalog, show a four-way voice alternation. Also, these languages have rigid nounmarking systems. Indonesian-type languages, on the other hand, have two grammatical voices at most, lack noun markers, and have fewer pronominal cases. In between these two types of symmetrical voice languages, there are languages with transitional features, such as fewer voice alternations and an incomplete noun-marking system. 3. Distribution of voice-alternation types case marking systems 3.1. The number of grammatical voices in WMP languages WMP languages located in Taiwan, the Philippines, northern Borneo, and northern Sulawesi are categorized as Philippine-type. Map 1 shows the number of voices in Taiwan and northern Philippines, where most of the languages have four-way voice
U TSUMI : D IVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 105 alternations. When we look at the languages further south in Mindanao, northern Borneo, and Sulawesi, the frequency of languages with three-way voice alternations increases, as shown in Map 2. In the languages of western and southern Indonesia, as shown in Map 3, a two-way voice alternation becomes prevalent. Legend: Number of grammatical voices Map 1: The number of grammatical voices in Taiwan and Northern Philippines
U TSUMI : D IVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 106 Map 2: The number of voices in southern Philippines, northern Borneo, and northern Sulawesi Map 3: The number of grammatical voices in West and South Indonesia
UTSUMI: DIVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 107 3.2. The number of pronominal cases and noun markers in WMP languages Philippine-type languages have a pronominal paradigm in which multiple cases are differentiated, mostly three (pivot, non-pivot, locative; or nominative, genitive, dative/locative), at least 2 cases (nominative and genitive/oblique). Noun markers encode corresponding cases on common nouns. In principle, the number of pronominal cases and noun markers coincide, but some languages have fewer noun markers than pronominal cases. Verb-initial word order is prevalent in Philippine-type languages, so two or more nominal arguments are placed next to one another in transitive constructions. Pronominal cases and noun markers indicate grammatical roles, so that a freer (flexible?) word order of arguments is allowed. Most Indonesian-type languages differentiate two pronominal cases, nominative and genitive, but in some languages, case distinction is limited to certain items only. Other languages do not exhibit case marking, but rather display other oppositions, such as free forms vs dependent forms. These languages rarely require noun markers. In order to distinguish grammatical roles, stricter word order, mostly subject-verb-object (or Actor-verb-Patient), is required. Transitional languages exhibit intermediate features; some require verb-initial word order, whereas others require subject-verb-object word order. Three pronominal cases are found in many of them, but the number of noun markers is often less than that of pronominal cases, or they are not strictly required. As shown in Maps 4 and 5, in Philippine-type languages and transitional languages located (spoken?) from Taiwan to northern Borneo and Sulawesi, a three-way pronominal case distinction is most common. Map 6 shows that many Indonesian-type languages in Sulawesi also have a three-way distinction, whereas those in other areas mostly show a two-way distinction.
U TSUMI : D IVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 108 Map 4: Number of pronominal cases in Taiwan and Northern Philippines Map 5: Number of pronominal cases in southern Philippines, northern Borneo, and northern Sulawesi
U TSUMI : D IVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 109 Map 6: Number of pronominal cases in West and South Indonesia In most Philippine-type languages, three noun markers are found, which distinguish nominative, genitive, and dative/locative cases. However, a few exhibit only two noun markers, as shown in Map 7. Transitional languages, too, exhibit two or three noun markers, as presented in Map 8. In the area where Indonesian-type languages are prevalent, languages with no noun markers are most common, whereas transitional areas find languages with two or three noun markers, as shown in Map 9. Map 7: Number of noun markers in Taiwan and northern Philippines
U TSUMI : D IVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 110 Map 8: Number of noun markers in southern Philippines, northern Borneo, and northern Sulawesi Map 9: Number of noun markers in southern and western Indonesia
UTSUMI: DIVERSITY IN GRAMMATICAL VOICE AND NOUN MARKING SYSTEMS 111 3.3. Correlation between the number of grammatical voices and pronominal/ nominal case marking The number of grammatical voices correlates with both the number of pronominal cases and the number of noun markers. Many Formosan languages and northern Philippine languages have four voices or more, transitional languages have three, and (while?) most Indonesian languages have two. The three-way pronominal case distinction is prevalent in Taiwan, the Philippines, and also in Sulawesi, but a two-way distinction is most often found in areas further south and west. Noun marking systems have a similar distributional trend; languages in Taiwan and the Philippines exhibit a three-way noun marker opposition, whereas those in western and southern Indonesia lack noun markers. Two-way distinctions are found in languages in the transitional areas. In short, languages with four grammatical voices tend to have three pronominal cases and three noun markers. Those with three grammatical voices most commonly have two to three pronominal cases and one to two noun markers, as is often the case with transitional languages. Those with two-way voice alternations seldom have pronominal case opposition or noun markers. Maps 10, 11, and 12 show the correlation between the number of grammatical voices and pronominal cases. The number of grammatical voices in Philippine-type languages is four, while the pronominal cases are mostly three. As a result, in most languages in Taiwan and northern Philippines, and northern Borneo, there are fewer pronominal cases than the number of grammatical voices, as presented in Map 10. In contrast, some languages in the southern Philippines and northern Sulawesi exhibit the same number of voices and pronominal cases, although others have more voices than pronominal cases, as shown in Map 11. In western and southern Indonesia, most languages have two grammatical voices and two pronominal cases, resulting in the same number of voices and pronominal cases, as in Map 12. No language has more pronominal cases than grammatical voices.
Proceedings of the sixth International Conference of Asian Geolinguistics 118 Dialect transition along the Perak River Khairul Ashraaf Saari (University Poly-Tech Malaysia) Nor Hashimah Jalaluddin (Linguistics Association of Malaysia) Harishon Radzi (National University of Malaysia) Abstract: Dialectal transition may occur when there is a gradual change in speech patterns as a result of geographic interaction and social mobility within a community. Phonetic variation, particularly observed along riverine corridors, reflects the dynamic interactions present in such areas. This paper aims to identify the distribution of dialects and analyze the zones of dialect transition occurring along the Perak River corridor. The research instrument focuses on dialectological investigation involving 957 informants from 59 villages situated along the Perak River. This study employs a descriptive geolinguistic approach to examine the causes of dialectal distribution in the study area, with isogloss maps generated using Geographic Information Systems (GIS). The analysis reveals the presence of seven distinct phonetic variants [bɑntaj], [bɑtɑ], [bɑntε], [bɑnte], [bɑtɑl], [bɑntɑ], and [bɑntɑl]. Hulu Perak District has been identified as the primary site of dialect convergence and the most significant transition zone, ultimately leading to the emergence of new phonetic variants as a result of dialectal diffusion in the region. Key words: dialect transition, Perak River, zone transition, lateral, Geographical Information System 1. Introduction Dialect refers to a linguistic variation that serves as a medium for conveying information used by a specific speech community. According to Mengrui Zhu (2023), a dialect not only functions as a tool for expressing thoughts and emotions but also plays a crucial role in the continuity and preservation of a community’s cultural identity. This reflects how language practices often mirror a community’s cultural values. Similarly, Asmah Haji Omar (2015) argues that the use of a dialect is inherited through family lineage and passed down from one generation to the next. Yule (1996) categorizes dialectal variation into two main branches: social dialects (speech influenced by social status) and geographical dialects (speech shaped by SAARI, Khairul Ashraaf, Nor Hashimah Jalaluddin, and Harishon Radzi. 2025. Dialect transition along the Perak River. In Nor Hashimah Jalaluddin, Hiroyuki Suzuki and Mitsuaki Endo (eds.) Proceedings of the sixth International Conference of Asian Geolinguistics, 118–129. doi: https://doi.org/10.5281/zenodo.17204650
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 119 regional community). Social dialects relate to linguistic features determined by factors such as social status, education, occupation, age, or gender. Geographical dialects, also referred to as regional dialects, are those spoken by native speakers within a specific area and are closely associated with the speakers’ residential zones. This is supported by Matthias Urban (2020), who observed phonological differences (i.e., pronunciation) among communities residing in highland and lowland areas. The analysis of dialect distribution may involve elements such as phonological variation, affixation, and unique expressions employed by specific communities. According to Trudgill (2011), in sociolinguistic typology, isolated areas tend to exhibit distinct linguistic features when compared to areas with high levels of interaction and communication. This view can be applied to upstream and downstream communities along a river corridor. As highlighted by Sau Heng Leong in J. Kathirithamby-Wells and John Villiers (1990), trade networks in Peninsular Malaysia historically operated along upstream-downstream models, where small-scale farmers in upstream regions transported their goods to collecting centres downstream. Consequently, riverine networks play a critical role in influencing the spread and diversity of dialects in such areas. Hence, this study seeks to identify dialect distribution patterns and analyze dialect transition zones along the Perak River. 2. Perak River The Perak River is the second-longest river in Peninsular Malaysia, traversing several major districts in the state of Perak, including Hulu Perak, Kuala Kangsar, Kinta, Perak Tengah, and Hilir Perak. Historically, the river has been central to migration patterns and the socio-economic development of local communities, serving as a cradle for early human civilization along its banks. According to Zuliskandar et al. (2015), Kinta was once a major destination for immigrants during the 19th century due to its flourishing tin mining industry. This economic expansion led to the growth of towns and cities such as Ipoh, Kampar, Gopeng, Batu Gajah, Pusing, Pasir Putih, Papan, Lahat, Menglembu, Jelapang, Tasek, Bercham, and Gunung Rapat. Civilizational progress along the Perak River is further influenced by its geographical proximity to the borders of Kedah, Southern Thailand, Kelantan, and Pahang. Social interactions arising from contact with neighboring states and external traders, such as those from Southern Thailand, have contributed significantly to the presence of diverse dialectal variants along the river. The overlapping of territorial boundaries, geographical features, and historical development in the region provides a
S AARI , ET AL .: D IALECT T RANSITION A LONG THE P ERAK R IVER 120 basis for understanding the rich and complex distribution of dialects along the Perak River corridor. Map 1: Map of the state of Perak Sumber: https://www.wonderfulmalaysia.com/map-state-perak-malaysia.htm 3. Research Methodology The exploration of dialect transition was conducted across villages situated along the Perak River corridor. A total of 59 villages were selected as research sites using a combination of systematic sampling and random sampling techniques. Systematic sampling refers to the selection of sample sites based on predetermined criteria, including fixed spatial intervals between villages. In contrast, random sampling involves selecting sample sites using a table of random numbers (Choong Wai Cheong, 2004).
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 121 A mixed sampling method was adopted to overcome the limitations of each individual approach. Systematic sampling may fail to consider topographical features such as rivers and hills or uneven village distribution, while random sampling may result in clustering of sample sites within certain areas, thereby overlooking broader geographical coverage. To address these issues, a combination of both methods was used, guided by several selection criteria, including: a) Random selection of study sites. b) Locations must be homogeneous (readily accessible) and clearly identified using GPS. c) Sites must not include newly developed settlements (e.g., FELDA, FELCRA) or villages not inhabited predominantly by Malay speakers (e.g., Orang Asli communities). d) Locations must be situated near the river corridor. Fieldwork interviews were conducted with 957 informants across the study area, consisting of three demographic groups: elderly, adults, and youth. On average, each village contributed 16 informants. The lexical item /bantal/ (meaning “pillow”) was selected as the primary focus for examining dialect distribution. This selection is justified by the presence of lateral consonants /-l/ and /-r/ in coda positions, which have been categorized as sporadic or unstable in the Perak River area (N. Habibah C.H. and Rahim Aman, 2020). A similar phenomenon has been observed along the Batanghari River in Sumatra, where Anderbeck K. R. (2008) noted irregular patterns of lateral consonants /-r/ and /-l/ in dialect variation. A descriptive geolinguistic approach was employed to analyze dialect distribution and the underlying factors influencing it. This approach links the research database to dialect mapping, as suggested by Khairul Ashraaf Saari (2019). The analysis was carried out descriptively, focusing on mapping dialect distribution patterns and identifying influencing factors using Geographical Information Systems (GIS) to generate isogloss maps based on the selected lexical data.
S AARI , ET AL .: D IALECT T RANSITION A LONG THE P ERAK R IVER 122 Map 2: Map of the Study Area along the Perak River 4. Research Analysis Historically, the Perak River has served as a vital artery for communication and trade since the era of the ancient Perak Sultanate. This river functioned as a principal route for the movement of people, goods, and cultural practices, and it was also one of the earliest sites of Malay settlement. Economic activities such as tin mining, agriculture, and fishing along the river contributed to societal diversification and dialectal interaction. This view is supported by Mior Ahmad Noor (2002), who asserted that any ruling power that managed to control a river’s estuary and course could consequently dominate the socio-economic activities of the surrounding region. Additionally, the development of road infrastructure and modern communication systems has facilitated greater social mobility and enhanced intercommunity interaction.
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 123 Fieldword enabled classification of seven phonetic variants of the lexical item bantal (pillow) along the Perak River corridor. These variants are as follows: [bɑntaj], [bɑtɑ], [bɑntε], [bɑnte], [bɑtɑl], [bɑntɑ], and [bɑntɑl]. Table 1 below presents the frequency distribution of each variant across the respective districts along the Perak River. Table 1: Frequency Ratio of Phonetic Variant Distribution along the Perak River NO DISTRICT VARIANT OF /bantal/ L1bɑntaj L2bɑtɑ L3bɑntε L4bɑnte L5bɑtɑl L6bɑntɑ L7bɑntɑl FREQUENCY RATIO L1 L2 L3 L4 L5 L6 L7 1. Hulu Perak / / / / / 71.43% 2. Kuala Kangsar / / / 42.86% 3. Kinta / / / / 57.14% 4. Perak Tengah / / / / 57.14% 5. Hilir Perak / / 28.57% The table demonstrates that the Hulu Perak district exhibits the most diverse dialectal spread, with five variants recorded at a frequency of 71.43%. This is followed by Kinta and Perak Tengah, each with four variants at 57.14%. The Kuala Kangsar district recorded three variants at 42.86%, while Hilir Perak exhibited the least variation with two variants at 28.57%. 4.1. Phonetic Variants Analysis 4.1.1. Phonetic Variant [bɑntaj] This variant is associated with the Kedah dialect, evident in the phonological change from /-al/ to [-ɑj] at the surface level. As Asmah Haji Omar (2015) explains, in the Kedah dialect, the lateral /l/ typically appears only in prevocalic or intervocalic positions, not in coda. Therefore, the Standard Malay form /-al/ is realized as [-ɑj] in this dialect. The variant was notably present in Hulu Perak, including Kampung Simpang Pulai, and to a lesser extent in Perak Tengah and Kinta. Its discovery in Kampung Simpang Pulai offers a revision to Nor Hashimah Jalaluddin’s (2018) view, which posited that Kedah dialect influence in Perak occurred primarily via riverine and coastal routes. In contrast, this study suggests that dialect diffusion also occurred via land-based routes,
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 124 supported by the absence of significant topographical barriers such as highlands or dense forests. Kinta, as a key destination for immigrants during the 19th-century tin mining era, witnessed rapid urban growth and settlement development, facilitating the inflow of non-local dialects. This supports the argument that modernization and intercultural contact contributed to the integration of the [bɑntaj] variant in the region. 4.1.2. Phonetic Variant [bɑtɑ] This variant, predominantly found in Hulu Perak, is attributed to the Patani dialect influence. According to Nur Habibah C. H. and Rahim Aman (2020), the variant emerged from the migration of Southern Thai Malay (Patani) communities into the upper Perak River region. Phonologically, this variant shows lateral deletion and nasal deletion in homorganic nasal-obstruent clusters. Its concentrated presence in Hulu Perak, which borders Southern Thailand, indicates linguistic diffusion due to cross-border interaction. Sociologically, the variant reflects historical migration driven by conflict and oppression in Thailand, resulting in the resettlement of Patani Malays in areas such as Lenggong and Grik (Nor Hashimah Jalaluddin, 2015). 4.1.3. Phonetic Variant [bɑntε] Statistical data show that [bɑntε] is widespread along the Perak River and is present in all districts except Kerian. This variant reflects segment coalescence, where the final syllable /-al/ becomes [ε]. Zaharani Ahmad (2006) identifies this as a hallmark of the Perak dialect, representing a systematic phonological process. The variant is especially dominant in Parit and Kuala Kangsar, with the latter considered its primary center of distribution. Historically, Kuala Kangsar functioned as Perak’s administrative and trade capital, and its strong local identity facilitated intergenerational retention of this dialect. The variant was institutionalized as part of administrative communication, further reinforcing its linguistic stability. 4.1.4. Phonetic Variant [bɑnte] Found primarily in Kuala Kangsar, especially in Kampung Rambong and Kampung Laneh, this variant arises in a dialectal transition zone between Kedah and Perak dialects (Asmah Haji Omar, 2015). The district’s location between northern and southern Peninsular Malaysia supports its transitional role. The variant [bɑnte] appears to derive from [bɑntε] (Perak) and [bɑntaj] (Kedah). The phonological change from [-ε] to [-e] reflects vowel narrowing, where [-e], with
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 125 the features [+front, +mid-close], replaces [-ε], [+front, +mid-open]. This development is consistent with Schmidt’s (1871) wave theory, which describes the gradual spread and interaction of linguistic features across dialect zones. 4.1.5. Phonetic Variant [bɑtɑl] This variant, minimally present in Hulu Perak, represents another transition form, situated between Standard Malay [bɑntɑl] and the Patani dialect [bɑtɑ]. As Ismail Hussein (1978) notes, these two dialects exert significant influence on the region. The presence of [bɑtɑl] may be attributed to ethnic convergence and historical events, particularly the Pangkor Treaty of 1874, which brought about British administrative influence. This resulted in the spread of Standard Malay and its interaction with existing dialects, contributing to the emergence of hybrid forms such as [bɑtɑl] (Lizawati Ramli, 2015). 4.1.6. Phonetic Variant [bɑntɑ] This variant appears sporadically along the Perak River in districts such as Hulu Perak, Kuala Kangsar, Kinta, and Perak Tengah. It represents a blended form of [bɑntɑl] (Standard Malay) and [bɑtɑ] (Patani dialect), consistent with Ismail Hussein’s (1978) observation of Perak’s linguistic duality: influenced by both northern Patani and southern Johor dialects. Its phonological features developed along the river’s upstream-downstream gradient: [bɑtɑ] in upstream Hulu Perak, transitioning into [bɑntɑ] in midstream areas, and finally into [bɑntɑl] in downstream Hilir Perak. 4.1.7. Phonetic Variant [bɑntɑl] This is the Standard Malay form, exhibiting no phonological transformation. It is predominantly found in Hilir Perak, located downstream. The spread of this variant is driven primarily by migration from Selangor and Kuala Lumpur, where Standard Malay is the norm. Mohd Fadzil Abdul Rashid et al. (2012) reported consistent migration into Perak—especially Batang Padang, Perak Tengah, and Hilir Perak—from 1980 to 2000. Smaller-scale migration also affected Manjung, Kerian, and Kuala Kangsar. These movements support the presence of [bɑntɑl] in southern Perak, demonstrating how external migration influences local dialects through linguistic displacement.
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 126 5. Research Findings The distribution of dialects in a given area is closely influenced by social and geographical factors. Social factors encompass migration, urban development, language standardization, and historical events that shape linguistic practices. Geographical factors include physical features such as mountain ranges, rivers, and political boundaries. Together, these factors shape dialectal diffusion along the Perak River. Based on fieldwork, the researcher identified several villages exhibiting high dialectal diversity (refer to Table 2). This variation arises from dialectal convergence zones, also referred to as transition zones. According to Chambers and Trudgill (1998), a transition zone is a community where neighboring dialects coexist and interact despite geographic barriers. These zones are linguistically dynamic, continually influenced by surrounding dialects. The Perak River study area, bordered by Kedah, Southern Thailand, Kelantan, and Pahang, serves as a strategic linguistic corridor for such interactions. Table 2: Diversity of Phonetic Variant Distribution along the Perak River District Village Name [bɑntɑj] [bɑtɑ] [bɑntɛ] [bɑnte] [bɑtɑl] [bɑntɑ] [bɑntɑl] HULU PERAK Kg Pahit Tengah 18% 27% 0 0 18% 18% 18% Kg Luat Tengah 0 29% 14% 0 14% 14% 29% Kg Simpang Pulai 44% 11% 22% 0 0 11% 11% Kg Selarong 17% 17% 0 0 17% 17% 33% KUALA KANGSAR Kg Perlop 1 11% 5% 42% 21% 0 0 21% KINTA Kg. Tualang Tujuh 22% 0 22% 11% 0 11% 33% PERAK TENGAH Kampung Baru Pulau 5% 0 33% 14% 0 10% 38% The data above confirms the presence of five phonetic variants within single villages, with Hulu Perak demonstrating the highest degree of variation, as seen in Kg Pahit Tengah, Kg Luat Tengah, Kg Simpang Pulai, and Kg Selarong. Kuala Kangsar,
SAARI, ET AL.: DIALECT TRANSITION ALONG THE PERAK RIVER 127 Kinta, and Perak Tengah also show variation in individual villages. Map 3 illustrates these findings visually using color-coded labels . This pattern further highlights Hulu Perak as a key dialect transition zone, supported by its proximity to Southern Thailand, Kelantan, and Kedah. Ongoing interaction among these border communities has led to linguistic convergence. In addition to geography, migration plays a critical role. Asmah Haji Omar (2015) noted the presence of Patani migrants in Larut Matang and Hulu Perak, who fled regional conflict and settled near river basins and highlands in the Titiwangsa Range (Nor Hashimah Jalaluddin, 2015). Map 3: Dialect Transition Zones along the Perak River The spread of dialects along the Perak River demonstrates significant diversity, especially in upstream and midstream areas. Dialect diffusion via language