Linguistic Forum Published by MARS Publishers Published by Licensee MARS Publishers. Copyright: © the author(s). This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/). CrossLinguistic Overlap Vocabulary: A Data Approach to Correspondence: Maria Isabel Maldonado Garci <
[email protected] Received: June 25, 2025 Accepted: Abstract This study presents a computational methodology for identifying shared vocabulary between English and Spanish to support second language acquisit calculated to determine the percentage languages. Both Spanish and English use the Roman script, comparison. English has received multiple loanwords from Latin, from which Spanish also derives. This study analys es the most frequent vocabulary in both languages to assess the similarity level and extract similar lexical items. Th corresponding to 1594 shared lexical terms method to compare the 3000 highest frequency terms of the basic vocabulary of English and Spanish, identifying shared vocabulary Keywords: Crosslanguage similarities, acquisition 1. Introduction The second language acquisition literature is language pairs. Spanish and English considered language relatives. The lexical ISSN (Online) 2707 Volume 7, Issue 1, 2025 http://doi.org/ 10.5281/zenodo Published by Licensee MARS Publishers. Copyright: © the author(s). This article is an open access article conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/). Linguistic Overlap EnglishSpanish Vocabulary: A Data -Driven Approach to Lexical Similarity Research Article Maria Isabel Maldonado Garci a spanishprofess[email protected] m> Incharge/Professor, Institute of Languages and Linguistics, University of the Pakistan Publication Details Accepted: November 25, 2025 Published: November This study presents a computational methodology for identifying shared vocabulary between English and Spanish to support second language acquisit ion. Similarity indexes (S.I.) calculated to determine the percentage of orthographic and semantic overlap between the two Spanish and English use the Roman script, which allowed for string received multiple loanwords from Latin, from which Spanish also derives. es the most frequent vocabulary in both languages to assess the similarity level and extract similar lexical items. Th e results show that a lexical similarity level shared lexical terms was calculated using a lexicostatistical method to compare the 3000 highest frequency terms of the basic vocabulary of English and Spanish, identifying shared vocabulary as a pedagogical tool. language similarities, shared vocabulary , cognates, loanwords, The second language acquisition literature is extensive , especially in comparative studies of English , both members of the IndoEuropean language language relatives. The lexical - orthographic comparison of the Spanish and English ISSN (Online) 2707 -5273 Volume 7, Issue 1, 2025 10.5281/zenodo .17772134 Pages 73-92 Published by Licensee MARS Publishers. Copyright: © the author(s). This article is an open access article conditions of the Creative Commons Attribution (CC BY) license Spanish Similarity Incharge/Professor, Institute of Languages and University of the Punjab, Punjab, November 30, 2025 This study presents a computational methodology for identifying shared vocabulary between ion. Similarity indexes (S.I.) were of orthographic and semantic overlap between the two which allowed for string -based received multiple loanwords from Latin, from which Spanish also derives. es the most frequent vocabulary in both languages to assess the similarity level lexical similarity level of 53.13% statistical computational method to compare the 3000 highest frequency terms of the basic vocabulary of English and , cognates, loanwords, second language , especially in comparative studies of European language family, are orthographic comparison of the Spanish and English
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 74 Linguistic Forum 7(1), 2025 languages is relevant since these languages have large numbers of speakers, who are not only native but also speak them as a second language. For example, in the USA, Spanish is the second most spoken language (Ethnologue, 2023). Due to the rise of English because of globalization, technology, and trade worldwide, many Spanish speakers need to teach it. The Indo-European family is the most prominent family of languages by speaker numbers. According to Ethnologue, 439 languages are under its umbrella, and the speakers compose 46% of the world's population. They are located, in the majority, in Europe and South Asia (Ethnologue, 2023). According to Instituto Cervantes (2022), the most spoken languages worldwide are Chinese, English, Spanish, Hindi-Urdu, and Arabic. Instituto Cervantes (2022) says about it: "[…] the projections indicate that the number of potential users of Spanish will continue to increase in absolute terms until 2068, the year in which it will exceed 726 million people, with different degrees of language proficiency. […] The driving force behind the growth of the Spanish-speaking community will be the increase in speakers with native proficiency in this language, which currently represents 6.3% of the world's population (496,573,842 people)." The Spanish language belongs to the Italic branch of the Indo-European family. Forty-seven million three hundred seventy thousand five hundred forty-two speakers speak Spanish in Spain, and 559 million worldwide speakers in more than forty countries (Ethnologue, 2023). In its classification, Ethnologue describes it as "Indo-European, Italic, Romance, Italo-Western, Western, Gallo-Iberian, Ibero-Romance, West Iberian, Castilian." It describes it as "SVO; prepositions; genitives, relatives after noun heads; articles, numerals before noun heads; adjectives before or after noun heads depending on whether it is evaluative or descriptive; question word initial; (C(C))V(C); nontonal. Silbo Gomero whistled a variety of Spanish used in the Canary Islands. "Braille script. Latin script, primary usage." On the other hand, in its classification, English is listed as having 582 million speakers in the United Kingdom, and 334,800,758 speakers in countries such as the USA, Australia, and Canada, where English is a native language (Lewis, 2015). According to Ethnologue (2019), it is the most spoken language in more than a hundred countries, spoken either as a native or SL. It is also the official language of many nations and has become an international language. Ethnologue describes it like this: "Indo-European, Germanic, West, English," Its main features are: "SVO, prepositions, genitives after noun heads, articles, adjectives, numerals before noun heads, question word initial, word order distinguishes subject, object, indirect objects, given and new information, topic and comment, active and passive, causative, comparative, consonant and vowel clusters; nontonal." (Ethnologue, 2019). English uses the Latin alphabet as a script. Ethnologue states: "Braille script. Deseret Alphabet was developed in 1854 with limited usage until 1877. Latin script, primary usage. Shavian (Shaw) script, no longer in use" (Ethnologue, 2019). In order to teach foreign languages swiftly, basic vocabularies were collected to aid students in learning languages. They include the most frequently used words in a particular language. Various basic vocabularies have been published. In the English language, Ogden and Richards' Basic English list was collected (1920-1968). For French, Français Fondamental was published in 1959.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 75 Linguistic Forum 7(1), 2025 For Spanish, Vocabulario del Español Hablado (1975) was published by Luis Marquez, although, previously, Víctor García Hoz wrote his Vocabulario Usual, Vocabulario Común, Vocabulario Fundamental in Madrid in 1953. (2004, pp. 55-56). This study is conducted to identify the everyday lexical items in the high-frequency vocabularies of Spanish and English by comparing them to accelerate second language learning. Similarity indexes (S.I.) among languages are calculated to determine the percentage of similarities in the semantic structure of a language pair. Spanish and English use the Roman script, and English, a Germanic language, has received numerous loanwords from Latin, the language from which Spanish is derived. The present study focuses on the identification and quantification of shared vocabulary between English and Spanish to compare the 3000 most frequent words both languages to identify orthographic and semantic cognates. The resulting list is collected to support students of English and Spanish as a second language acquire core vocabulary efficiently. 1.2 Latin loanwords in English It is widely known that English borrowed much of its vocabulary from Latin. According to Wollman (1993), Old English borrowed between 600 and 700 loanwords, 800 more were borrowed during various periods of the Brittonic languages (Cornish, Welsh, and Breton), and five hundred (early) Latin loanwords are commonly identified in West Germanic languages. Wollman (1993) considers the time of European Christianization, specifically the 6th century, a significant date for borrowing these loanwords. Kavtaria (2011) states that the sixth century AD marked the peak of loanword borrowing due to Christianization. During this time, words like mass, monk, bishop, apostle, angel, altar, and other lexical items related to the workings of the church structure and services were borrowed. Kavtaria also mentions that other words entered the English language through French. However, loanwords from Latin were borrowed in the first century BC during the Roman and sub-Roman periods. Another period of intense borrowing was after the Battle of Hastings (1066), which marked the Norman Conquest in the eleventh century, when Latin words entered the English language through French. Words like state, government, justice, crime, prison, parliament, council, power, court, library, lesson, autumn, dinner, uncle, table, plate, river, and others made their way into the language. England continued to borrow words during the Renaissance (16th and 17th centuries). According to Kavtaria (2011), 65%-70% of English words are borrowed from various languages. In most cases, the reason for borrowing loanwords was the deficit theory, which means the absence of certain lexical items to denote specific realities. For example, potatoes and tomatoes were introduced to England by the Spanish, who had obtained them in the Americas. The English language adopted these words and adapted them from the Spanish (patata and tomate), which had borrowed them from Nahuatl (a language of South America) and other native-American languages. 1.3 Objectives of the Study The study explores the orthographic overlap between two Indo-European languages: English and Spanish, to support second language acquisition. The methodology identifies shared high-
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 76 Linguistic Forum 7(1), 2025 frequency terms from the 3,000 most common words in each language. For this reason, the objectives are: 1. To identify high-frequency shared vocabulary between Spanish and English. 2. To develop a shared vocabulary list based on orthographic and semantic similarity that supports SLA for English and Spanish language learners. 1.4 Research Questions Accordingly, the research questions were formulated: 1. What is the percentage of similarity in English and Spanish? 2. Which high-frequency vocabulary is common in English and Spanish? 2. Literature Review 2.1 Lexical Similarity and Cognates Cognate identification is essential in historical linguistics to establish genealogical relationships between languages (Campbell, 2004). Ethnologue defines lexical similarities: "The percentage of lexical similarity between two linguistic varieties is determined by comparing a set of standardized wordlists and counting those forms that show similarity in form and meaning. Percentages higher than 85% usually indicate a speech variant likely a dialect of the language with which it is being compared" (Lewis, 2013). The preliminary studies on S.I. calculation were conducted by Swadesh in 1948 using a lexicostatistical method in various studies (Swadesh, 1948, 1950, 1952, 1955). Those studies were conducted to find family relations between languages. However, in this paper, the similarity analysis was conducted by recognizing cross-language cognates to identify lexical items that are common or similar in both languages. The varieties used for the comparison were the Mexican variety of Spanish and the American English, which was utilized since they have the most significant number of speakers—similarity levels. Between English and Spanish, the answer has yet to be revealed. Nevertheless, other figures are already extracted; for example, the S.I. between Portuguese and Spanish is 89%, while between English and Portuguese, it is 20.4% (MaldonadoGarcia & Borges, 2014). These indices were calculated by comparing word lists. 2.2 Methods for Measuring Lexical Similarity To determine similarity levels between terms, we look at their level of overlap (Mackay &Kondrak, 2005, pp. 40-47). Brew, C., and McKelvie, D. (1996) had already utilized a corpus of word pairs for extraction for lexicography. Costa et al. (2005) and Sherkina-Lieber (2004:108-121& 2008: 192200) also researched the facilitating effects of cognate words in bilinguals later. Allen D., & Conklin, K. (2013), who studied phonetic and semantic overlap (synonymy) between English and Japanese, Djistra et al. (2010) who investigated bilingual word recognition, and
Cross-Linguistic Overlap EnglishSpanish www.linguisticforum.com language similarity between Dutch and English and Costa et al. who did not study lexical similarity, but the resistance in cognates between Spanish and English and the facilitating effects of cognate in bil ingual speech production. The method utilized in this research was set by Bla and Levenshtein (1965, pp. 707710). Furthermore, other manners of cognate identification were employed by Holmes and Ramos (1993), List (2012), and Rama, Kolachina and other scholars. Word similarity calculation methods are usually based on orthography (Wagner & Fischer, 1974), which is possible only when a language pair uses a particular script, or phonetic (Hall & Dawling, 1980), usually based on the IPA whe These lexical similarities are mainly calculated based on genetic relations (Maldonado Garcia, 2013a, 2013b). However, identification can also occur through strings and synonymy similarities (Maldonado-Garcia & Yapici, 20 14). String analysis of the alignment of the word symbols can be used to calculate the similarity. When orthography is utilized, the symbols' basic order and alignment are identical in the compared pair of languages. In this manner, word similarity is com our analysis, these symbols are orthographic characters. The figure shows some Indo languages and their similarity indexes. Figure 1: Some IndoEuropean Languages and their Similarity Ind Source: Lewis (2013), English 2.3 Latin Loanwords and their role in English The similarity index between English and German is higher than that of other languages because English and German belong to the same family of languages, the Germanic family. Portuguese and French (languages that derive from Latin) have a lower S.I. The reason is that they do not belong to the Germanic family but to the Romance family. In this regard, the S them is comparatively lower, even with the numerous loanwords from Latin that English has received. The Russian language belongs to the East Slavic family, so the S.I. is also lower. Higher S.I.s indicate that the languages are genetic relative similarity percentage between English and Spanish. 2.4 Exis ting gaps and study motivation The motivation behind this methodology is the scarcity of studies investigating similarities in terms of percentages (mai nly conducted by Ethnologue). Another is that basic vocabulary lists are an effective aid for second language learning that teachers can use in their language classes, as shown in Ramirez, Chen and Pasquarella (2013), Hancin (2011). Spanish Vocabulary. . . 77 Linguistic Forum language similarity between Dutch and English and Costa et al. who did not study lexical similarity, but the resistance in cognates between Spanish and English and the facilitating effects of ingual speech production. The method utilized in this research was set by Bla 710). Furthermore, other manners of cognate identification were Ramos (1993), List (2012), and Rama, Kolachina an d and other scholars. Word similarity calculation methods are usually based on orthography (Wagner & Fischer, 1974), which is possible only when a language pair uses a particular script, or phonetic (Hall & Dawling, 1980), usually based on the IPA whe n languages use different scripts. These lexical similarities are mainly calculated based on genetic relations (Maldonado Garcia, 2013a, 2013b). However, identification can also occur through strings and synonymy similarities 14). String analysis of the alignment of the word symbols can be used to calculate the similarity. When orthography is utilized, the symbols' basic order and alignment are identical in the compared pair of languages. In this manner, word similarity is com pared by substituting, inserting, and deleting symbols. For our analysis, these symbols are orthographic characters. The figure shows some Indo languages and their similarity indexes. European Languages and their Similarity Ind exes English -Portuguese calculated by Maldonado Garcia & Borges (2014) 2.3 Latin Loanwords and their role in English The similarity index between English and German is higher than that of other languages because belong to the same family of languages, the Germanic family. Portuguese and French (languages that derive from Latin) have a lower S.I. The reason is that they do not belong to the Germanic family but to the Romance family. In this regard, the S them is comparatively lower, even with the numerous loanwords from Latin that English has received. The Russian language belongs to the East Slavic family, so the S.I. is also lower. Higher S.I.s indicate that the languages are genetic relative s. This study aims to determine S.I., or the similarity percentage between English and Spanish. ting gaps and study motivation The motivation behind this methodology is the scarcity of studies investigating similarities in terms nly conducted by Ethnologue). Another is that basic vocabulary lists are an effective aid for second language learning that teachers can use in their language classes, as shown Pasquarella (2013), Hancin -Bhatt and Nagy (2004), and Dressler LinFo Linguistic Forum 7(1), 2025 language similarity between Dutch and English and Costa et al. who did not study lexical similarity, but the resistance in cognates between Spanish and English and the facilitating effects of ingual speech production. The method utilized in this research was set by Bla ir (1990) 710). Furthermore, other manners of cognate identification were d Kolachina (2013), and other scholars. Word similarity calculation methods are usually based on orthography (Wagner & Fischer, 1974), which is possible only when a language pair uses a particular script, or n languages use different scripts. These lexical similarities are mainly calculated based on genetic relations (Maldonado Garcia, 2013a, 2013b). However, identification can also occur through strings and synonymy similarities 14). String analysis of the alignment of the word symbols can be used to calculate the similarity. When orthography is utilized, the symbols' basic order and pared by substituting, inserting, and deleting symbols. For our analysis, these symbols are orthographic characters. The figure shows some Indo -European Maldonado Garcia & Borges (2014) The similarity index between English and German is higher than that of other languages because belong to the same family of languages, the Germanic family. However, Portuguese and French (languages that derive from Latin) have a lower S.I. The reason is that they do not belong to the Germanic family but to the Romance family. In this regard, the S .I. between them is comparatively lower, even with the numerous loanwords from Latin that English has received. The Russian language belongs to the East Slavic family, so the S.I. is also lower. Higher s. This study aims to determine S.I., or the The motivation behind this methodology is the scarcity of studies investigating similarities in terms nly conducted by Ethnologue). Another is that basic vocabulary lists are an effective aid for second language learning that teachers can use in their language classes, as shown Nagy (2004), and Dressler et al.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 78 Linguistic Forum 7(1), 2025 The Oxford Dictionary lists cognate as having the exact linguistic derivation as another' (and 'formally related; connected: related to or descended from a common ancestor'), (Oxford Dictionary, 2016). For the List, cognate identification is calculated through the edit distance based on similarity. This edit distance is calculated through the alignment matches or mismatches (List, 2012, p. 125). This study responds to the above-mentioned gap by applying computational similarity measures to achieve the identification of shared vocabulary of both languages to support teachers and learners by highlighting useful, high-utility terminology that is easily recognizable. 3. Methodology 3.1 Research Design The approach employed in the study is a quantitative, computational, corpus-based approach intended to identify shared high-frequency vocabulary between English and Spanish. Various studies by Maldonado Garcia and Borges (2014), Maldonado Garcia and Gavrishyk (2016), Chohan and Maldonado Garcia (2019), and Chohan and Maldonado Garcia (2022) use a string similarity algorithm as used in this study. The fundamental objective is to calculate lexical similarity utilizing the edit distance and then categorize the word pairs according to their level of orthographic similarity. The analysis produced a pedagogical list of useful common lexical items to support the teaching and learning of English and Spanish. 3.2 Data Collection In this case, the corpora utilized have two sets of 3000 high-frequency words, one for each language. The data collection was facilitated through an internet search. The search yielded a variety of results. The Oxford 3000 American English (most frequent words) list, an existing corpus, was selected for the study. It constituted Corpus A. The words in this corpus were later correlated with the equivalent terms in Spanish obtained from Corpus B, for which another existing corpus was selected. This is A Frequency Dictionary of Spanish: Core Vocabulary for Learners (Davies, 2006). 3.3 Steps of the Methodology 3.3.1 Algorithm Selection After the word list compilation, a high-frequency words corpus of 3000 words in English its equivalent list in Spanish were compared using the Levenshtein algorithm, used by Kessler (1995) and Rama, Kolachina, and Kolachina (2013) as well as other researchers for the computational dialect comparison. The algorithm calculates the edit distance, which refers to the minimum number of processes necessary to transform one string into another simply through a character's substitution, insertion, or deletion. The edit distance between the symbols in a string a and b is given by lev a,b (|a|,|b|):
Cross-Linguistic Overlap EnglishSpanish www.linguisticforum.com Figure 2: The Algorithm (Levenshtein) With the abovementioned formula, the following computed example was drawn: Figure 3: Computed illustration of Levenshtein distance between Cognates can now be identified through orthographic Categories were set according to the level of edit distance. This is significant because the strings' length may differ and present the same edit distance. This would manifest in different levels of similarity even when the distance is the same. The higher the similarity, An edit distance of zero means that the pairs are exactly the same. In this regard, the lower the edit distance and the longer the string, the higher the similarity percentage. 3.3.2 Distance Computation Initially, a 3000-word highfrequency word list of English was obtained, followed by the Spanish equivalents. Then, the English language high Excel spreadsheet for symbol comparison with Spanish in column 2. Finally, the edit dista calculation was placed in column 3. 3.3.3 Analytical Model and C ategory At this point, several categories for assessing similarity were mapped with an adapted method that combines Blair (1990) a nd a model mentioned in Mackay drew inspiration for the alignment of the vocabularies and the setting of categories. The following categories were implemented. Spanish Vocabulary. . . 79 Linguistic Forum (Levenshtein) Source: (Levenshtein 1965) mentioned formula, the following computed example was drawn: Computed illustration of Levenshtein distance between the sequences ‘gota’ Cognates can now be identified through orthographic and string alignment. set according to the level of edit distance. This is significant because the strings' length may differ and present the same edit distance. This would manifest in different levels of similarity even when the distance is the same. The higher the similarity, the lower the edit distance. An edit distance of zero means that the pairs are exactly the same. In this regard, the lower the edit distance and the longer the string, the higher the similarity percentage. frequency word list of English was obtained, followed by the Spanish equivalents. Then, the English language high - frequency words were placed in column 1 on an Excel spreadsheet for symbol comparison with Spanish in column 2. Finally, the edit dista calculation was placed in column 3. ategory Assignment At this point, several categories for assessing similarity were mapped with an adapted method that nd a model mentioned in Mackay and Kondrak (2005) drew inspiration for the alignment of the vocabularies and the setting of categories. The following LinFo Linguistic Forum 7(1), 2025 mentioned formula, the following computed example was drawn: the sequences ‘gota’ & ‘goto’ set according to the level of edit distance. This is significant because the strings' length may differ and present the same edit distance. This would manifest in different levels of the lower the edit distance. An edit distance of zero means that the pairs are exactly the same. In this regard, the lower the edit frequency word list of English was obtained, followed by the Spanish frequency words were placed in column 1 on an Excel spreadsheet for symbol comparison with Spanish in column 2. Finally, the edit dista nce At this point, several categories for assessing similarity were mapped with an adapted method that Kondrak (2005) , from which we drew inspiration for the alignment of the vocabularies and the setting of categories. The following
Cross-Linguistic Overlap EnglishSpanish www.linguisticforum.com C.1.: High similarity a. Equal terms. Those strings present a distance of zero mean a 100% similarity index. b. Terms that present one difference, and the remaining characters are positioned in the exact location in the strings. c. Terms with distance two because they differ on two characters. C. 2: Moderate similarity d. Terms that present three or more characters C. 3: No similarity e. Terms not similar orthographically. Figure 4: Process for d etermining the similarity level 3.3.4 Ethical Considerations As this study involved computational analysis of public language corpora, hu not involved. Therefore, an ethical conducted using publicly available linguistic data. No personal, or proprietary data were accessed or used at any stage. The process adhered to that all materials were freely accessible, non research purposes. The corpora used followed the principles of size, representativeness and bal across semantic domains and representativeness relat Spanish Vocabulary. . . 80 Linguistic Forum a. Equal terms. Those strings present a distance of zero mean a 100% similarity index. Terms that present one difference, and the remaining characters are positioned in the exact c. Terms with distance two because they differ on two characters. d. Terms that present three or more characters that are different. e. Terms not similar orthographically. etermining the similarity level As this study involved computational analysis of public language corpora, hu man particip ethical review committee approval was not required. conducted using publicly available linguistic data. No personal, or proprietary data were accessed or used at any stage. The process adhered to international ethical standards for data usage that all materials were freely accessible, non - identifiable and intended for linguistic and educational research purposes. The corpora used followed the principles of size, representativeness and bal across semantic domains and representativeness relat ed to vocabulary usage in every LinFo Linguistic Forum 7(1), 2025 a. Equal terms. Those strings present a distance of zero mean a 100% similarity index. Terms that present one difference, and the remaining characters are positioned in the exact man particip ants were committee approval was not required. This study was conducted using publicly available linguistic data. No personal, or proprietary data were accessed international ethical standards for data usage , ensuring identifiable and intended for linguistic and educational research purposes. The corpora used followed the principles of size, representativeness and bal ance ed to vocabulary usage in every day language.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 81 Linguistic Forum 7(1), 2025 This approach strengthens the fairness, reliability and pedagogical relevance of the outcomes obtained. 3.3.5 Limitations of the Study This study presents various key limitations: 1. Focus on orthographic similarity: As both, English and Spanish share the same (Roman) script, some similar words that differ in spelling may not have been identified as cognates. 2. High frequency vocabulary only: While this selection aids in pedagogical relevance, it excludes semantically essential vocabulary. 3. No contextual or syntactic analysis: The study does not account for part-of-speech differences or grammatical roles which may affect the usage of certain words. 4. No machine or AI data was used. 4. Analysis of the Data 4.1 Corpora Comparison In Maldonado Garcia and Borges (2014), the S.I. calculation was performed by collecting 500 highfrequency words in English. In this case, the corpora utilized have two sets of 3000 high-frequency words, one for each language, and an internet search facilitated data collection. The search yielded a variety of results. The Oxford 3000 American English (most frequent words) list was selected for the study. This corpus is a high-frequency list that constituted corpus A. The words in this corpus were later correlated with the equivalent terms in Spanish obtained from Corpus B (A Frequency Dictionary of Spanish. Core Vocabulary for Learners, Davies, 2006). The list contains the highest frequency 3000 English words. Their equivalents in Spanish are included in the first 5000 words. Corpus A was set in the first column, while Corpus B was in the second column. For example, the closest equivalent of the word century is centuria, but this word was not present in Corpus B as it is not a high-frequency word; rather, the word siglo, with the same meaning, was, and so this term was placed in the second column even though centuria is more similar to century than siglo is. In the third column, the edit distance was calculated and documented like this:
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 88 Linguistic Forum 7(1), 2025 D) Terms with distance 6: Appearance-apariencia, approval-aprobación, assignment-asignación, attemptintentar, attend-asistir, bike-bicicleta, citizen-ciudadano, collection-colección,commitment-compromiso, consistent-coherente, constantly-constantemente, cooking-cocinando, correctly-correctamente, deliberatelydeliberamente, depressing-deprimente, designer-diseñador, directly-directamente, disadvantage-desventaja, effective-eficaz, effectively-efectivamente, emphasize-enfatizar, entertain-entretener, exactly-exactamente, examination-examen, government-gobierno, guide-guía, gym-gimnasia, humorous-humorístico, immediatelyinmediatamente, impressive-impresionante, initially-inicialmente, midnight-medianoche, obviouslyobviamente, occasionally-ocasionalmente, package-paquete, perfectly-perfectamente, phone-teléfono, photographer-fotógrafo, probably-probablemente, procedure-procedimiento, properly-propiamente, rapidlyrápidamente, react-reaccionar, refuse-rechazar, reject-rechazar, related-relacionado, requirement-requisito, retired-retirado, simply-simplemente, specifically-específicamente, stretch-estirar, suggestion-sugerencia, surprised-sorprendido, translate-traducir, unconscious-inconsciente, unexpected-inesperado, unfair-injusto (58 pairs). E) Terms with distance 7: Announcement-anuncio, apparently-aparentemente, approximatelyaproximamente, clearly-claramente, entertainment-entretenimiento, extremely-extremadamente, frequentlyfrecuentemente, incredibly-increíblemente, necessarily-necesariamente, northern-delNorte, pointed-puntiagudo, possibly-posiblemente, recently-recientemente, similarity-semejanza, surely-seguramente, surprisingsorprendente, unemployed-desempleado (17pairs). F) Terms with distance 8 and above: According-deacuerdo, basketball-baloncesto, planning-planificación, significantly-significativamente, disappointing-decepcionante, discovery-descubrimiento (6 pairs). Category 3 A) Terms not orthographically similar: he remaining words in the list fall outside the identified similarity thresholds and therefore do not exhibit orthographic proximity to the target items. Table 3: Similar pairs and their distance DISTANCE # PAIRS 0 144 1 347 2 405 3 308 4 195 5 112 6 58 7 17 8 6 Total 1,594 The lexical similarity analysis revealed that 1,594 pairs present certain levels of similarity and are useful for second language acquisition.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 89 Linguistic Forum 7(1), 2025 Table 4: Total Lexical Similarity Index S.I. ENGLISH-SPANISH LANGUAGES ENGLISH-SPANISH Pairs 1,594 S.I. 53.13% Table 5: Improved Lexical S.I. German French Russian Portuguese Spanish English 60% 27% 24% 20.40% 53.13%* *Not genetic similarity, only lexical similarity 5. Conclusion The study aimed to reveal the SI between Spanish and English and to extract the similar lexical items in both languages. The similarity index calculated can be used as a learning strategy for students of Spanish as an L2 (English L1) or English as an L2 (Spanish as an L1). Hence, a statistically driven algorithm was used to compute lexical similarity among two distantly genetically related languages. The comparison was based on the similarity of the segments. Categories were mapped where the degree of variation augments progressively as the Levenshtein or edit distance augments simultaneously. The higher the number of variations, the higher the edit distance between the string segments. In this case, the genetic relationship of the languages was not calculated as the fundamental goal was to extract common vocabulary in both languages, which in this case comes from loanwords borrowed by English. The results yielded a list of shared terminology in Spanish and English that can be useful for students of both languages. In this context, the similarity between both languages is not necessarily due to cognacy, but rather to a large number of loanwords, mainly from Latin and to a minor degree from Greek, that have been assimilated into the English language at some point in history. Funding: This study was not funded in any shape or form by any party. Conflict of Interest: The author declares that he has no conflict of interest. Bio-note: Maria Isabel Maldonado Garcia holds a PhD in General Linguistics from UNED, Spain. Recognized for her outstanding contributions, she has been honored with the prestigious Fatima Jinnah National Pride Award. Currently serving as the Inchargeand Professor at the Institute of Languages and Linguistics, University of the Punjab, Maria Isabel's expertise and dedication continue to enrich the field of linguistics and education. Her remarkable journey and significant achievements make her a trailblazer in her domain, inspiring generations to come.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 90 Linguistic Forum 7(1), 2025 References Allen, D. B., and Conklin, K. (2013). Cross-linguistic similarity and task demands in JapaneseEnglish bilingual processing. PloS one 8(8), e72631.doi:10.1371/journal.pone.0072631 Blair, F. (1990).Survey on a shoestring: A manual for small-scale language surveys. Publication 98. Dallas: Summer Institute of Linguistics, University of Texas at Arlington. Brew, C., and McKelvie, D. (1996). Word-pair extraction for lexicography. In Proceedings of the 2nd International Conference on New Methods in Language Processing, pp. 45-55. Campbell, L. (2004). Historical Linguistics: An Introduction. MIT Press. Centro Virtual Cervantes (2022) El español en el mundo. Anuario 2022. Instituto Cervantes. Accessed 07-02-2024 https://cvc.cervantes.es/lengua/anuario/anuario_22/informes _ic/p01.htm Chohan, M. N., & Maldonado García, M. I. (2022). Phonemic Comparison of Majhi and Shahpuri Dialects of Punjabi. Jahan-e-Tahqeeq, 5(1), 159-168. Chohan, M. N., &Maldonado García, M. I. (2019). Phonemic comparison of English and Punjabi. International Journal of English Linguistics, 9(4), 347-357. Costa, A., et al.(2005). On the facilitatory effects of cognate words in bilingual speech production." Brain and language, 94(1), 94-103. Davies, M. (2006).A frequency dictionary of Spanish: Core vocabulary for learners. NY: Routledge. Djikstra, T., Koji M., Brummelhuis, B., Sappelli, M. & Baayen, H. (2010). How cross-language similarity and task demands affect cognate recognition. Journal of Memory and language, 62(3): 284-301. Dressler, C., Carlo, M. S., Snow, C. E., August, D., & White, C. E. (2011). Spanish-speaking students' use of cognate knowledge to infer the meaning of English words. Bilingualism: Language and cognition, 14(2), 243-255. Dryer, M. S., & Haspelmath, M. (2011). The World Atlas of Language Structures online. Munich: Max Planck Digital Library. Eberhard, D. M., G. F. Simons, and Charles D. Fennig (eds.). (2023). Ethnologue: Languages of the World. Twenty-sixth edition. Dallas, Texas: SIL International. Online version: http://www.ethnologue.com. Hall, P. A.V., & Dowling. G. R. (1980). Approximate string matching." ACM computing surveys (CSUR), 129(4): 381-402.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 91 Linguistic Forum 7(1), 2025 Hancin-Bhatt, B., & Nagy, W. (1994). Lexical transfer and second language morphological development. Applied Psycholinguistics, 15(3), 289-310. Haensch, G. (2004). Los diccionarios del español en el siglo XXI (Vol. 10).Universidad de Salamanca. Holmes, J., & Ramos, R. G. (1993). False friends and reckless guessers: Observing cognate recognition strategies. Second language reading and vocabulary learning, 86-108. Kavtaria, M. (2011). Greek and Latin Loan Words in English Language (Tendencies of Evolution). International Journal of Arts & Sciences 4(18) 255-264. Kessler, B. (1995). Computational dialectology in Irish Gaelic. In Proceedings of the seventh conference on European chapter of the Association for Computational Linguistics, San Francisco: Morgan Kaufmann Publishers Inc., 60-66. Levenshtein, V. I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, 10(8): 707-710. Lewis, M. P., G. F. Simons, & C. D. Fennig, (Eds.). (2013)Ethnologue: Languages of the World. 17th ed. Dallas, TX: SIL International. Ethnologue. .http://www.ethnologue.com/language /spa ; http://www.ethnologue.com/ language/eng; https://www.ethnologue.com/about/ lang uage-info. List, J. M. (2012, April). LexStat: Automatic detection of cognates in multilingual wordlists. In Proceedings of the EACL 2012 Joint Workshop of LINGVIS & UNCLH (pp. 117-125). Mackay, W., & Kondrak, G. (2005, June). Computing word similarity and identifying cognates with Pair Hidden Markov Models. In Proceedings of the Ninth Conference on Computational Natural Language Learning (CoNLL-2005) 40-47). Maldonado García, M. I., & Gavrishyk, E. (2016). False Friends in Urdu and Russian. Journal of Critical Inquiry, 48. Maldonado García, M. I., & Borges de Souza, A. M. (2014). Lexical similarity level between English and Portuguese. ELIA, 14, 145-163. Maldonado García, M. I., & Yapici, M. (2014). Common vocabulary in Urdu and Turkish language: A case of historical onomasiology. PakistanVision, 15(1), 193. Maldonado García, M. I. (2013a).Comparación del léxico básico del español, el inglés y el urdu.Unpublished Doctoral Dissertation. Madrid: Universidad Nacional de Educación a Distancia. Maldonado García, M. I. (2013b). Estudio etimológico de cuatro pares de cognados en español y urdu. Revista Iberoamericana de Lingüística (RIL), 8, 61-74.
Cross-Linguistic Overlap English-Spanish Vocabulary. . . LinFo www.linguisticforum.com 92 Linguistic Forum 7(1), 2025 Oxford English Dictionary(2024).http://www.oxforddictionaries.com/definition/english/ cognate? q=cognate. Retrieved on 17-1-2024. Rama, Taraka, Kolachina, Prasant & Kolachina Sudheer.(2013).Two Methods for Automatic Identification of Cognates.” In Proceedings of Quantitative Investigations in Theoretical Linguistics (QITL-5). Leuven, Belgium, 76. Ramírez, G., Chen, X. & Pasquarella, A. (2013). Cross-linguistic transfer of morphological awareness in Spanish-speaking English language learners: The facilitating effect of cognate knowledge. Topics in Language Disorders 33(1): 73-92. Real Academia Española (2023).Diccionario de la lengua española[Dictionary of the Spanish Language] (23rd ed.). Madrid, Spain. Serfaty, J. and Serrano, R. (2024), Practice Makes Perfect, but How Much Is Necessary? The Role of Relearning in Second Language Grammar Acquisition. Language Learning, 74. 218-248. https://doi.org/10.1111/lang.12585 Sherkina-Lieber, M. (2004). The cognate facilitation effect in bilingual speech processing; the case of Russian-English bilingualism. Cahiers linguistics d’Ottawa, 32: 108-121. Sherkina-Lieber, M. (2008). The cognate facilitation effect is a frequency effect: Evidence from Russian-English bilingualism. In G. Zybatow, L. Szucsich, U. Junghanns, & R. Meyer (Eds.), Formal Description of Slavic Languages: The Fifth Conference, Leipzig 2003.Frankfurt & Main: Peter Lang.192-200. Swadesh, M. (1948).The time value of linguistic diversity. Paper presented at the Viking Fund Supper Conference for Anthropologists, New York. Swadesh, M. (1950). Salish internal relationships. International Journal of American Linguistics, 16(4), 157-167. Swadesh, M. (1952). Lexico-Statistic Dating of Prehistoric Ethnic Contacts: With Special Reference to North American Indians And Eskimos. Proceedings of the American Philosophical Society, 96(4): 452-463. Swadesh, M. (1955). Towards Greater Accuracy in Lexicostatistic Dating.” International Journal of American Linguistics, 21(2): 121-137. Wagner, R. A., and Michael J. Fischer. (1974). The string-to-string correction problem. Journal of the ACM (JACM)219(1): 168-175. Wollmann, A. (1993) Early Latin loan-words in Old English. Anglo-Saxon England22, 1-26.