i42

Volume 22. Issue 42. Spring 2027

Thematic Section: East Asian languages in learner corpora of Spanish and other Romance languages: New needs, new tools

Guest Editors of the thematic section: Nobuo Ignacio López-Sako and Cristóbal Lozano (University of Granada)

Important Dates:

  • Deadline for article submission: February 15, 2027
  • Notification of review or acceptance: March 31, 2027
  • Publication date: April 2027

*Preliminary proposals with a 400-word abstract will be accepted until November 30, 2026, by writing to guest editor Nobuo Ignacio López-Sako <nilsako@ugr.es>.

 

Learner corpora are defined as electronic collections of (quasi-)natural textual data from foreign language (FL) or second language (L2) learners, assembled according to explicit design criteria (Granger 2017, 2024). These corpora began to develop in the 1980s based on design models applied to native language (L1) corpora, thereby giving rise to the discipline of Learner Corpus Research (LCR).

Although experimental methods have traditionally been preferred in Second Language Acquisition (SLA) research (Granger, 2017, 2024; Lozano, 2021a, 2021b, 2022; Lozano & Mendikoetxea, 2013; Mendikoetxea, 2014; Myles, 2015), nowadays L2 learner corpora have become a benchmark for the study of FL/L2 learning and acquisition (Granger, 2017, 2024; Lozano, 2021a, 2021b, 2022; Lozano, López-Sako & Montaño, 2026; Parodi, 2010; Rojo, 2021; Sampedro Mella, 2026). This is due to the possibility of accessing large databases of real language collected with methodological rigor and following strict design criteria (Lozano, 2022; Lozano & Mendikoetxea, 2013; Sampedro Mella, 2026; Sinclair, 2005; Tracy-Ventura & Paquot, 2021; Tracy-Ventura et al., 2021). Especially since the 2000s, in the so-called second-generation LCR (Granger, 2024), learner corpus-based research has surged in number and diversity in terms of objects of study (lexicon, morphology, syntax, discourse, pragmatics), L1–L2 language combinations, corpus typology and design (longitudinal, multimodal, developmental, etc.), text typology and modality (monologic vs. dialogic, written vs. spoken, online, emails, etc.), and analytical methods (e.g., Key Word In Context, KWIC) and statistical models used (Chi-square, linear mixed-effects regression models, Bayesian models, etc.) (see Granger, 2024, for a comprehensive review).

 

Although learner corpora have developed significantly, this growth has been dominated by FL/L2 English. Currently, 60% of the 210 corpora listed in the Learner Corpora Around the World (LCAW) database of UC Louvain correspond to English. Of the remaining 40%, Spanish accounts for barely 7.6%, despite the growing interest in learning FL/L2 Spanish (Moreno Fernández & Álvarez Mella, 2025) and its status as the second language of global communication after English (Muñoz-Basols & Fernández Muñoz, 2019). Other neighboring languages have even lower representation (German: 5.7%; French: 5.2%; Italian: 6.7%; Portuguese: 1.9%).

 

On the other hand, as Granger (2024, p. 254) highlights, efforts are being made to diversify and expand the number of participants’ L1s in corpora to allow for more robust and varied interlanguage comparisons (Lozano, 2021a, 2021b; Lozano, López-Sako & Montaño, 2026). For instance, the International Corpus of Learner English (ICLE) has expanded from eleven different L1s in its first version to 26 in version 3 (Granger, 2024). The Corpus Escrito de Español L2 (CEDEL2) has grown considerably: it began with two languages in version 1, but it expanded to 11 in version 2 and 16 in version 3 (Lozano, López-Sako & Montaño, 2026), while the version 2 of the Corpus de Aprendices de Español (CAES) contains 11 native languages (Rojo & Palacios, 2022).

However, a remaining challenge is the geolinguistic distribution of L1s. Specifically, despite the increasing importance of East Asia on the global stage, the representation of Asian languages in corpora of FL/L2 Spanish and other Romance languages remains scarce. In the case of Spanish, to the best of our knowledge, only eleven corpora include East Asian languages (for info of the main ones, see Calero Hernández et al., n.d.; Campillos Llanos, 2014; Cestero Mancera & Penadés Martínez, 2009; Hincapié, 2018; Hincapié & Rubio, 2017; Lozano, 2021a, 2021b, 2022; Lozano, López-Sako & Montaño, 2026; Rojo & Palacios, 2022; Valverde, 2023; Yamada et al., 2020), and their presence in Italian, French, and Portuguese corpora is even lower.

It is essential to emphasize the need and opportunity to boost the development of corpora of Asian L1 languages, considering the growing interest for FL Spanish in the region. In China, there is a significant surge in interest in learning Spanish (Hidalgo Gallardo & Yang, 2020, p. 96; Zhang, 2024, p. 350), and the number of Spanish departments at universities has risen from 15 in 2000 to 106 in 2023 (Ministry of Education, Vocational Training and Sports of Spain (MEFPD), 2025, p. 139). In Japan, the number of universities teaching Spanish increased from 109 in 1985 to 228 in 2021 (Moyano López, Hiroyasu & Yamaura, 2025, p. 20), and in South Korea, it is taught in 48 high schools and 33 universities (Lumbreras Cobo & Rodríguez García, 2020, p. 127).

Against this backdrop, it is crucial to facilitate the work of researchers, teaching practitioners, and teaching materials designers of FL/L2 Spanish and other Romance languages by providing them with linguistic databases and analytical tools that are useful for their research and educational contexts. Therefore, the aim of this thematic section (East Asian languages in learner corpora of Spanish and other Romance languages: New needs, new tools) is to give visibility to corpus-based studies where:

  1. the L1 is one or more East Asian languages, such as Chinese, Japanese, and Korean;
  2. the L2 is Spanish (or another Romance language such as French, Italian, Portuguese, etc.).

Submissions should be methodologically focused, rigorously explaining the design principles of the corpus introduced (e.g., data collection instruments, learner and task variables, proficiency levels, data collection process, data collection platform, data cleaning, etc.). Authors are also encouraged to illustrate which linguistic phenomena could be investigated with these corpora. Contributions presenting empirical findings from corpus-based research will not be accepted, but illustrative/demonstrative studies in the field of linguistic research, education (e.g., Data-Driven Learning, DDL; instructional materials design), or both may be included to showcase their usefulness.

Contributions may be written in any of the languages in which this journal is published.

 

References

Campillos Llanos, L. (2014). A Spanish Oral Learner Corpus for Computer-Aided Error Analysis. Corpora, 9(2), 207–238. https://doi.org/10.3366/cor.2014.0058

Calero Hernández, M. A., et al. [online]. Corpus de Interlengua Española de Aprendices Sinohablantes. https://cineas.udl.cat [July 30, 2026]

Cestero Mancera, A. M., & Penadés Martínez, I. (2009). Corpus de textos escritos para el análisis de errores de aprendices de E/LE (CORANE). Universidad de Alcalá. Retrieved from: https://mele.web.uah.es/wp-content/uploads/2020/12/Cuadernillo-Corpus.pdf

Granger, S. (2017). Learner corpora in foreign language education. In S. Thorne & S. May (Eds.), Language and technology. Encyclopedia of language and education (3rd ed.) (pp. 427–440). Springer. https://doi.org/10.1007/978-3-319-02237-6_33

Granger, S. (2024). From early to future learner corpus research. International Journal of Learner Corpus Research, 10(2), 247–279. https://doi.org/10.1075/ijlcr.00050.gra

Hidalgo Gallardo, M., & Yang, X. (2017). Enseñar ELE en la China continental. De un precario inicio a un prometedor futuro. In M. C. Méndez Santos & M. M. Galindo Merino (Eds.), Atlas de ELE. Geolingüística de la enseñanza del español en el mundo. Vol. II. Asia Oriental (pp. 89–117). EnClaveELE. Retrieved from http://www.todoele.net/atlas-ele

Hincapié, D. (2018). Corpus de Aprendientes de Español como Lengua Extranjera y Segunda Lengua (CAELE/2): el componente escrito. Forma y Función, 31(2), 129–143. https://doi.org/10.15446/fyf.v31n2.74659

Hincapié, D., & Rubio, R. (2017). Diseño y construcción del CAELE/2: Base para una planificación curricular. Hechos y Proyecciones del Lenguaje, 23(1), 42–52.

Lozano, C. (2021a). Corpus textuales de aprendices para investigar sobre la adquisición del español LE/L2. In M. Cruz Piñol (Ed.), E-Research y español LE/L2: Investigar en la era digital (pp. 138–163). Routledge. http://doi.org/10.4324/9780429433528-9

Lozano, C. (2021b). Generative approaches. In N. Tracy-Ventura & M. Paquot (Eds.), The Routledge handbook of second language acquisition and corpora (pp. 213–227). Routledge.

Lozano, C. (2022). CEDEL2: Design, compilation and web interface of an online corpus for L2 Spanish acquisition research. Second Language Research, 38(4), 965–983. https://doi.org/10.1177/02676583211050522

Lozano, C., López-Sako, N. I., & Montaño, J. (2026). Aportaciones de la versión 3 del corpus CEDEL2 de español L2/LE. Variación, 3, 68–100. https://doi.org/10.30827/3020.9854rvcl.3.2026.36933

Lozano, C., & Mendikoetxea, A. (2013). Learner corpora and second language acquisition: The design and collection of CEDEL2. In A. Díaz-Negrillo, N. Ballier, & P. Thompson (Eds.), Automatic treatment and analysis of learner corpus data (pp. 65–100). John Benjamins. https://doi.org/10.1075/scl.59.06loz

Lumbreras Cobo, D., & Rodríguez García, O. (2017). Enseñar ELE en Corea del Sur. In M. C. Méndez Santos & M. M. Galindo Merino (Eds.), Atlas de ELE. Geolingüística de la enseñanza del español en el mundo. Vol. II. Asia Oriental (pp. 119–147). EnClaveELE. Retrieved from http://www.todoele.net/atlas-ele

Mendikoetxea, A. (2014). Corpus-based Research in Second Language Spanish. In K. L. Geeslin (Ed.), The Handbook of Spanish Second Language Acquisition (pp. 11–29). Wiley & Blackwell.

Ministerio de Educación, Formación Profesional y Deportes (2025). El mundo estudia español 2024. Unidad de Acción Educativa Exterior.

Moreno Fernández, F., & Álvarez Mella, H. (2025). Demografía del español en el mundo 2025. In El español en el mundo. Anuario del Instituto Cervantes 2025 (pp. 30–112). Instituto Cervantes. Retrieved from https://cvc.cervantes.es/lengua/anuario/anuario_25/el_espanol_en_el_mundo_anuario_instituto_cervantes_2025.pdf

Moyano López, J. C., Hiroyasu, Y., & Yamaura, A. (2025). La enseñanza del español en las universidades japonesas. Observatorio Global del Español. Instituto Cervantes. Retrieved from https://cervantes.org/es/sobre-nosotros/publicaciones/ensenanza-espanol-universidades-japonesas

Muñoz-Basols, J., & Hernández Muñoz, N. (2019). El español en la era global: agentes y voces de la polifonía panhispánica. Journal of Spanish Language Teaching, 6(2), 79–95. https://doi.org/10.1080/23247797.2020.1752019

Myles, F. (2015). Second Language Acquisition Theory and Learner Corpus Research. In S. Granger, G. Gilquin, & F. Meunier (Eds.), The Cambridge Handbook of Learner Corpus Research (pp. 309–332). Cambridge University Press.

Parodi, G. (2010). Lingüística de corpus: de la teoría a la empiria. Iberoamericana.

Rojo, G. (2021). Introducción a la lingüística de corpus en español. Routledge.

Rojo, G., & Palacios, I. (2022). CAES (Corpus de Aprendices de Español) (v. 2.1). http://galvan.usc.es/caes

Sampedro Mella, M. (2026). Los corpus y la lingüística del corpus del español: desarrollos recientes. Variación, 3, 4–12. https://doi.org/10.30827/3020.9854rvcl.3.2026.37868

Sinclair, J. (2005). How to build a corpus. In M. Wynne (Ed.), Developing linguistic corpora: A guide to good practice (pp. 79–83). Oxbow Books.

Tracy-Ventura, N., & Paquot, M. (2021). Second Language Acquisition and Corpora. An overview. In N. Tracy-Ventura & M. Paquot (Eds.), The Routledge handbook of second language acquisition and corpora (pp. 1–8). Routledge.

Tracy-Ventura, N., Paquot, M., & Myles, F. (2021). The future of corpora in SLA. In N. Tracy-Ventura & M. Paquot (Eds.), The Routledge handbook of second language acquisition and corpora (pp. 409–424). Routledge.

Valverde, P. (2023). El corpus de aprendices japoneses CELEN y su aplicación a la docencia y la investigación en ELE. TEISEL: Tecnologías para la Investigación en Segundas Lenguas, 3, Article e-a_hr01, 1–31. http://doi.org/10.1344/teisel.v3.42898

Yamada, A., Davidson, S., Fernández-Mira, P., Carando, A., Sagae, K., & Sánchez-Gutiérrez, C. (2020). COWS-L2H: A corpus of Spanish learner writing. Research in Corpus Linguistics, 8(1), 17–32. https://doi.org/10.32714/ricl.08.01.02

Zhang, Q. (2024). Análisis del sistema educativo para la enseñanza del español como lengua extranjera en China. Del Español. Revista de Lengua, 2, 349–369. https://doi.org/10.33776/dlesp.v2.8046