The Corpus of English–Tagalog Code-switching: an integrative corpus of linguistic contact effects

Publication type: 
Article
Author(s): 
Aaron Santa Maria & Renata Enghels
Citation: 

Santa Maria A. & Enghels R. (2026) The Corpus of English–Tagalog Code-switching: an integrative corpus of linguistic contact effects. Scientific Data.

Description: 

The Corpus of English–Tagalog Code-switching (CEnTaCS) is a new corpus designed to document naturalistic English–Tagalog bilingual speech. Locally referred to as Taglish, it is a widely-used yet understudied contact variety in the Philippines. Recorded in 2025 in Metro Manila, the corpus comprises three interrelated datasets: individual story retelling recordings, cognitive control data based on the Arrow Flanker task, and sociolinguistic questionnaire data. CEnTaCS offers fine-grained data on code-switching across clausal, lexical, and morphemic levels, with particular attention to intra-word code-switching, a salient but still insufficiently documented feature of Taglish. Beyond code-switching, the corpus captures a broad range of language contact phenomena, including borrowing, calquing, convergence, and interference. By making these data systematically available, the corpus supports an integrated account of mixed-language practices that brings together structural, cognitive, and sociolinguistic perspectives. CEnTaCS thus enables researchers to investigate how contact-induced structures are constrained linguistically, how they relate to bilingual processing, and how they index social meaning and identity in a community where bilingualism is deeply embedded in everyday life.

Year of publication : 
2026
Magazine published in: 
Scientific Data