There is a newer version of the record available.

Published November 29, 2025 | Version v1

Background data for: Ordinal random forests in language data analysis

Description

This dataset is associated with a methodological study on the use of ordinal random-forest models for language data analysis:

  • Mühlbauer, Michael & Lukas Sönning. 2025. Ordinal random forests in language data analysis.

 

Description

This dataset contains information on the usage preferences of speakers of Maltese English with regard to 63 pairs of lexical expressions. These pairs (e.g. truck-lorry or realization-realisation) are known to differ in usage between BrE and AmE (cf. Algeo 2006). The data were elicited with a questionnaire that asks informants to indicate whether they always use one of the two variants, prefer one over the other, have no preference, or do not use either expression (see Krug and Sell 2013 for methodological details). Usage preferences were therefore measured on a symmetric 5-point ordinal scale. Data were collected between 2008 to 2023, as part of a larger research project on lexical and grammatical variation in varieties of English. The current dataset, which informs a methodological study on modeling strategies for ordinal data, include 1,391 speakers and 84,892 ratings in total. It also provides information about a number of socio-demographic variables, including gender, year of birth, age, highest qualification, and the language(s) used at home while growing up. 

 

Questionnaire administration

The questionnaire used for data elicitation, which is included in this post ("lexical_questionnaire.pdf"), was administered in printed form and filled in with a pen. Data were mostly gathered by German university students, who approached (potential) informants and asked if they were willing to take part in an anonymous survey that would be used for scientific purposes only; no financial compensation was offered. Once informants gave their verbal consent to participate in the study, they would either fill in the questionnaire on their own, or they were assisted by the data collector. Participants were predominantly recruited in the streets of Malta, i.e. in various public places (e.g. parks, busses, university campus, etc.) as part of several field trips (between 2008 and 2023), each in connection with a university-level seminar on the varieties of English spoken and written in Malta. Prior to data acquisition, students received basic instructions on how to avoid (or handle) potentially problematic sitations, e.g. ensuring that respondents understood the task and the relevant meaning of lexical expressions (see "Explanation/comment" column in questionnaire), allowing informants to withhold sensitive biographical information, etc. Due to the fact that questionnaires were also administered to university students, either in a lecture hall or on campus, the age distribution of participants in the dataset is uneven, with a notable peak at around 20 years of age.

 

Data and file overview

  • malta_lexical_data_mixfabOF.tsv  tab-separated data table containing the ratings provided by the 1,391 informants
  • item_labels.tsv  tab-separated data table with information about the full set of 68 item pairs in the questionnaire
  • lexical_questionnaire.pdf  questionnaire used for data elicitation

 

Data-specific information for: item_labels.tsv

UTF-8-encoded, tab-separated data table with 69 rows and 5 columns. The columns represent the following variables:

  • item  item pair ID (links to file "malta_lexical_data_mixfabOF.tsv", which includes the rating data)
  • AmE_variant  the more/exclusively/traditionally American English variant
  • BrE_variant  the more/exclusively/traditionally British English variant
  • AmE_short  short(er) label for American English variant, used for plotting
  • BrE_short  short(er) label for British English variant, used for plotting

 

Data-specific information for: malta_lexical_data_mixfabOF.tsv

UTF-8-encoded, tab-separated data table with 84,893 rows and 11 columns. The columns represent the following variables:

  • id  anonymized participant ID
  • year  year of data collection
  • age  age of participant in years
  • gender  self-reported gender of participant ("Female", "Male")
  • highest_qualification highest level of qualification (completed or ongoing)
    • "Higher Education"
    • "No/Other Qualification"
    • "Professional/Vocational"
    • "Secondary Education"
    • "Work-Based Learning"
  • language_home  language(s) used at home while growing up, where "(Other)" indicates an optional third language
    • "English (Other)"
    • "English, Maltese (Other)"
    • "Maltese, English (Other)"
    • "Maltese (Other)"
  • ratio  proportion of life spent in Malta
  • item  item pair ID
  • rating  numeric version of the rating provided by the participant
    • +2  I always use this (British English) expression
    • +1  I use this (British English) expression more often
    • 0  I have no preference
    • -1  I use this (American English) expression more often
    • -2  I always use this (American English) expression
  • date_of_birth : year of birth  
  • rating_factor  second numeric version of the rating provided by the participant (for straightforward conversion to a factor)
    • 5  I always use this (British English) expression
    • 4  I use this (British English) expression more often
    • 3  I have no preference
    • 2  I use this (American English) expression more often
    • 1  I always use this (American English) expression

 

References

Algeo, John. 2006. British or American English: A handbook of word and grammar patterns. Cambridge: Cambridge University Press.

Krug, Manfred & Katrin Sell. 2013. Designing and conducting interviews and questionnaires. In Manfred Krug & Julia Schlüter (eds.), Research methods in language variation and change, 69–98. Cambridge: Cambridge University Press.

Files

lexical_questionnaire.pdf

Files (7.0 MB)

Name Size Download all
md5:de6f36e9be9c60a3e6a7c0a8af3d8510
3.5 kB Download
md5:c04328be5a55d5674407d82ce59f22a6
160.9 kB Preview Download
md5:d04ad30db47eebe0fdc3620b43c9ac11
6.8 MB Download

Additional details

Funding

Deutsche Forschungsgemeinschaft
548274092