The University of Pittsburgh English Language Institute Corpus (PELIC)
Description
This is the first public release of the dataset from the University of Pittsburgh English Language Institute Corpus (PELIC). PELIC is a publicly available 4.2-million-word learner corpus of written texts. These texts were collected in an English for Academic Purposes (EAP) context over seven years in the University of Pittsburgh’s Intensive English Program and were produced by over 1100 students with a wide range of linguistic backgrounds and proficiency levels. PELIC is longitudinal, offering greater opportunities for tracking development in a natural classroom setting. In addition to the data, the PELIC repository contains corpus statistics and tutorials on how to access and analyze the data.
Notes
Files
ELI-Data-Mining-Group/PELIC-dataset-v1.0.zip
Files
(230.7 kB)
Name | Size | Download all |
---|---|---|
md5:6fe0a2f5c551b13cab17f827d6590cb8
|
230.7 kB | Preview Download |
Additional details
Related works
- Is supplement to
- https://github.com/ELI-Data-Mining-Group/PELIC-dataset/tree/v1.0 (URL)
Funding
- U.S. National Science Foundation
- Toward a Decade of PSLC Research: Investigating Instructional, Social, and Learner Factors in Robust Learning through Data-Driven Analysis and Modeling 0836012