There is a newer version of this record available.

Dataset Open Access

The University of Pittsburgh English Language Institute Corpus (PELIC)

Alan Juffs; Na-Rae Han; Ben Naismith

This is the first public release of the dataset from the University of Pittsburgh English Language Institute Corpus (PELIC). PELIC is a publicly available 4.2-million-word learner corpus of written texts. These texts were collected in an English for Academic Purposes (EAP) context over seven years in the University of Pittsburgh’s Intensive English Program and were produced by over 1100 students with a wide range of linguistic backgrounds and proficiency levels. PELIC is longitudinal, offering greater opportunities for tracking development in a natural classroom setting. In addition to the data, the PELIC repository contains corpus statistics and tutorials on how to access and analyze the data.

Corpus homepage:
Files (230.7 kB)
Name Size
230.7 kB Download
All versions This version
Views 159134
Downloads 1411
Data volume 6.4 MB2.5 MB
Unique views 132115
Unique downloads 1310


Cite as