Published January 2025 | Version v2.0

CEREAL I, el Corpus del Español REAL

  • 1. ROR icon German Research Centre for Artificial Intelligence
  • 2. ROR icon University of Bologna

Description

Content:

CEREAL v2 (visit the project website) is a document-level corpus of documents in Spanish extracted from  Colossal OSCAR. Each document in the corpus is classified according to its country of origin. CEREAL covers 24 countries where Spanish is spoken. Following OSCAR, we provide our annotations with CCO license, but we do not hold the copyright of the content text, which comes from OSCAR and therefore from Common Crawl.  

The process to build the corpus and its characteristics can be found in:

Cristina España-Bonet and Alberto Barrón-Cedeño. "Elote, Choclo and Mazorca: on the Varieties of Spanish." In proceedings of the 2024 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL 2024), Mexico City, Mexico, June 2024.

In order to reproduce the results of the paper, please, use v1 of the corpus. The corpus used to train the classifier and the sentence-level version of CEREAL is available at
https://zenodo.org/records/11390829

Files Description:

See the README.txt file

Files

README.txt

Files (121.0 GB)

Name Size
md5:334e5a328a6b4a454df6cd86002c937f
6.4 MB Download
md5:fa3bf56412f26b3d9fbe07f965e2d6b7
17.5 GB Download
md5:d485a60a99dea9b0a12fdcf314f6c1d9
546.2 MB Download
md5:d8015896a48ae4b49cfa63d3b54ef2ec
610.3 MB Download
md5:98aaadd110ff7b2f0b62f03be20f31e0
8.7 GB Download
md5:25c01346b707bc56de37cd65d5e5b015
6.1 GB Download
md5:303f8160cc3a93cc69edeb420cabc4cd
458.8 MB Download
md5:3983b95636e69578fb768b4cebe06ce9
1.4 GB Download
md5:ddb78040d23da21efac8205a2c2671f4
818.7 MB Download
md5:e298f96f9204b963c6d027caedfaeaab
1.1 GB Download
md5:6ce2b370b7620646b21418ab842bcb26
19.7 GB Download
md5:c3f79767391820c2a91fdc5e7dc29120
19.8 GB Download
md5:fecdb881af6d6c56978259dac8c9432c
19.6 GB Download
md5:a207ad0ee5962d34343998ca38b120be
557.2 MB Download
md5:28ee90dd5d62be578c7b37d182ae25be
124.3 MB Download
md5:5392522c7cd9645444ced05d6d597cf3
34.1 MB Download
md5:58809e4ac2eb84f266007e46fd468e22
371.3 MB Download
md5:deac342fc2943989361a3837b46488e6
423.7 MB Download
md5:962d548cdb4c4aee3559b5a4931be362
14.9 GB Download
md5:b749fac2ee896044b70d6ac0d7b60a98
342.5 MB Download
md5:bf15914a5dedaf1307b154b38249b3d1
296.0 MB Download
md5:627a6d35a5273d648feea0a4c2837a69
3.5 GB Download
md5:1973cc0c228097929084060a04196176
3.1 MB Download
md5:759d41670536b00c611dd8ac4f4f9503
99.2 MB Download
md5:5a659d99a8aff1d9c81ad17e4e88b489
506.5 MB Download
md5:6fa4b4bf0dd88ef1f97a6716d677268c
271.7 MB Download
md5:faa6385922c62b7ccf267d6927cafa81
216.5 MB Download
md5:57be53af1d2c7a924695f724e7d26a10
1.6 GB Download
md5:db7b3c98821b4c953a5b21b0ce155c4d
1.4 GB Download
md5:54110ca2b9c9fda8ceba416ffd9ed71a
464.6 kB Preview Download
md5:b6e5be739784b031b0a672b9b1d2fad9
6.5 kB Preview Download

Additional details

Additional titles

Subtitle (English)
Document-level corpus

Related works

Is described by
Other: https://cereal-es.github.io/CEREAL/ (URL)
Is supplemented by
Dataset: 10.5281/zenodo.11390829 (DOI)
Model: https://huggingface.co/cristinae/cereal (URL)

Software

Repository URL
https://github.com/cristinae/docTransformer
Programming language
Python