Published March 9, 2026 | Version 1.0

Śmigiel: corpus of human-written and machine-generated text fragments in Polish

  • 1. ROR icon University of Warsaw
  • 2. Institute of Computer Science, Polish Academy of Sciences
  • 3. ROR icon Universitat Pompeu Fabra
  • 4. ROR icon Institute of Computer Science

Description

Introduction

Śmigiel (Spotting Machine-Generated Text from LLMs for Polish) is the first open dataset for training and evaluating machine-generated text (MGT) in Polish. It includes a collection of human-written text (HWT) fragments from six domains, which are used to prompt text generation by eight language models capable of producing credible Polish text. In addition to the raw corpus of over 462K generated texts, we also release a cleaned source- and domain-balanced dataset suitable for training and evaluating MGT detectors.

The dataset is described in an article presented at LREC 2026 conference (Śmigiel Dataset: Laying Foundations for Investigating Machine-Generated Text Detection in Polish) and was used in the shared task at PolEval 2025.

Human-written text (HWT)

To compile a corpus of Polish HWT passages, we use several recent datasets with permissive licenses:

Machine-generated text (MGT)

To compile a corpus of MGT passages, we prompt the following LLMs:

Śmigiel Dataset

The published data covers two stages of Śmigiel:

  • Raw text, including full HWT fragments and MGT generations,
  • Postprocessed text, following cleaning, sampling and trimming to obtain a clean and unbiased dataset for the purpose of training and evaluating MGT detection models.

The descriptions below are based on articles covering Śmigiel dataset and the shared task, so you should look at them for more information.

Raw text

The raw text is included in three files:

  • RAW_train_testalpha.csv:  fragments in the literature, reviews, social and wikipedia domains, including generations by all models except Llama-3.3-70B-Instruct -- used for the training portion and the test-alpha),
  • RAW_testbeta.csv: the same domains as train, but the generations are from Llama-3.3-70B-Instruct,
  • RAW_testgamma.csv: all of the fragments in the news and parlamint domains.

The files are in a CSV format with the following fields:

  • numerical ID,
  • HWT fragment,
  • prefix obtained from the HWT fragment,
  • full prompt provided to an LLM,
  • source dataset,
  • domain,
  • number of tokens generated,
  • sampling strategy,
  • text generated by the model,
  • LLM model used.

Postprocessed text

The postprocessed text includes the following portions:

  • training: 80% of the fragments in the literature, reviews, social and wikipedia domains, including generations by all models except Llama-3.3-70B-Instruct,
  • testing data split in one of two ways:
    • according to the source:
      • test_alpha: 10% of the fragments in the same domain as training, including generations by the same models,
      • test_beta: 10% of the fragments in the same domain as train, but using the Llama-3.3-70B-Instruct generations,
      • test_gamma: all of the fragments in the news and parlamint domains.
    • according to their use in shared task:
      • testA: 50% of test_alpha, 33% of test_beta and 33% of test_gamma,
      • testB: 50% of test_alpha, 67% of test_beta and 67% of test_gamma,

Thanks to this structure, the model trained on the training portion can be tested on data from the same distribution (test_alpha), from an unseen model (test_beta), from an unseen domain (test_gamma) or from a mixtures of these (testA, testB).

Each portion is saved in three files:

  • <name>.txt: newline-separated fragments,
  • <name>.key: newline-sparated labels (0: HWT, 1: MGT)
  • <name>.meta: TSV file with additional information: text checksum, source model (or human) and sampling strategy.

Statistics

The Śmigiel dataset is balanced across the main categories (MGT and HWT) and across text domains and LLM sizes.

It includes 32K HWT and 32K MGT examples. The MGT portion contains 10K samples from small LLMs, 10K from medium LLMs, and 12K from large LLMs. For domains, there are about 5.5K HWT and MGT samples each from four types: literature, reviews, social media, and wikipedia. Śmigiel also includes two unseen test domains: news (2.6K HWT and 2.6K MGT) and parlamint (7K samples per category).

More

For more information, you can:

This work was supported by the Ramón y Cajal grant RYC2024-050327-I, funded by the Spanish State Research Agency (MICIU/AEI/10.13039/501100011033) and by the European Social Fund Plus (ESF+) of the European Union. We also gratefully acknowledge Polish high-performance computing infrastructure PLGrid (HPC Centers: ACK Cyfronet AGH) for providing computer facilities and support within computational grant no. PLG/2025/018019.

Files

alpha.txt

Files (1.3 GB)

Name Size
md5:540b7ee742a49a212f559196128f0a67
9.0 kB Download
md5:d38cb5694502903c98a72fefd1f1a060
248.4 kB Download
md5:630f5268f00fa0e91ef2340a478efc5d
3.4 MB Preview Download
md5:4215005320b37c33c9153a8e9d36d964
8.7 kB Download
md5:0ba416b68da5e9abf6f81dc0d2a93132
241.5 kB Download
md5:46af67ff1d6e00e8a5d6918a6ee28100
3.7 MB Preview Download
md5:1b0441fc3bf76173a04cadf8d625b3cf
41.0 kB Download
md5:265a5244e9108c56287f3baa19e4f10f
1.1 MB Download
md5:3c12b47a8e4023672209c4ccd82a1b03
10.9 MB Preview Download
md5:75efe85272078576279fa82bb3332bf7
12.6 MB Preview Download
md5:1d608cbb56190859cb6c50edf98bcc37
294.1 MB Preview Download
md5:81153c44f29b6b914a61a8a271afd1d1
890.6 MB Preview Download
md5:e9bd479a8e51956905a3ece7f22e974c
20.7 kB Download
md5:92b77891401c5acdcb22603892c54d94
635.6 kB Download
md5:84c09c2940aed283b492f6b86ff8ef2c
6.5 MB Preview Download
md5:3c40f9bd269abdb9ca65663fe2e3d5a0
36.9 kB Download
md5:8819010843161683a75ccd78584d0bfa
1.1 MB Download
md5:37db3e20ded63a2eda03c572a9661566
11.2 MB Preview Download
md5:182a6e2361d0c10bef3dc08629f07090
71.5 kB Download
md5:65babf84858255b5848c52f79fdbfb07
2.2 MB Download
md5:707c2e1c55a0c17ebddcae60bdfd8636
26.8 MB Preview Download

Additional details

Related works

Is described by
Publication: 10.63317/3p7ghe9pfm8v (DOI)

Funding

Agencia Estatal de Investigación
Ramón y Cajal fellowship RYC2024-050327-I

Dates

Available
2026-04-28

Software

Repository URL
https://github.com/JStrebeyko/code-of-smigiel
Programming language
Python