Published June 26, 2025 | Version v1

Training Data Size Sensitivity in Unsupervised Rhyme Recognition: Data & replication code

  • 1. ROR icon Czech Academy of Sciences, Institute of Czech Literature
  • 2. ROR icon University of Tartu
  • 3. ROR icon Charles University
  • 4. ROR icon Tilburg University
  • 5. ROR icon University of Basel
  • 6. ROR icon University of Ljubljana
  • 7. ROR icon The Institute of the Polish Language of the Polish Academy of Sciences

Description

Data & replication code for "Training Data Size Sensitivity in Unsupervised Rhyme Recognition"

  • train.py
    • script to train RhymeTagger models
  • tag.py
    • script to perform annotations with RhymeTagger models
  • replication.ipynb
    • replication code
  •  [corpora]
    • text corpora in json files
  • [gold standard]
    • random samples ([texts]) and their annotations provided by two human annotators ([annotation])
  • [models]
    • RhymeTagger models trained on samples of different sizes
  • [output]
    • annotations provided by RhymeTagger models
  • [llms]
    • annotations provided by three different LLMs (GPT4-o, Claude 3.7 Sonnet, DeepSeek-V3)

Files

rt.zip

Files (1.6 GB)

Name Size
md5:b1eea963ac69c34462f7573a66a2af64
1.6 GB Preview Download