Published July 7, 2024 | Version v1

Clustering Tasks and Decision Trees with Elegiac Poets

  • 1. Universidad Nacional de la Plata

Contributors

  • 1. Universidad Nacional de la Plata

Description

The dataset contains files generated during a Natural Language Processing (NLP) and automatic text analysis task. Attached is a Jupyter notebook with the complete code, along with several Excel files (.xlsx) containing organized information. Additionally, there are three folders that include files generated during the Silhouette calculation, K-means clustering, and feature extraction using decision trees.

The three folders are:
1. Silhouette Calculation: Contains PNG images of Silhouette plots for various analysis configurations.
2. K-means Clustering: Contains pickle (.pkl) files with features and labels for each combination of excluded author, n-gram type, n-gram range, and matrix type.
3. Feature Extraction: Contains CSV files with lists of documents by cluster and the most important features along with information gain and information gain ratio metrics.

Other file formats included in the dataset are:
- CSV files containing Silhouette scores, optimal clustering results, cluster assignments, and optimal cluster assignments.
- PNG images of scatter plots colored by author and by cluster.
- Pickle files containing the top features extracted during the analysis.

Files

CHR2024_Aarhus_Cluster_Decision_Trees.zip

Files (50.2 MB)

Name Size Download all
md5:d852dfeb2e0928d4c11ada635e266013
50.2 MB Preview Download

Additional details

Dates

Created
2024-07-07

Software

Programming language
Python