Published August 16, 2023 | Version 1.0.0

Dataset of Mastodon toots using the hashtag #Fairdata

Authors/Creators

  • 1. Nele

Contributors

Supervisor:

  • 1. University of Potsdam

Description

This dataset provides mastodon toots that are using the hashtag FAIR-Data. This dataset is supposed to provide the foundation of further network analysis around the topic FAIR-Data. 

Data Collection

Data was harvested using the Mastodon API. For each Mastodon server listed in the dataset, the API was employed to retrieve posts tagged with "fairdata". Up to 5,000 posts were collected per request, utilizing the API's pagination feature to obtain all available posts for that hashtag. Each Mastodon server was queried separately, and the results were stored in distinct CSV files.

Potential Duplicates

Given that Mastodon operates as a federated network, a post made on one instance can be replicated across different instances. This implies that the same post might appear in the data from multiple Mastodon servers, likely accounting for the duplicates observed in the dataset.

Analysis of Federation Dynamics: Duplicates can reveal which content is shared between servers and which servers are most active in the federation. In this context, duplicates could provide valuable insights for the network analysis.

The columns are:

  1. id - A unique identifier for each post.
  2. created_at - Timestamp indicating when the post was created.
  3. content - The content or message of the post.
  4. account - Detailed information about the account that created the post (ID, username, URL).
  5. replies_count - Number of replies to the post.
  6. reblogs_count - Number of reblogs or shares of the post.
  7. favourites_count - Number of favorites or likes the post has received.
  8. language - Language in which the post is written.
  9. mentions - Any mentions in the post (for example, other users).
  10. tags - Tags associated with the post.
  11. emojis - Emojis used in the post.
  12. category - Category of the post.

Notes

This dataset is licensed under the Creative Commons Attribution-NonCommercial (CC-BY-NC) license. This means you are free to share (copy and redistribute) the material in any medium or format and adapt (remix, transform, and build upon) the material. However, you must give appropriate credit, provide a link to the license, and indicate if changes were made. You may not use the material for commercial purposes.

Files

categorized_toots_preprocessed.csv

Files (25.1 MB)

Name Size Download all
md5:1eca15144a535ff1d01470da33d4ba4d
25.1 MB Preview Download