Published April 6, 2026 | Version 3.1.1
Other Open

Web Scraping with R

Authors/Creators

  • 1. The Language Technology and Data Analysis Laboratory (LADAL), The University of Queensland, Australia

Description

This tutorial introduces web scraping in R using the rvest and xml2 packages, covering HTML structure, CSS selectors, navigating multi-page websites, handling pagination, and storing scraped text and data for downstream analysis. It is aimed at researchers in corpus linguistics and digital humanities who want to collect text data from websites programmatically. This tutorial is part of the Language Technology and Data Analysis Laboratory (LADAL), a free, open-access research infrastructure at the University of Queensland. LADAL provides tutorials, tools, and courses for researchers working with language data. All materials are freely available at https://ladal.edu.au and are part of the Language Data Commons of Australia (LDaCA), funded by ARDC and NCRIS.

Files

Files (252.0 kB)

Name Size Download all
md5:337a9128ef8c9f6243ace55dd68437f9
252.0 kB Download

Additional details