Global monthly tuna, tuna-like and shark catch (levels 1-2) and fishing effort (level 0) datasets (1950-2024) at 1° and 5° spatial resolution
Authors/Creators
Contributors
Data collector (5):
- 1. Indian Ocean Tuna Commission
- 2. International Commission for the Conservation of Atlantic Tunas
- 3. Western and Central Pacific Fisheries Commission
- 4. Inter-American Tropical Tuna Commission
- 5. Commission for the Conservation of Southern Bluefin Tuna
Description
Global Tuna Atlas (GTA) - Dataset Description
This Zenodo record contains the complete collection of datasets required to reproduce the Global Tuna Atlas (GTA) processing workflow, together with the Docker image used to generate the processed products. The archive contains both the original public datasets collected from the five tuna Regional Fisheries Management Organizations (t-RFMOs) - CCSBT, IATTC, ICCAT, IOTC and WCPFC - and the processed catch and effort datasets generated by the GTA workflow, covering the period 1950-2024 on 1° or 5° spatial grids with monthly temporal resolution.
Lower levels of processing have been officially endorsed by FIRMS and are also published on Zenodo (see FIRMS Global Tuna Atlas datasets, DOI: 10.5281/zenodo.5745958). FIRMS datasets currently deal only with catches and Level 0 data - a global dataset kept as close as possible to what is published on t-RFMO websites — including a lower spatio-temporal resolution version giving the best estimates of total (nominal) catches per year and per ocean.
Data structure
All Global Tuna Atlas datasets comply with a common data format aligned with the CWP Reference Harmonization standard (FAO CWP RH), described in a JSON schema available on GitHub. Catch data are stratified by: month, species, gear_type, fishing_fleet, fishing_mode, geographic_identifier, measurement_unit, measurement (catch), measurement_type (landings or retained catches), and measurement_processing_level (original samples or processed data). A descriptive label column is added for each coded field.
Included datasets
Raw input data (all_raw_data_GTA.zip)
Contains all datasets required to reproduce the workflow from the original public sources: catch and effort datasets from CCSBT, IATTC, ICCAT, IOTC and WCPFC; nominal catch datasets; species, gear, fleet and measurement codelists; conversion factors; spatial reference datasets; mapping tables; workflow configuration files; and all auxiliary resources used by the processing workflow. These files are the exact inputs used to generate every processed GTA product.
Nominal dataset
Annual catches aggregated by ocean and species. This is the official reference dataset used throughout the workflow and the target against which georeferenced catches are raised.
Level 0 catch dataset (IRD Level 0)
The harmonized global georeferenced catch dataset. Original t-RFMO reporting is preserved as closely as possible while field names, codelists and metadata are converted to the common CWP-based data model.
Level 1 catch dataset (IRD Level 1)
Extends Level 0 by harmonizing catch measurement units. Catches reported in numbers are converted to weight (tons) using official IOTC conversion factors, or historical conversion factors developed by Alain Fonteneau (IRD) for the other t-RFMOs. Only the five major species with reliable conversion factors are retained: yellowfin tuna, skipjack tuna, bigeye tuna, albacore, southern bluefin tuna, and swordfish.
Level 2 catch dataset (IRD Level 2)
IRD Level 2 denotes the series of processing steps applied by IRD to convert and raise the georeferenced catch data (Level 1) to match nominal dataset values. It is not a final, static product but a processing level focused on this raising/conversion step. Although some steps mirror those used in the FIRMS Level 0 product (DOI: 10.5281/zenodo.5745958), the entire workflow was rerun to integrate early adjustments to IATTC shark and billfish data prior to final aggregation.
This dataset compiles monthly global catch data for tuna, tuna-like species and sharks from 1950 through 2023, stratified according to the latest CWP standards update: month, species, gear_type (reporting fishing_gear), fishing_fleet (reporting country), fishing_mode (type of school used), geographic_identifier (1° or 5° grid cell), measurement_unit (weight or number), measurement (catch), measurement_type (landings or retained catches), and measurement_processing_level (original samples or processed data), with a label column added to each coded field for descriptive metadata.
Exception - coarser strata for month and geographic_identifier because catches are raised to match nominal data at the strata level judged most probable when an exact match isn't available, a significant share of records end up aggregated at a resolution coarser than a single month or a 1°/5° grid cell — in some cases geographic_identifier covers a much larger area, up to a large portion of an entire ocean. This is a direct consequence of the raising procedure, not a data-quality artifact, but it means the nominal spatio-temporal resolution advertised for Level 2 does not hold uniformly across all records. Notable treatments:
- Catches and numbers are raised to nominal only for exactly matching strata, or otherwise to the strata judged most probable to correspond, since every georeferenced catch stratum should match a nominal catch.
- IATTC purse seine data are distributed as three separate files (tuna, billfish, shark) because tuna is reported by both observers and logbooks, while shark and billfish are observer-only; shark and billfish catches are therefore raised to the fishing effort reported in the tuna dataset (new in v4; previously done at FIRMS Level 0).
- Catches reported in numbers of fish are converted to weight based on nominal data (since v5).
- Strata where tonnage catches were raised to match nominal data have had their corresponding number-of-fish values removed (since v5).
- In some strata nominal data exceeds georeferenced data; this likely reflects differences in aggregation methods and warrants further review with data providers.
Intended use: this dataset improves understanding of fish counts at Level 0 and the extent of georeferenced coverage. It is not suitable for precisely georeferencing catches by country or fleet, and should not be used for studies of fishing-zone legality or quota management — it offers a georeferenced footprint approximating reported biomass, but substantial locational uncertainty remains.
Notable difference from the previous version: missing IATTC data at the 5° resolution have been added.
Level 0 effort dataset (IRD Level 0)
Preserves all publicly available georeferenced fishing effort observations from the five t-RFMOs, processed with the same workflow as the catch data but with different parametrization. Effort is reported in 23 measurement units; only a small number of mappings between similar t-RFMO units have been established (see fdi-mappings). Remaining units are kept unconverted to preserve semantic richness, since each reflects different operational aspects depending on gear, fleet behavior and reporting RFMO. No higher processing level is currently distributed for effort; further aggregation is left to end-users based on their scientific goals.
Known duplication: for ICCAT, and for WCPFC purse-seine data, the same effort may be reported multiple times under different units, since no official conversion between them exists (e.g., Hours.FAD and Hours.FSC may partially inform Hours.STD). For WCPFC, SETS records are associated with a fishing_mode while DAYS records are not, so duplicates do not always share the same fishing_mode. Limited harmonization (e.g., NET-days vs. Nets) has not yet been implemented but may be considered in future releases.
Reproducibility
- Docker image (
gta-workflow.tar.gz): complete software environment to reproduce the workflow from raw inputs without installing R or additional dependencies. - GitHub repository: firms-gta/geoflow-tunaatlas (DOI: 10.5281/zenodo.14039665), with full code, materials and documentation on the impact of each processing step.
- Shiny app:
ghcr.io/firms-gta/tunaatlas_pie_map_shiny_cwp_database:latest, a Docker-based visualization tool for exploring catch data in CWP format. - R package: bastienird/CWP.dataset, for manipulating CWP-standard data and generating structured plots and reports.
Contact
For a customized version of the Global Tuna Atlas with specific filters or adjustments for particular research questions, please contact the maintainers directly.
Notes
Files
Report__level2.pdf
Files
(4.2 GB)
| Name | Size | |
|---|---|---|
|
md5:901f645ffc2e2c7db7197ddf9102530e
|
194.7 MB | Preview Download |
|
md5:d4798ba82d37f7c41b09aa49ced3f179
|
391.1 MB | Preview Download |
|
md5:96f38f4139177733e1f7ee4309781bc6
|
308.8 MB | Preview Download |
|
md5:916464aa10733e20dea791257d1ba15a
|
625.0 MB | Preview Download |
|
md5:dbf99a1589f221b3cf9092d9df8df9a9
|
331.1 MB | Preview Download |
|
md5:7144138f1d2ac4387adcc8334b9e4528
|
2.3 GB | Download |
|
md5:acc950c0601795e750be17234bbbeda9
|
10.4 MB | Preview Download |
|
md5:8dee05e5811089f622abc0eee1d27677
|
8.0 MB | Preview Download |
|
md5:86d0ef05ed06f86674b34b8391e76ee0
|
9.3 MB | Preview Download |
Additional details
Related works
- Is described by
- Computational notebook: 10.5281/zenodo.14025108 (DOI)
- Requires
- Workflow: 10.5281/zenodo.21109008 (DOI)
Funding
Software
- Repository URL
- https://github.com/firms-gta/geoflow-tunaatlas
- Programming language
- R , Dockerfile
- Development Status
- Active