DECAFIA: A Field-Collected Coffee Leaf Disease Detection Dataset from Santander, Colombia (CoffeeLeaf-CO)
Authors/Creators
Description
DECAFIA / CoffeeLeaf-CO is a field-collected object detection dataset of
coffee leaf disease and pest damage, assembled for Colombian Andean
agronomic conditions and released with a per-image provenance manifest.
VERSION 2.0.0 — CORRECTED RELEASE
This version supersedes version 1. It corrects the image count, the class
scheme, the annotation count, the source attribution and the baseline
metrics. See "Changes from version 1" below. Users of version 1 should
migrate to this release.
--------------------------------------------------
CONTENTS
--------------------------------------------------
2,315 images with YOLO-format bounding box annotations
12,794 annotated instances across 3 detection classes
1,113 background images (healthy leaves, empty label files)
provenance_manifest.csv — per-image source declaration
DATASET_CARD.md — full counts, splits and exclusion rules
CORRECTIONS.md — log of annotation corrections applied
--------------------------------------------------
DETECTION CLASSES
--------------------------------------------------
ID 0 — roya | Hemileia vastatrix (coffee rust)
ID 1 — coco | Curculionidae weevil defoliation (Compsus sp. /
Epicaerus sp.), locally "coco" or "vaquita"
ID 2 — minador | Leucoptera coffeella (coffee leaf miner)
Healthy leaves are NOT a detection class. They are included as background
examples with empty label files, following standard practice for reducing
false positives in object detection. A "healthy" verdict is derived from
the absence of detections above the confidence threshold, not from a class
prediction.
The coco class is, to the authors' knowledge, absent from all publicly
available coffee leaf datasets (BRACOL, JMuBEN, RoCoLe). Its taxonomic
identification follows Constantino et al. (2013), Manual del Cafetero
Colombiano, Vol. 2, pp. 261-306, Cenicafé,
https://doi.org/10.38141/cenbook-0026_25
--------------------------------------------------
SPLITS
--------------------------------------------------
train : 1,618 images | 8,950 instances | 779 background
val : 348 images | 1,896 instances | 167 background
test : 349 images | 1,948 instances | 167 background
Splits are stratified by class presence, seed 42. Verified: zero filename
collisions and zero SHA256 collisions across splits.
--------------------------------------------------
IMAGE SOURCES — FULL DECLARATION
--------------------------------------------------
Every image is attributed individually in provenance_manifest.csv
(filename, split, source, sha256, annotation count, class ids).
Source 1 — Original field collection (1,705 images, 74%)
Prefixes: SANAS_NUEVAS, SANAS_SOCORRO, COCO_RECORTE, COCO_M_A,
COCO_Muy_A, COCO_P_A, ROYA_P_A, ROYA_MA, ROYA_Muy_A
Coffea arabica plantation, El Socorro, Santander, Colombia.
Captured with consumer Android smartphones under natural field lighting,
with no controlled capture setup. All annotations original to this work.
Source 2 — Silva et al. 2020, Brazil (610 images, 26%), CC BY 4.0
Prefixes: ROYA_FONDO, ROYA_RECORTES, MINEIRO, MINADO_RECORTES, MINADOR
Silva, L. B. et al., "Rust and leaf miner in coffee crop (Coffea arabica)",
Mendeley Data, 2020. https://doi.org/10.17632/vfxf4trtcg.5
Redistributed and independently re-annotated under CC BY 4.0 terms.
Per class, by instance count:
coco 100% Source 1
roya 60% Source 1 / 40% Source 2
minador 0% Source 1 / 100% Source 2
By image count, roya images are 52.6% Source 1 / 47.4% Source 2.
Minador performance therefore measures cross-country generalization
rather than Colombian in-field performance.
REMOVED IN THIS VERSION — RoCoLe (454 images)
Prefix: SANAS_ROCOLE. Parraga-Alava, J. et al., "RoCoLe: A robusta coffee
leaf images dataset", Data in Brief, 2019.
https://doi.org/10.17632/c5yvn32dzg
These healthy-leaf images were present in version 1 and have been removed
so that RoCoLe remains usable as an independent external validation set.
Their removal also eliminated five duplicate image pairs, three of which
spanned different splits.
--------------------------------------------------
ANNOTATION
--------------------------------------------------
Tool: VGG Image Annotator (VIA) with Segment Anything Model (SAM) assistance
Format: YOLO bounding boxes, one .txt per image, normalized coordinates
Inter-annotator agreement: Cohen's kappa = 0.91
Two independent annotators with adjudication of disagreements
--------------------------------------------------
BASELINE
--------------------------------------------------
YOLOv8m fine-tuned from MS COCO pre-training.
Seed 42, 100 epochs, imgsz 640, batch 16, AdamW, linear LR decay
(lr0 0.01, lrf 0.001), workers 0.
ultralytics 8.4.148, torch 2.11.0+cu128, Python 3.14, RTX 5070 Laptop.
Training wall-clock time: 1 h 21 m.
Evaluated on the held-out TEST split (349 images, 1,948 instances):
Class | P | R | F1 | mAP50 | mAP50-95
roya | 0.8364 | 0.7931 | 0.8142 | 0.8619 | 0.5700
coco | 0.9177 | 0.9087 | 0.9132 | 0.9638 | 0.7921
minador | 0.8687 | 0.8987 | 0.8834 | 0.9449 | 0.8258
all | 0.8740 | 0.8670 | | 0.9235 | 0.7293
164 of 167 background images (98.2%) produce zero detections at
confidence >= 0.50, supporting the three-class plus background design.
Weights: https://huggingface.co/estebanr25/decafia
Code: https://github.com/estebanr25/decafia-research
--------------------------------------------------
CHANGES FROM VERSION 1
--------------------------------------------------
Version 1 described 2,786 images, 15,181 instances and four classes, and
reported a baseline of mAP50 92.4% / mAP50-95 76.6%. Those figures are
superseded for the following reasons, all documented in this release:
1. The declared class "hojas" (whole-leaf outline) carried zero annotated
instances in the data used to train the released model. Because classes
with no instances are excluded from the mAP mean, the reported figure
was an average over three classes, not four.
2. The 15,181 instance count corresponded to a different annotation set
than the one used for the released model, which contained 12,810
instances.
3. Source attribution was incomplete: ROYA_FONDO, ROYA_RECORTES,
MINADO_RECORTES and MINADOR originate from Silva et al. (2020) and were
not attributed as such in version 1. This release corrects the
attribution for all 610 externally sourced images.
4. The dataset contained five duplicate image pairs, three of them spanning
different splits, which compromised the independence of the evaluation.
5. One spurious annotation (a coco box in MINEIRO_111.jpg, test split) was
removed; see CORRECTIONS.md.
--------------------------------------------------
CITATION
--------------------------------------------------
Rosas Ruiz, L. E., Salom Medina, A. F., & Barrero Perez, J. G. (2026).
DECAFIA: A Field-Collected Coffee Leaf Disease Detection Dataset from
Santander, Colombia (CoffeeLeaf-CO) (Version 2.0.0) [Data set]. Zenodo.
--------------------------------------------------
LICENSE
--------------------------------------------------
Creative Commons Attribution 4.0 International (CC BY 4.0).
Images from Source 2 are redistributed under their original CC BY 4.0
license with attribution as stated above.
Files
CoffeeLeaf-CO-v2.zip
Files
(167.5 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:2ba46bd8f7a23708d982ff578760f1fa
|
167.5 MB | Preview Download |
Additional details
Related works
- Is derived from
- Dataset: 10.17632/vfxf4trtcg.5 (DOI)
- Is supplemented by
- Software: https://github.com/estebanr25/decafia-research (URL)
- References
- Dataset: 10.17632/c5yvn32dzg (DOI)
- Book: 10.38141/cenbook-0026_25 (DOI)
Software
- Repository URL
- https://github.com/estebanr25/decafia-research
- Programming language
- Python
- Development Status
- Active