Published September 14, 2026 | Version 2.0.0

DECAFIA: A Field-Collected Coffee Leaf Disease Detection Dataset from Santander, Colombia (CoffeeLeaf-CO)

Description

DECAFIA / CoffeeLeaf-CO is a field-collected object detection dataset of
coffee leaf disease and pest damage, assembled for Colombian Andean
agronomic conditions and released with a per-image provenance manifest.

VERSION 2.0.0 — CORRECTED RELEASE
This version supersedes version 1. It corrects the image count, the class
scheme, the annotation count, the source attribution and the baseline
metrics. See "Changes from version 1" below. Users of version 1 should
migrate to this release.

--------------------------------------------------
CONTENTS
--------------------------------------------------
2,315 images with YOLO-format bounding box annotations
12,794 annotated instances across 3 detection classes
1,113 background images (healthy leaves, empty label files)
provenance_manifest.csv — per-image source declaration
DATASET_CARD.md — full counts, splits and exclusion rules
CORRECTIONS.md — log of annotation corrections applied

--------------------------------------------------
DETECTION CLASSES
--------------------------------------------------
ID 0 — roya     | Hemileia vastatrix (coffee rust)
ID 1 — coco     | Curculionidae weevil defoliation (Compsus sp. /
                  Epicaerus sp.), locally "coco" or "vaquita"
ID 2 — minador  | Leucoptera coffeella (coffee leaf miner)

Healthy leaves are NOT a detection class. They are included as background
examples with empty label files, following standard practice for reducing
false positives in object detection. A "healthy" verdict is derived from
the absence of detections above the confidence threshold, not from a class
prediction.

The coco class is, to the authors' knowledge, absent from all publicly
available coffee leaf datasets (BRACOL, JMuBEN, RoCoLe). Its taxonomic
identification follows Constantino et al. (2013), Manual del Cafetero
Colombiano, Vol. 2, pp. 261-306, Cenicafé,
https://doi.org/10.38141/cenbook-0026_25

--------------------------------------------------
SPLITS
--------------------------------------------------
train : 1,618 images |  8,950 instances |   779 background
val   :   348 images |  1,896 instances |   167 background
test  :   349 images |  1,948 instances |   167 background

Splits are stratified by class presence, seed 42. Verified: zero filename
collisions and zero SHA256 collisions across splits.

--------------------------------------------------
IMAGE SOURCES — FULL DECLARATION
--------------------------------------------------
Every image is attributed individually in provenance_manifest.csv
(filename, split, source, sha256, annotation count, class ids).

Source 1 — Original field collection (1,705 images, 74%)
Prefixes: SANAS_NUEVAS, SANAS_SOCORRO, COCO_RECORTE, COCO_M_A,
COCO_Muy_A, COCO_P_A, ROYA_P_A, ROYA_MA, ROYA_Muy_A
Coffea arabica plantation, El Socorro, Santander, Colombia.
Captured with consumer Android smartphones under natural field lighting,
with no controlled capture setup. All annotations original to this work.

Source 2 — Silva et al. 2020, Brazil (610 images, 26%), CC BY 4.0
Prefixes: ROYA_FONDO, ROYA_RECORTES, MINEIRO, MINADO_RECORTES, MINADOR
Silva, L. B. et al., "Rust and leaf miner in coffee crop (Coffea arabica)",
Mendeley Data, 2020. https://doi.org/10.17632/vfxf4trtcg.5
Redistributed and independently re-annotated under CC BY 4.0 terms.

Per class, by instance count:
  coco     100% Source 1
  roya      60% Source 1 / 40% Source 2
  minador    0% Source 1 / 100% Source 2
By image count, roya images are 52.6% Source 1 / 47.4% Source 2.
Minador performance therefore measures cross-country generalization
rather than Colombian in-field performance.

REMOVED IN THIS VERSION — RoCoLe (454 images)
Prefix: SANAS_ROCOLE. Parraga-Alava, J. et al., "RoCoLe: A robusta coffee
leaf images dataset", Data in Brief, 2019.
https://doi.org/10.17632/c5yvn32dzg
These healthy-leaf images were present in version 1 and have been removed
so that RoCoLe remains usable as an independent external validation set.
Their removal also eliminated five duplicate image pairs, three of which
spanned different splits.

--------------------------------------------------
ANNOTATION
--------------------------------------------------
Tool: VGG Image Annotator (VIA) with Segment Anything Model (SAM) assistance
Format: YOLO bounding boxes, one .txt per image, normalized coordinates
Inter-annotator agreement: Cohen's kappa = 0.91
Two independent annotators with adjudication of disagreements

--------------------------------------------------
BASELINE
--------------------------------------------------
YOLOv8m fine-tuned from MS COCO pre-training.
Seed 42, 100 epochs, imgsz 640, batch 16, AdamW, linear LR decay
(lr0 0.01, lrf 0.001), workers 0.
ultralytics 8.4.148, torch 2.11.0+cu128, Python 3.14, RTX 5070 Laptop.
Training wall-clock time: 1 h 21 m.

Evaluated on the held-out TEST split (349 images, 1,948 instances):

  Class    | P      | R      | F1     | mAP50  | mAP50-95
  roya     | 0.8364 | 0.7931 | 0.8142 | 0.8619 | 0.5700
  coco     | 0.9177 | 0.9087 | 0.9132 | 0.9638 | 0.7921
  minador  | 0.8687 | 0.8987 | 0.8834 | 0.9449 | 0.8258
  all      | 0.8740 | 0.8670 |        | 0.9235 | 0.7293

164 of 167 background images (98.2%) produce zero detections at
confidence >= 0.50, supporting the three-class plus background design.

Weights: https://huggingface.co/estebanr25/decafia
Code:    https://github.com/estebanr25/decafia-research

--------------------------------------------------
CHANGES FROM VERSION 1
--------------------------------------------------
Version 1 described 2,786 images, 15,181 instances and four classes, and
reported a baseline of mAP50 92.4% / mAP50-95 76.6%. Those figures are
superseded for the following reasons, all documented in this release:

1. The declared class "hojas" (whole-leaf outline) carried zero annotated
   instances in the data used to train the released model. Because classes
   with no instances are excluded from the mAP mean, the reported figure
   was an average over three classes, not four.
2. The 15,181 instance count corresponded to a different annotation set
   than the one used for the released model, which contained 12,810
   instances.
3. Source attribution was incomplete: ROYA_FONDO, ROYA_RECORTES,
   MINADO_RECORTES and MINADOR originate from Silva et al. (2020) and were
   not attributed as such in version 1. This release corrects the
   attribution for all 610 externally sourced images.
4. The dataset contained five duplicate image pairs, three of them spanning
   different splits, which compromised the independence of the evaluation.
5. One spurious annotation (a coco box in MINEIRO_111.jpg, test split) was
   removed; see CORRECTIONS.md.

--------------------------------------------------
CITATION
--------------------------------------------------
Rosas Ruiz, L. E., Salom Medina, A. F., & Barrero Perez, J. G. (2026).
DECAFIA: A Field-Collected Coffee Leaf Disease Detection Dataset from
Santander, Colombia (CoffeeLeaf-CO) (Version 2.0.0) [Data set]. Zenodo.

--------------------------------------------------
LICENSE
--------------------------------------------------
Creative Commons Attribution 4.0 International (CC BY 4.0).
Images from Source 2 are redistributed under their original CC BY 4.0
license with attribution as stated above.

Files

CoffeeLeaf-CO-v2.zip

Files (167.5 MB)

Name Size Download all
md5:2ba46bd8f7a23708d982ff578760f1fa
167.5 MB Preview Download

Additional details

Related works

Is derived from
Dataset: 10.17632/vfxf4trtcg.5 (DOI)
Is supplemented by
Software: https://github.com/estebanr25/decafia-research (URL)
References
Dataset: 10.17632/c5yvn32dzg (DOI)
Book: 10.38141/cenbook-0026_25 (DOI)

Software

Repository URL
https://github.com/estebanr25/decafia-research
Programming language
Python
Development Status
Active