Machine learning for communicating water risk and predicting public quality categories
Authors/Creators
Description
Water-quality indices are not only technical summaries but communication devices: they translate complex multivariate monitoring data into a few public categories used to understand and act upon environmental risk. This article reinterprets the Canadian Council of Ministers of the Environment Water Quality Index (CCME-WQI) as a communication instrument and asks how faithfully its public categories can be reconstructed from raw physicochemical data, and what is lost when a continuous risk gradient is compressed into a public message. Using 400 observations from 12 stations of Peru’s National Water Authority in the Llave River Basin, supervised classifiers and a stacking ensemble were trained to recover the three public categories (Excellent, Good, Fair/Poor), while silhouette analysis and principal-component analysis probed the data structure. The categories proved only moderately recoverable: the best model reached 65% exact accuracy, a ceiling no method surpassed, yet about 91% within-one-category accuracy, almost never confusing excellent with degraded water. Nearly half the observations sit at the index ceiling and most of the rest sit just below a threshold, so distinct public messages attach to near-identical water; a near-zero category silhouette and PCA confirm a high-dimensional continuum that the labels cut across. We argue that this gap between gradient and label is the price of legibility, and discuss how category boundaries frame public risk perception and the unequal capacity of Quechua- and Aymara-speaking communities to access these messages.
Files
Art_XX_Machine learning for communicating water risk....pdf
Files
(914.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:ebbf41849ecc65c93c7e1b24097e7f6c
|
914.0 kB | Preview Download |