Published January 29, 2020 | Version v1

Adversarial Robustness Guarantees for Classification with Gaussian Processes

Description

We investigate adversarial robustness of Gaussian Process classification (GPC) models. Specifically, given a compact subset of the input space $T\subseteq \mathbb{R}^d$ enclosing a test point $x^*$ and a GPC trained on a given dataset $\mathcal{D}$, we aim to compute the minimum and the maximum classification probability for the GPC over all the points in $T$.
In order to do so, we show how functions lower- and upper-bounding the GPC output in $T$ can be derived, and implement those in a branch and bound optimisation algorithm. For any error threshold $\epsilon > 0$ selected \emph{a priori}, we show that our algorithm is guaranteed to reach values $\epsilon$-close to the actual values in finitely many iterations.
We apply our method to experimentally investigate the robustness of GPC models on a 2D synthetic dataset, the SPAM dataset and a subset of the MNIST dataset, providing comparisons of different GPC training methods, and show how our method can be used for interpretability analysis. Our empirical analysis suggests that GPC robustness increases with more accurate posterior estimation.

Notes

This project has received funding from the European Union's Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No 722022

Files

Aistats_2020_almost_cr.pdf

Files (829.1 kB)

Name Size Download all
md5:47d647dd9d6965abb9d469981c7b1d63
829.1 kB Preview Download

Additional details

Funding

European Commission
AffecTech - Personal Technologies for Affective Health 722022