Published August 10, 2018 | Version v1

Can deep neural networks rival human ability to generalize in core object recognition?

Description

Humans are thought to transfer their knowledge well to unseen domains. This putative ability to generalize is often juxtaposed against deep neural networks that are believed to be mostly domain-specific. Here we assessed the extent of generalization abilities in humans and ImageNet-trained models along two axes of image variations: perturbations to images (e.g., shuffling, blurring) and changes in representation style (e.g., paintings, cartoons). We found that models often matched or exceeded human performance across most image perturbations, even without being exposed to such perturbations during training. Nonetheless, humans performed better than models when image styles were varied. We thus asked if there was any linear decoder that, when applied on model features, would rectify model performance. By adding examples from all representation styles to decoder training, we found that models matched or surpassed human performance in all tested categories. Our results indicate that ImageNet-trained model encoding space is sufficiently rich to support suprahuman-level performance across multiple visual domains.

Files

cnn_2018_submission.pdf

Files (241.0 kB)

Name Size Download all
md5:f16feb25526aa1213e1df3e33d309652
241.0 kB Preview Download

Additional details

Funding

European Commission
DEEPCEPTION - Visual perception in deep neural networks 705498