Crowding in humans is unlike that in convolutional neural networks

Ben Lonnqvist; Alasdair D F Clarke; Ramakrishna Chakravarthi

doi:10.1016/j.neunet.2020.03.021

Crowding in humans is unlike that in convolutional neural networks

Neural Netw. 2020 Jun:126:262-274. doi: 10.1016/j.neunet.2020.03.021. Epub 2020 Mar 27.

Authors

Ben Lonnqvist¹, Alasdair D F Clarke², Ramakrishna Chakravarthi³

Affiliations

¹ Business School, University of Aberdeen, United Kingdom of Great Britain and Northern Ireland. Electronic address: ben.lonnqvist.16@abdn.ac.uk.
² Department of Psychology, University of Essex, United Kingdom of Great Britain and Northern Ireland. Electronic address: a.clarke@essex.ac.uk.
³ School of Psychology, University of Aberdeen, United Kingdom of Great Britain and Northern Ireland. Electronic address: rama@abdn.ac.uk.

PMID: 32272430
DOI: 10.1016/j.neunet.2020.03.021

Abstract

Object recognition is a primary function of the human visual system. It has recently been claimed that the highly successful ability to recognise objects in a set of emergent computer vision systems-Deep Convolutional Neural Networks (DCNNs)-can form a useful guide to recognition in humans. To test this assertion, we systematically evaluated visual crowding, a dramatic breakdown of recognition in clutter, in DCNNs and compared their performance to extant research in humans. We examined crowding in three architectures of DCNNs with the same methodology as that used among humans. We manipulated multiple stimulus factors including inter-letter spacing, letter colour, size, and flanker location to assess the extent and shape of crowding in DCNNs. We found that crowding followed a predictable pattern across architectures that was different from that in humans. Some characteristic hallmarks of human crowding, such as invariance to size, the effect of target-flanker similarity, and confusions between target and flanker identities, were completely missing, minimised or even reversed. These data show that DCNNs, while proficient in object recognition, likely achieve this competence through a set of mechanisms that are distinct from those in humans. They are not necessarily equivalent models of human or primate object recognition and caution must be exercised when inferring mechanisms derived from their operation.

Keywords: Convolutional neural networks; Crowding; Object recognition.

MeSH terms

Artificial Intelligence*
Crowding*
Humans
Neural Networks, Computer*
Pattern Recognition, Automated / methods*
Pattern Recognition, Visual / physiology
Photic Stimulation / methods*
Recognition, Psychology / physiology
Visual Perception / physiology