Introduction

Ned Block

https://doi.org/10.1093/oso/9780197622223.003.0001

Pages

1–60

Abstract

This chapter introduces key concepts of perception, cognition, concept, proposition, high-level properties, low-level properties, iconic representations, nonconceptual state, nonpropositional state, associative agnosia, apperceptive agnosia, the global broadcasting approach to consciousness, the recurrent processing view of consciousness, the distinction between perception and a minimal immediate direct perceptual judgment, core cognition, the difference between intrinsic and derived intentionality, peripheral inflation, fragile visual short-term memory, working memory, slot vs. pool models of working memory, conceptual engineering, the language of thought, and the default mode network. It explains the three-layer methodology of the book: It starts with prescientific ways of thinking of perception and cognition, using them to identify apparent indicators of perception and cognition; then considers whether the indicators depend on constitutive properties of perception and cognition or mere symptoms; and then leverages those conclusions to find the constitutive features of perception and cognition. The chapter explains how work on the neural and psychological basis of consciousness can be repurposed to isolate the psychological and neural basis of perception. The chapter ends with a consideration of consequences outside philosophy of mind of the views presented.

Keywords: conceptpropositionformatcontentfragile short-term memoryworking memoryagnosia

Subject

 Philosophy of Mind

Collection: Oxford Scholarship Online

What is the difference between seeing and thinking? Is the border between seeing and thinking a joint in nature in the sense of a fundamental explanatory difference? Is it a difference of degree? Does thinking affect seeing, or, rather, is seeing “cognitively penetrable”? Are we aware of faces, causation, numerosity, and other “high-level” properties or only of the colors, shapes, and textures that—according to the advocate of high-level perception—are the low-level basis on which we see them? How can we distinguish between low-level and high-level perception, and how can we distinguish between high-level perception and perceptual judgment? Is there evaluative perception or is evaluation a matter of emotion and perceptual judgment? Is perception conceptual and propositional? Is perception iconic or more akin to language in being discursive? Is seeing singular? Which is more fundamental, visual attribution or visual discrimination? Is all seeing seeing-as? What is the difference between the format and content of perception, and do perception and cognition have different formats? Is perception probabilistic and, if so, why are we not normally aware of this probabilistic nature of perception? Does perception require perceptual constancies? Are the basic features of mind known as “core cognition” a third category in between perception and cognition? Are there perceptual categories that are not concepts? Where does consciousness fit in with regard to the difference between seeing and thinking? What is the best theory of consciousness and does the perception/cognition border have any relevance to which theories of consciousness are best? These are the questions I will be exploring in this book. I will be exploring them not mainly by appeals to “intuitions,” as is common in philosophy of perception, but by appeal to empirical evidence, including experiments in neuroscience and psychology.

I will orient the discussion around the question of a joint in nature between perception and cognition resting on differences in format and kind of representation that have been the subject of a great deal of controversy in recent years. Perception is constitutively nonconceptual, nonpropositional, and iconic, but cognition has none of these properties constitutively.

Claims that perception is iconic, nonconceptual, or nonpropositional have been advocated—and opposed—for many years. Stephen Kosslyn is perhaps the most notable advocate of recent years for the iconicity of both perception and mental imagery, but many others have also advocated such views (Block, 1981, 1983a, 1983b; Burge, 2010bCarey, 2009Kosslyn, 1980Kosslyn, Pinker, Schwartz, & Smith, 1979Kosslyn, Thompson, & Ganis, 2006). Zenon Pylyshyn has been a notable opponent of iconicity for both perception and mental imagery (Pylyshyn, 19732003). Similarly, many have advocated nonconceptual and/or nonpropositional perception (Burge, 2010a; Carey, 20092011bCrane, 1988Evans, 1982; Peacocke, 19861989). And there are many opponents (McDowell, 1994Strong, 1930Wittgenstein, 1953).

The intended contribution of this book is not that perception is nonconceptual, nonpropositional and iconic, but the elaboration of what that view comes to, engagement with the evidence for and against it; and using this picture of perception to refute widely held theories of consciousness, the global workspace theory and the higher order thought theory, and to argue for a new reason to think there can be phenomenal consciousness without access consciousness. I think I have new evidentially based arguments for other familiar theses. For example, Chapter 6 is devoted to an argument for non-conceptual color perception based on developmental psychology. In the first few chapters, I will be especially concerned with how to distinguish low-level perception from high-level perception and how to distinguish high-level perception from perceptual judgment.

This book is all about evidence. I aim to avoid pronouncements and intuitions. I will also explore the relation between these claims about format, content, and state to modularity and consciousness, and rebut arguments that misconceive the border between perception and cognition. (The content of a representation is the way it represents the world to be, the way the world has to be for the representation to be accurate. The format of a representation is the structure of its representational vehicle.)

I also aim to avoid cherry-picking evidence. When I know of evidence that goes against my claims, I will introduce it.

To say that perception is constitutively X is to say that it is in the nature of perception to be X. The evidence I will present that perception constitutively has certain properties applies most clearly to actual creatures that perceive rather than possible creatures. The evidence I will be talking about concerns the way actual perceptual mechanisms work. Occasionally I will talk about consequences for robot perception, though I am less certain about those claims.

Although I am arguing for certain constitutive properties of perception, my evidence is almost entirely concerned with vision. I believe the points I am making apply at least to all the spatial senses. There is good reason to include smell in the spatial senses (Smith, 2015). Humans can track odors across grass blindfolded, and their tracking deteriorates if they are deprived of the use of one nostril. Humans can also identify the direction of a smell via stereo-olfaction without moving (Jacobs, Arter, Cook, & Sulloway, 2015).

But I will not be talking about the nonspatial senses except in asides such as this one. I will not be talking about proprioception, the sense of balance, thermoception (the sense of temperature, kinesthesia (the sense of movement), chronoception (the sense of time) or others of the perhaps 21 senses (Durie, 2005).

Although there are many types of perception, a wide variety of them obey the same laws of perception such as Weber’s Law (that the discriminability of two stimuli is a linear function of the ratios of the intensities of the two stimuli) and Stevens’s Power Law (that says that perceived intensity is proportional to actual intensity raised to an exponent, where the exponent differs according to stimulus type). Stevens’s Power Law has been shown to apply not only to various forms of visual and auditory intensity but also to many other kinds of perception and sensation. A recent textbook chapter lists the exponents for the following kinds of perception: electric shock, warmth on arm, heaviness for lifted weights, pressure on arm, cold on arm, vibration, loudness of white noise, loudness of 1 KHz tone, and brightness of white light (Zwislocki, 2009). In sum, although my evidence is almost entirely from vision, there is a prima facie case to be made that many of our perceptual modalities have similar underlying natures.

One feature of the treatment of these ideas that will emerge in Chapters 4 and 6 is that perceptual and cognitive states can share the same or at least similar contents, nonconceptual and nonpropositional in the case of perception, conceptual and propositional in the case of cognition. So nonconceptual content is not a kind of content. Chapter 6 will use an extended example in terms of color contents.

Jerry Fodor argued for a joint in nature between perception and cognition based on the distinction between modular (perception) and nonmodular (cognition) processing. The modularity thesis says perception is a fast, inflexible, automatic, domain-specific system that is informationally encapsulated from other systems, has a fixed neural architecture, a characteristic ontogenetic pace of development, and processes that are themselves largely opaque to other systems (Fodor, 1983). This book argues that the joint in nature between perception and cognition does not depend on modularity, and more specifically that there is a joint and there also is considerable penetration of perception by cognition. Still, there is something to the idea that perception is modular, with only restricted kinds of cognitive penetration. In Chapter 9, I will critique a recent proposal in the spirit of modularity by E. J. Green, but I am friendly to the general approach.

How do we know that we are perceiving a face as a face—as opposed to perceiving a face as having certain colors, lines, curves, textures, shapes, and the like—all low-level properties? To answer that question, we need methods of distinguishing high-level from low-level perception.

Something can look blue, look like a face, look expensive, or look like a piano. But are these kinds of looking all perceptual as opposed to judgmental overlays on perception? The perceptual representation of blue is low-level, whereas the perceptual representation of faceness is high-level. Low-level visual representations are products of sensory transduction that are causally involved in the production of other (mid- and high-level) visual representations and include representations of contrast, spatial relations, motion, texture, brightness, and color. (Transduction is conversion of signals received by sense organs into neural impulses.) Another low-level property is spatial frequency (roughly, “stripiness”—see Figure 1.1.). Representations at a slightly higher level, sometimes characterized as mid-level, include representations of shapes that indicate corners, junctions, and contours (Long, Konkle, Cohen, & Alvarez, 2016). High-level representations include representations of recognizable objects and object-parts, but also causation and numerosity. Some think that conscious perception is never high-level, for example Alex Byrne, Adam Pautz, and Jesse Prinz (Pautz, 2021; Prinz, 2002Siegel & Byrne, 2016) and that what happens when it seems that something looks like a face is that we perceive lower-level properties while judging that certain high-level properties apply. Those who advocate high-level perception are often said to advocate rich as opposed to thin perception (Siegel, 2010Siegel & Byrne, 2016).

 Superimposed low-frequency and high-frequency images. From close up you see Albert Einstein (high-frequency image), but from far away (or if you squint) you see Marilyn Monroe (low-frequency image). Any curve can be decomposed into component sine waves. The spatial frequency of the curve depends on the spatial frequencies of those sine waves. See the Wikipedia article on spatial frequency at https://en.wikipedia.org/wiki/Spatial_frequency. Thanks to Aude Oliva for the figure. (See Oliva, Torralba, & Schyns, 2006.)

Figure 1.1

 Superimposed low-frequency and high-frequency images. From close up you see Albert Einstein (high-frequency image), but from far away (or if you squint) you see Marilyn Monroe (low-frequency image). Any curve can be decomposed into component sine waves. The spatial frequency of the curve depends on the spatial frequencies of those sine waves. See the Wikipedia article on spatial frequency at https://en.wikipedia.org/wiki/Spatial_frequency. Thanks to Aude Oliva for the figure. (See Oliva, Torralba, & Schyns, 2006.)

Open in new tabDownload slide

Although I will be arguing at length that we do perceptually represent some high-level properties, I agree with the skeptics that it is a mistake to postulate rich content solely on the ground that a perceiver can visually recognize something. For example, I can visually recognize that something is a pipe wrench without seeing it as such.

One view of the rich/thin debate that is not the one I am endorsing is that the thin view is one in which we perceive low-level properties “directly” and high-level properties “indirectly.” This picture, sometimes called the “layering” conception, is that “we see more abstract and worldly things in and by seeing simpler and more primitive ones” (Lycan, 2014, p. 7). I think we visually attribute faceness, causation, and numerosity directly.

High-level perception is to a large extent causally dependent on low-level perception but not totally dependent on low-level perception. For example, there are direct connections between subcortical structures like the amygdala and the high-level fusiform face area (Herrington, Taylor, Grupe, Curby, & Schultz, 2011). The amygdala is activated by fearful faces by a pathway that skips the low-level perceptual analysis of early visual cortex (McFadyen, Mermillod, Mattingley, Halász, & Garrido, 2017). Further, even to the extent that high-level perception is causally dependent on low-level perception, that doesn’t make high-level perception indirect in the sense that high-level percepts are composed of low-level percepts. High-level perception in the sense I am using is just the perceptual attribution of high-level properties.

What is the evidence that we visually represent some high-level properties? I will be addressing this question in more detail later, especially in Chapter 2, but I will give one line of evidence here having to do with visual agnosia (a term due to Freud). My argument will be superficially similar to one given by Tim Bayne (2009), and one reason for introducing the argument in the introductory chapter is that the difference between my argument and Bayne’s illustrates a key feature of the methodology of this book.

A nineteenth-century classification system of agnosias due to Lissauer (Shallice & Jackson, 1988) that is still useful distinguishes between apperceptive agnosia, in which subjects have problems with grouping of perceptual elements, and associative agnosia, in which grouping is normal or close to normal but recognition is not. Apperceptive agnosia is now more commonly referred to as “visual form agnosia.” (The latter term is used in the second [but not the first] edition of Martha Farah’s classic book [Farah, 2004].) Apperceptive agnosics can often see color and texture but cannot see shapes and often have trouble telling whether they are seeing one object or two objects. They have difficulty copying drawings even when they can draw from memory with their eyes closed.

Associative agnosics can often copy drawings well without knowing what they are copying. Associative agnosia is characterized by failure to apply high-level visual properties, even though associative agnosics often have normal recognition through nonvisual sensory modalities and intact low-level vision. For example, an associative agnosic might be unable to visually recognize that something is a dog despite being able to recognize that something is a dog haptically and despite having the concept of a dog. Hans-Lukas Teuber described associative agnosia as “percepts stripped of their meaning” (quoted in Goodale & Milner, 2005, p. 13).

In addition to broad visual associative agnosia there are also specialized associative agnosias such as prosopagnosia, the inability to recognize faces despite intact low-level perception. Color agnosia, a type of associative agnosia, is discussed in Chapters 4 and 6. A recent paper reported an even more specific agnosia, agnosia for the digits ‘2’ through ‘9.’ The patient could recognize ‘0’ and ‘1’ and letters of the alphabet. The patient was an engineering geologist with a degenerative brain disease. Although he could not recognize the eight mentioned digits, he could do mental arithmetic. He came up with a different system of representing the eight digits and set up his computer to use the new numerals on the screen so that he could keep working (Kean, 2020Schubert et al., 2020).

Apperceptive agnosics don’t demonstrate the existence of low-level without high-level perception, since what they lack is a kind of low-level perception involving low-level grouping. However, patients whose perceptual grouping is normal but have associative agnosia have low-level perception without high-level perception. The existence of associative agnosics is an excellent reason to believe that there is high-level perception and high-level perceptual content.

I think this point establishes that there is high-level perceptual content, but it does not show that there is high-level perceptual phenomenology, that is conscious high-level content (cf. Bayne, 2009). But separating the issue of high-level perceptual phenomenology into the question of high-level perceptual content and whether that content can be conscious allows us to get a grip on the latter question. And this is the point of methodology that I am illustrating here.

It is widely agreed among those who study the neuroscience of consciousness that specific phenomenal contents are based at least in part in the brain circuits that process that kind of content. For example, visual brain area MT+ processes motion perception, both 3D and 2D motion. The evidence is partly correlational—even illusory motion and motion aftereffects involve activation in MT+ (Tootell et al., 1995), but the conclusion is also supported by stimulation to MT+. Indeed the experience of specific directions of motion is produced by stimulation to distinct subareas of MT+ (Salzman, Murasugi, Britten, & Newsome, 1992).

Although we know that activation in MT+ is part of the neural basis of motion experience, what we don’t know is what else has to happen for (conscious) motion experience to occur. Though we don’t know, theorists do make claims about what else has to happen. For example, the global workspace theory (to be discussed in detail in Chapters 4 and 13) says that the motion representational contents must be “globally broadcast.”

What is global broadcasting? Perception sparks competition among neural “coalitions” in perceptual areas in the back of the head. (See Koch, 2004, for a readable account of neural coalitions and (Dehaene, 2014) for a readable presentation of the global workspace account.) These coalitions involve feedback/feed-forward loops known as recurrent activations. The recurrent activations in the back of the head trigger “ignition,” in which a winning neural coalition in perceptual areas links up with frontal circuits via “workspace neurons” that link the front and back of the head. The result is a systemwide mutually supporting neural coalition that advocates of the global workspace model describe as broadcasting in the global workspace. See Figure 1.2. According to the global workspace theory, representations of motion based in MT+ being globally broadcast is constitutive of perceptual consciousness of motion.

 Schematic diagram of the global workspace. Dark pointers added. I am grateful to Stan Dehaene for supplying this drawing.

Figure 1.2

 Schematic diagram of the global workspace. Dark pointers added. I am grateful to Stan Dehaene for supplying this drawing.

Open in new tabDownload slide

The “global workspace” model of consciousness (Dehaene, 2014) is illustrated in Figure 1.2. The outer ring indicates the sensory surfaces of the body. Circles are neural systems and lines are links between them. Filled circles are activated systems and thick lines are activated links. Activated neural coalitions compete with one another to trigger recurrent (reverberatory) activity, symbolized by the ovals circling strongly activated networks. Sufficiently activated networks trigger recurrent activity in cognitive areas in the center of the diagram and they in turn feed back to the sensory activations, maintaining the sensory excitation until displaced by a new dominant coalition. Not everyone accepts the global workspace theory as a theory of consciousness (including me), but it does serve to illustrate one kind (again, not the only kind) of competition among sensory activations that in many circumstances is “winner-takes-all,” with the losers precluded from consciousness. (As we will see in Chapter 7, there is a revised version of the global workspace theory, the global playground theory, that is arguably superior.)

I have argued that recurrent activations confined to the back of the head can be conscious without triggering central activation. Because of local recurrence and other factors, these are “winners” in a local competition without triggering global workspace activation (Block, 2007a). Strong recurrent activations in the back of the head normally trigger “ignition,” in which a winning neural coalition in the back of the head spreads into recurrent activations in frontal areas that in turn feed-back to sensory areas. See Figure 1.2. As Dehaene and colleagues have shown, such locally recurrent activations can be produced reliably with a strong stimulus and strong distraction of attention (Dehaene, Changeux, Nacchache, Sackur, & Sergent, 2006). (Since I am concerned in most chapters of this book with normal perception, I won’t say much about my disagreement with the model until we get to the chapter on consciousness. But for the record, I think the global workspace model is a better model of conceptualization than of consciousness.)

Here is the point: What is true for low-level perceptual representations such as the representation of motion is also true for high-level perceptual representations, such as face-representations based in the fusiform face area. When high-level perceptual representations are broadcast in the global workspace they give rise to high-level conscious phenomenology—according to the global workspace theory. So the global workspace account plus the fact that there is high-level perceptual content leads to the conclusion that there is high-level perceptual phenomenology. (I am assuming that the high-level representations are sometimes broadcast in the global workspace, but that is obvious enough.) So if we assume the global workspace account of consciousness, we can move from the high-level perceptual content shown by the agnosia evidence to high-level perceptual phenomenology.

The first-order recurrent activation account of conscious content is basically a truncated form of the global workspace account: It identifies conscious perception with the recurrent activations in the back of the head without the requirement of broadcasting in the global workspace (Block, 2005bLamme, 2003). First-order theories do not say that recurrent activations are by themselves sufficient for consciousness. These activations are only sufficient given background conditions. Those background conditions probably include intact connectivity with subcortical structures. (The cortex is the thin sheet covering the brain, the “gray matter.”) This kind of connectivity is disrupted under general anesthesia (Alkire, 2008Golkowski et al., 2019).

That is, according to the first-order recurrent activation account, the active recurrent loops in perceptual areas plus background conditions are enough for conscious perceptual phenomenology. So long as high-level representations participate in those recurrent loops, conscious high-level content is assured. So if we assume the first-order recurrent activation account of consciousness, we can move from the high-level perceptual content shown by the agnosia evidence to high-level perceptual phenomenology.

I favor the first-order point of view. If the first-order point of view is right, it may be conscious phenomenology that promotes global broadcasting, something like the reverse of what the global workspace theory of consciousness supposes. (“Something like”: First-order phenomenology may be a causal factor in promoting global broadcasting; but according to the global workspace theory, global broadcasting constitutes consciousness rather than being caused by it.)

I believe that the same line of thought will apply to any neuroscientific theory of consciousness. All will have to agree that perceptual representation of motion in MT+ plus something else—e.g., certain relations to other brain activations or to behavior—are the basis of conscious experience of motion. It is difficult to imagine a remotely plausible candidate for the “something else” that applies to low-level representations such as activations in MT+ but does not apply to high-level activations such as activations in the fusiform face area. This point is the core of my argument for high-level phenomenology.

Some will reject the application of this idea to high-level phenomenology, but that rejection will have to be based on an independent doctrine that there is no high-level perceptual phenomenology. For example, Jesse Prinz’s AIR theory (attended intermediate level representation) holds that conscious perception is a matter of mid-level perceptual representations being modulated by attention in a way that allows for availability to working memory—albeit indirectly, via encoding of high-level representation (Prinz, 2012). (Working memory will be discussed later in this chapter and in Chapters 5 and 6.) My point here is that Prinz’s theory fits the mold I have described for low-level contents such as motion contents, and he is only able to avoid the extrapolation to high-level phenomenology by explicit stipulation that only mid-level representations can be conscious.

Representationists (also known as representationalists or intentionalists) among philosophers, such as Alex Byrne and Michael Tye, would also be committed to high-level perceptual phenomenology if they accept my argument that there is high-level perceptual representation (Byrne, 2001Tye, 2019). Representationists hold that the phenomenology of perception is grounded in or determined by its representational content, so high-level content requires high-level phenomenology. (They require further conditions, e.g., that the high-level representations are poised for a role in thought and report, but there is no reason to suppose that these conditions do not apply to high-level representations.)

In sum, many widely accepted perspectives in both neuroscience and philosophy support the move from the existence of high-level perceptual content to the conclusion that there is high-level perceptual phenomenology.

Adam Pautz has argued that the hypothesis of visually representing clusters of low-level properties is methodologically superior to the hypothesis of representing high-level properties. The low-level account is alleged to be more uniform, applying both to faces and to Byrne’s “greebles,” invented stimuli that are used to study object perception. And the low-level account is alleged to be more parsimonious since the high-level account postulates an extra layer. However, these arguments can’t explain the evidence of the sort just presented from agnosias for specific high-level representation. Further, as mentioned earlier, high-level perception has causal sources that are causally independent of low-level perception so high-level perception cannot be reduced to low-level perception.

Now we can return to the use of associative agnosias in arguing for high-level phenomenology. As I mentioned, Tim Bayne (2009) uses associative agnosia to argue for high-level phenomenology. He notes that associative agnosics have low-level perceptual phenomenology, but that it is “extremely plausible” to suppose that the result of the agnosia is that “the phenomenal character of his visual experience has changed” (p. 391). However, introspective judgments about high-level phenomenology are hard to agree on, as anyone who has argued with an opponent about them realizes. How sure can anyone be that they have the experience of visually attributing faceness as opposed to visually attributing colors, shapes, and textures?

Further, as Robert Briscoe (2015) points out, those who do not believe in high-level phenomenology (such as Prinz) will regard this argument as question-begging. They will think that the associative agnosics have lost recognition of high-level properties without having lost any alleged phenomenology that is specific to high-level properties.

Note however, that my argument is not subject to the problem that Briscoe raises. I have used associative agnosia and its difference from visual form agnosia to argue for high-level perceptual representation—without assuming anything about phenomenology or consciousness. Then I have added that all the main contenders for scientific accounts of how phenomenology relates to perceptual representation have the consequence that there is high-level perceptual phenomenology.

I think this two-step argument for high-level phenomenology is more convincing than other arguments with the same conclusion (Peacocke, 1983Siegel, 2010). At least it replaces reliance on intuitions about the phenomenology of experience with appeal to the scientific literature.

I am mentioning this issue in the introductory section of Chapter 1 because it illustrates the utility of the methodology of this book. I am discussing perception without consideration of consciousness or phenomenology, and then, once certain conclusions about perception are “established,” I will be arguing—in Chapter 13—that they provide arguments against “cognitive” theories of consciousness such as the global workspace and higher order theories.

More discussion of what the distinction between low-level and high-level comes to in terms of the visual hierarchy will come in the discussion in Chapter 2 and in the section of Chapter 4 on Bayesian inference.

I will be giving further arguments that there is high-level perception, though I do not think we have high-level representation of every observable property—observable in the sense that we can detect its presence perceptually. For example, we can often tell visually whether something is expensive but I doubt that expensiveness is visually represented. My main focus, however, is the difference between seeing something as, for example being a face, and forming a minimal immediate direct perceptual judgment that it is a face, the latter being the most perception-like of cognitive states. (More on this below.) The main issue of this book is what that difference is and why it is explanatorily important. The difference between high-level and low-level perception mainly comes in because it can often be hard to distinguish cognitive representations from high-level perceptual representations.

One feature of the scientific approach to perception that frustrates some readers is that there are observable properties that probably are not represented in perception. I just mentioned the property of expensiveness. Other such properties are being a baseball bat or a CD case. Endre Begby describes Burge’s view that probably we don’t perceptually represent these properties as “oddly reductive” (Begby, 2011). I think that what seems odd to many people is that we can tell pretty well by looking whether something is a baseball bat and whether it is a CD case. We can even do a visual search for baseball bats and CD cases. The reader should keep in mind that the issue here is whether, when one searches for a baseball bat, one is searching on the basis of low-level properties such as color, shape, and texture or whether a representation of baseball bats functions directly in visual search. This is a real experimental issue in many cases. For example Rufin van Rullen argues on the basis of experimental evidence that when one searches for a face in a scene, one uses representations of low-level properties and does not use a high-level face representation directly in search (VanRullen, 2006).

The cognitive state that is hardest to distinguish from a perception is a minimal immediate, direct perceptual judgment based on that perception. If I am perceiving a face, I often simultaneously judge on the basis of that perception that there is a face. When I speak of immediate perceptual judgment, I am talking about a cognitive state that is noninferentially triggered by perception. A direct perceptual judgment is based solely in the perception with no intermediary, inferential or otherwise. I can have a direct immediate perceptual judgment that the two lines in the Muller-Lyer illusion are unequal even when I know that they are actually equal. Thus in the case of known illusions, we have opposite cognitions. We have a perceptual judgment that the lines are unequal and knowledge that they are equal.

Of course, the content of an immediate perceptual judgment has to do with the way the world presents itself in perception. More specifically, what I mean is a minimal immediate direct perceptual judgment. A minimal perceptual judgment conceptualizes each representational aspect of a perception and no more. Later in this book I will use “perceptual judgment” to mean minimal immediate direct perceptual judgment, often leaving out everything but the “perceptual judgment.”

Given my definition of “perceptual judgment,” you may wonder whether there really is a difference between perception and perceptual judgment. This whole book is aimed at showing that there is such a difference and what that difference is. Chapter 2 concerns markers of perception that are not markers of perceptual judgment. And the penultimate section of Chapter 3, “Bias: perception vs. perceptual judgment,” gives a simple nontechnical example of how one could get empirical evidence that a mental state is a perceptual judgment rather than a perception. That issue stems from some shootings by police of unarmed Black men in which the police seemed to be saying that they “saw” a gun. This section presents one item of evidence that such reported states were perceptual judgments rather than perceptions. Readers who are especially interested in that issue could read that section now, since it does not presuppose anything in the book that comes before it.

A concept of something in my terminology is a representation that functions to provide a way of thinking of what it is a concept of. Oedipus had at least two concepts—i.e., two ways of thinking of his mother, one we could describe as “Mom,” and another when he married her (perhaps Greek analogs of “Jocasta” or “Sweetie”). He then found out that, tragically, the two ways of thinking of her, the two concepts of her, were ways of thinking of the same person. Some uses of the term “concept” link it to discrimination and others to categorization (Connolly, 2011). As we will see in Chapter 6, there is perceptual categorization in the absence of the ability to use a representation in thought and reasoning. In my terminology, these cases exhibit nonconceptual perceptual categorization. The ways of representing that are involved in concepts play a role in propositional thought or reasoning—that is definitional in my use of the term “concept.”

What is meant by “cognition”? A prescientific characterization is that it constitutively involves capacities for propositional thought, reasoning, planning, evaluating, and decision-making. (I won’t be talking about the conative attitudes [e.g., wanting] or emotions.) Some insects—for example, the nonsocial wasp—can perceive but are not good candidates for propositional thought. Perceptual judgment (i.e., minimal immediate direct perceptual judgment) can be unconscious and automatic. (I will have little to say about emotions, moods, and other mental states that are neither a matter of just perception nor just cognition.)

My usage of the terms “cognition” and “perception” is consonant with much of the recent literature on perception (Firestone & Scholl, 2016aPylyshyn, 1999), but some restrict the term “perception” to what I am calling low-level perception (Linton, 2017) and others use “cognition” to encompass mid-level and high-level perception as both perception and cognition (Cavanagh, 2011). This usage is responsible for the term “visual cognition” to mean mid- and high-level vision. For example, here is how Patrick Cavanagh defines “visual cognition” (Cavanagh, 2011, p. 1538):

A critical component of vision is the creation of visual entities, representations of surfaces and objects that do not change the base data of the visual scene but change which parts we see as belonging together and how they are arrayed in depth. Whether seeing a set of dots as a familiar letter, an arrangement of stars as a connected shape or the space within a contour as a filled volume that may or may not connect with the outside space, the entity that is constructed is unified in our mind even if not in the image. The construction of these entities is the task of visual cognition.

Cavanagh seems to exclude low-level perception with the phrase “base data,” and the example of a familiar letter brings in high-level perception.

Consider the Necker cube (pictured in Chapter 2). It can be seen as having a front surface pointing to the lower left or, alternatively, the upper right. These are mid-level differences since they have to do with computation of surface representations. By Cavanagh’s definition, the difference counts as cognitive—and also a difference in mid-level vision. This way of talking would preclude a joint in nature between perception and cognition, so I will stick with an anchor for “cognition” in propositional thought, reasoning, planning, evaluating, and decision-making. (See the section on conceptual engineering later in this chapter for a discussion of principles of clarifying the notions of perception and cognition.)

Conceptions of cognition that link it to thinking and reasoning have been called the conservative approach by Cecilia Heyes (Bayne et al., 2019). Of the conservative view, she says (p. R611) it “has a venerable history in Western thought but it’s out of kilter with contemporary scientific practice. It implies that much of the research done by those who identify as cognitive scientists—for example, work on the behaviour of plants, shoals of fish and swarms of bees—has nothing to do with cognition.”

Although using the term “cognition” in the “liberal” manner that Heyes prefers may have some advantages and it might identify some kind of a joint in nature, it won’t be the joint between seeing and thinking. Individual bees may have some thought, but the swarms of bees and shoals of fish don’t and plants don’t.

Perhaps the most significant difficulty for the perception/cognition joint has been “core cognition,” systems that appear to constitute a mental kind that is both paradigmatically perceptual and paradigmatically cognitive. The existence of a third category does not in itself impugn a joint. Glasses are normally rigid in the manner of solids but have the amorphous structure of liquids. Glasses do flow, more quickly at some temperatures than others, and they do not have the “shear” properties of solids. Glasses have a level of molecular organization intermediate between that of liquids and that of solids (Curtin, 2007). Indeed, one newly discovered form of glass resembles solids in that the molecules cannot change orientation but resembles liquids in that the molecules can move freely (in the sense of translational motion) in all directions (Roller, Laganapan, Meijer, Fuchs, & Zumbusch, 2021). No one would think that the existence of these intermediate types impugns the important explanatory significance of the division of matter into gas, liquid, and solid.

However, the problem posed by core cognition goes beyond providing a third category. The problem is that perception and cognition each are defined by a set of properties that have their own explanatory unity and core cognition purports to spoil that unity by combining fundamental features of both perception and cognition.

Two of the most dramatic cases of core cognition are mental representation of causation and approximate numerosity (Carey, 2009Carey & Spelke, 1994). I will argue in Chapter 12 that in these cases there are perceptual analogs of our cognitive representations of causation and numerosity and that some of the key phenomena can be seen as simply perceptual. Also, some can be seen as cognitive phenomena that utilize perceptual materials, as when we think about something by thinking about what it looks like. In their writings on core cognition, Susan Carey and Elizabeth Spelke have advanced a view of core cognition as combining both perceptual and cognitive features in a single type of representation. I will be arguing for a different view, in which the category of “core cognition” is best thought of as mixing fundamentally different kinds of representations in a single system. In a homogeneous mixture, like air or sugar water, the distinct kinds are not easily discernable. In a heterogeneous mixture, like smoke (gases plus particles) or salad dressing (oil plus vinegar), there are different kinds that are discernable. I am saying that core cognition may be a heterogeneous mixture of perception and cognition.

There are other heterogeneous mixtures of perception and cognition. Here is an example of the use of perceptual materials in cognition: when we use a perceptual simulation involving perceptual representation of a sofa and a doorway to think about whether the sofa will fit through the doorway. It is well established that we use the mental imagery system to do something often described as “mental rotation” in such reasoning, and the mechanisms of mental imagery are substantially (but not completely [Kaski, 2002]) shared with perception (Block, 1983a). Another example: the use of perceptual color representations to consider the question of which green color is darker, that of a Christmas tree or that of a frozen pea. In this case, the reasoner uses two different perceptual concepts of color shades. These perceptual concepts are conceptualized percepts, i.e., percepts that have been incorporated into a conceptual structure. Another type of conceptualized percept would be a descriptive concept with a slot for perceptual materials (Balog, 2009bBlock, 2006Papineau, 2002).

But why don’t these examples and the existence of conceptualized percepts show there is no joint in nature? The answer is that the joint is between perception on the one hand and types of states that may use perceptual materials, but not constitutively, on the other. This argument will be explored in greater detail later in the section Conceptual Engineering (p. 46). I will argue that perception should be restricted to what might be called pure perception, perception that does not occur as part of a judgment or as part of a working memory representation. (Working memory is a mental scratch pad that will be described in detail later, mainly in Chapters 5 and 6.) Pure perception is perception without any cognitive envelope.

Perceptual simulations used in cognition do not have all the properties of perception. I will present evidence in Chapter 2 that a basic kind of competition between perceptual representations that is characteristic of perception does not obtain with perceptual representations in working memory. Further, there is reason to believe that the perceptual representations of colors used in working memory representations do not allow for fine-grained colors, the minimal shades of perception.1 Finally, perceptual simulations used in working memory do not have the phenomenology characteristic of conscious perception. (Compare hearing a phone number and rehearsing it while you look for your phone.) These three points indicate deep differences between perception and perceptual materials used in working memory.

One consequence of the thesis of this book is that certain kinds of machine “vision” are not really cases of vision or even of perception. This issue will be explored in Chapter 6, where I will discuss a robot whose cameras output propositional representations concerning the properties of pixels in the camera’s sensor. If those representations are treated as premises in inferences on the basis of which the robot infers the environmental causes of those stimulations, then the robot would have direct sensory production of thought without perception.

One upshot of the ideas in this book for epistemology is that it is a mistake to think of the way perception justifies perceptual belief in terms of inference from perceptual contents to belief contents. Another upshot has to do with intentionality. The kind of intentionality involved in perception is different from although perhaps the source of the intentionality of belief (Neander, 2017Shea, 2018).

Consciousness

As noted earlier, throughout most of this book, I will be concerned with perception rather than conscious perception. The border I am talking about is the border between perception—whether conscious or unconscious—and cognition—whether conscious or unconscious. In part, this focus stems from the perception literature (cf. Teufel & Nanay, 2017). Much of the experimental work on perception and cognition does not address the issue of whether the effects are conscious or unconscious. But in part this focus reflects the view that perception is a natural kind that includes both conscious and unconscious perception (see Block, 2016aBlock & Phillips, 2016; and Phillips, 2015, for both sides on this view). (A similar idea applies to the relation between conscious and unconscious control of behavior (Suhler & Churchland, 2009).)

In endorsing the methodological priority of perception over conscious perception, I am not endorsing a view that perception is metaphysically and explanatorily prior (Miracchi, 2017) to consciousness. There are many conscious states that are not perceptual. Perception and consciousness are overlapping natural kinds, neither of which has priority over the other.

Of course the phenomenology of perception is conscious and is part of . . . perception! Some visual phenomena occur in both conscious and unconscious perception, for example the phenomenon of binocular rivalry to be described in Chapter 2. Also, illusory contours probably are present in unconscious vision, since, as will be described in Chapter 2, they occur in insects, and occur so early in visual processing in primates that they probably partially precede conscious perception. However, other perceptual phenomena may only occur in conscious vision, for example, perceptual pop-out.

In Chapter 6 I will argue for nonconceptual color perception in children. Then in Chapter 13 I will leverage that point to argue that these children have phenomenal-consciousness of color without access-consciousness of color.

Pure perception

Any perception that we notice will inevitably involve conceptualization—as part of the cognitive process of noticing the perception. But there may be creatures whose perceptions are always pure, never involving cognition. Wasps have well-developed visual systems. For example, wasp vision exhibits many of the standard visual constancies. Peter Godfrey-Smith has noted that despite their excellent vision, wasps have short lives with highly stereotyped action patterns and show little or no sign of a cognitive or affective life (Godfrey-Smith, 2017).

This is true of the solitary wasp species, for example the sphex wasp. There are thousands of wasp species, with the social wasps being more sophisticated. One social wasp (the paper wasp) has been shown to be capable of a kind of transitive inference-like behavior (Tibbetts, Agudelo, Pandit, & Riojas, 2019). Bees have been shown to learn matching relations between numerosities of 3, 2, and 1 items and symbols, and also between symbols and numerosities (Howard, Avarguès-Weber, Garcia, Greentree, & Dyer, 2019a). But bees trained on matching symbols to numerosities did not generalize to matching numerosities to symbols. And conversely. Although bees have been shown capable of such kinds of symbolic associations with number, they failed the transitivity test. Tibbetts et al. speculate that sensitivity to transitive relations evolved in the paper wasp because its social structure involves dominance hierarchies among individuals. Bees do not have social ranks of the same sort, and of course solitary wasps have no social ranks. My discussion below concerns solitary wasps, not social wasps.

Jumping spiders have been shown to distinguish between biological and nonbiological motion (De Agrò, Rößler, Kim, & Shamble, 2021). De Agrò et al. used displays of 11 moving dots. If the dots are on human joints, displays look like moving humans to human observers. De Agrò et al. constructed similar displays based on spider joints and presented spiders with those displays as well as displays based on moving geometrical figures and random displays. They found that the spiders looked longer at the nonbiological motion. This difference may reflect high-level perception as shown in humans of biological motion. Alternatively, it could reflect cognition.

Flying insects are subject to evolutionary pressure to reduce the size of their brains (Godfrey-Smith, 2017). Godfrey-Smith speculates (in a talk at NYU) that the lifestyle of wasps, combining excellent vision, short lifespan, and limited capacity for learning, pairs with perceptually controlled stereotyped action sequences rather than cognition.

Wasps are capable of “classical” (“Pavlovian”) associative conditioning but have not been shown to be capable of the kind of instrumental or operant conditioning in which a reward (or punishment) affects the probability or strength of the behavior that led to the reward (Godfrey-Smith, 2017Perry, Barron, & Cheng, 2013). In classical conditioning, a light stimulus paired with food can lead to an appetitive response in the presence of the light without the food. In operant (or “instrumental”) conditioning, a voluntary response that leads to food is then used by the animal to get food, as when a rat learns to press a bar to get a reward. Operant conditioning is more complex than classical conditioning; indeed it is usually considered to include classical conditioning as a component (Perry et al., 2013). (The pairing of the food with the bar press elicits an appetitive response that the animal voluntarily harnesses to get food.) A creature that does not have operant conditioning may have no cognition either.

The stereotyped action patterns of the sphex wasp were famously described by Wooldridge (Dennett, 1984Hofstadter, 1979Wooldridge, 1963).

When the time comes for egg laying, the wasp Sphex builds a burrow for the purpose and seeks a cricket which she stings in such a way as to paralyze but not kill it. She drags the cricket into the burrow, lays her eggs alongside, closes the burrow, then flies away, never to return. In due course, the eggs hatch and the wasp grubs feed off the paralyzed cricket, which has not decayed, having been kept in the wasp equivalent of a deepfreeze. To the human mind, such an elaborately organized and seemingly purposeful routine conveys a convincing flavor of logic and thoughtfulness—until more details are examined. For example, the wasp’s routine is to bring the paralyzed cricket to the burrow, leave it on the threshold, go inside to see that all is well, emerge, and then drag the cricket in. If, while the wasp is inside making her preliminary inspection, the cricket is moved a few inches away, the wasp, on emerging from the burrow, will bring the cricket back to the threshold, but not inside, and will then repeat the preparatory procedure of entering the burrow to see that everything is all right. If again the cricket is removed a few inches while the wasp is inside, once again the wasp will move the cricket up to the threshold and re-enter the burrow for a final check. The wasp never thinks of pulling the cricket straight in. On one occasion this procedure was repeated forty times, always with the same result.” (Wooldridge, 1963, pp. 82–83)2

Sphex wasps lay their eggs in crickets but apparently are not able to count the crickets they have placed in the nest. They drag the crickets into the nest by the antennae but if the antennae are cut off they do not drag the cricket in by the legs.

Wasps show no sign of “wound tending,” in which a damaged animal will protect and groom the affected part, making them good candidates for zombies with no phenomenal consciousness (Godfrey-Smith, 2016b). If a wasp tears a part of its body, it carries on, taking no notice of and without favoring the affected part. This suggests that wasps do not feel pain. Of course, the lack of pain mechanisms is compatible with other forms of consciousness attaching to mechanisms that wasps do have—such as perception. But given that consciousness may have evolved originally as a motivating force, the lack of that force in its most salient use is an indication of lack of consciousness. Bees have been shown not to prefer water with morphine in it when part of a leg is chopped off, though as far as I know this has not been tested with wasps (Groening, Venini, & Srinivasan, 2017). Still, the result with bees does suggest that pain may not be part of insect physiology.

Ants show “social” wound tending (Frank, Wehrhahn, & Linsenmair, 2018) in the sense that nest-mates groom the injuries of other ants, greatly decreasing the chance of infection. Interestingly, lightly wounded ants will adjust their behavior, acting as if their injuries are worse than they are around nest-mates. Note that social wound tending does not indicate conscious pain on the part of the wounded ant.

Given all this, the likelihood that (solitary) wasps have anything that we could call concepts seems low, so they are excellent candidates for nonconceptual perception. (Cf. Burge, 2010aPeacocke, 2001a). (Disagreements about related issues can be found in Bermudez, 1994, and Peacocke, 1992b2002.) At the very least, the susceptibility to evidence of the claim that the wasps have perception without cognition shows it is not a conceptual truth that perception requires concepts or cognition. Of course I am assuming that there is no hidden contradiction in the claim that solitary wasps have perception without concepts or cognition.

To avoid misunderstanding, note that I am not arguing that we know or even have very good reason to think that wasps have perception without concepts, cognition, or consciousness. Rather, wasps are candidates for such a status.

I expect that some philosophers will be tempted to deny that wasps perceive at all, if it really is true that their perception is not conscious and if they have no cognitive states that could be thought of as perceptual judgments. However, common sense and science both tell us that wasps perceive (Burge, 2010b). The wasp has a visual system that in rough outline is like ours. It uses its eyes to see its prey and uses vision to track and ambush it. The wasp sees that the cricket is no longer on the threshold of the burrow and that is why the wasp moves the cricket to the threshold. There have been debates about whether humans have unconscious perception, but those debates revolve around the issue of whether unconscious perceptual representations are at the “individual” level or whether they are subindividual (Block & Phillips, 2016Peters, Kentridge, Phillips, & Block, 2017Phillips, 2018). In the case of the wasp, I don’t see how it could be claimed that the wasp’s visual states are not at the individual level, given their role in guiding the wasp’s actions.

If I am right that wasps perceive colors and shapes and that their perceptions have color contents and shape contents, then there is a kind of perception that is not dependent on or derived from concepts.

As we will see in Chapter 5, there are some views of concepts that can accept that the Sphex wasp has no cognition but still maintain that it has perceptual concepts in virtue of the wasp’s ability to perceptually identify and track objects (Green & Quilty-Dunn, 2017Mandelbaum, 2017Quilty-Dunn, 2020). The disagreement here is not merely verbal but reflects different views of what empirical science tells us are the right categories for understanding the mind. I do agree with Green, Quilty-Dunn, and Mandelbaum, though, and disagree with advocates of a cognitive view of concepts (Camp, 2009) on the importance of stimulus independence as a rough guide to what representations are concepts. More on that topic later in this chapter.

What is a joint?

A joint is a fundamental and explanatorily significant difference between the kinds that are separated by the joint. Perception and cognition function differently in the mental economies of organisms that have them and in how they themselves are to be explained. For example, cognitive states—but not perceptual states—are formed by processes of reasoning and affect other states by processes of reasoning. The formation of perceptual states—but not cognitive states—involves direct effects of “constancy” mechanisms (Burge, 2010a). Perception functions to provide us with information about what is happening in the nearby environment now, whereas cognition functions in reasoning about the news provided by perception so as to decide what to do and to plan for the future.

To understand a difference in kinds, it may help to know what a kind is. There are many theoretical disputes about kinds (Franklin-Hall, 2015Kitcher, 2007Taylor, 2020, 2022). Rather than start with any theoretical position on what a joint is, I propose to frame the discussion of joints around actual examples of joints. Examples: gravitational/electromagnetic forces, hydrogen/helium, lepton/quark, liquid/solid, animals/plants. These examples show us that a joint is compatible with considerable causal interaction across the joint. Plants and animals interact and even co-evolved. For example, colors of flowers co-evolved with color vision in bees. And as mentioned earlier, a joint is compatible with intermediate or indeterminate cases. Red algae, mushrooms, slime mold, and giant kelp are neither plants nor animals. Viruses are neither alive like animals nor inanimate like rocks. Joints can be a matter of clusters of properties. For example, both animals and plants have vacuoles, storage bags of the cells, but plants tend to have one large vacuole whereas animal cells have one or more small vacuoles. There are some binary differences though. Both animal and plant cells have a cell membrane but only plants have a cell wall. And animal cells do not have chloroplasts, except temporarily when they eat plants. Although the joint between liquids and solids is explanatorily important, its role in explanations in physics and chemistry is background not foreground in current disputes and I expect that the same is true of the joint between perception and cognition.

Perception is iconic, nonconceptual, and nonpropositional, constitutively, or at least these features are explanatorily deep properties of perception. (See the next section.) These properties are necessary but not jointly sufficient for perception. For example, hallucination has all these properties but is not perception. And imaginings—visual simulations— with all three properties can be used to decide which peg fits in the hole. Perhaps the cognitive process of deciding what fits in the hole could not have occurred without the visual simulation and in that sense that particular token cognitive process essentially has all three properties. But cognition per se does not require any of the three properties.

I will list cases of the three properties that are not cases of perception in this footnote so as to easily refer back to it later.3 The project of this book is to elucidate some characteristics that are fundamental to perception that are not fundamental to cognition. I am not aiming at conditions that are both necessary and sufficient for perception—nor conditions that are necessary and sufficient for cognition.

Although I am not aiming at conditions that are both necessary and sufficient I am not ruling out such conditions at least for perception. And I firmly reject views such as the “cluster concept” view of perception that are incompatible with necessary and sufficient conditions.

William Alston offered a cluster concept view of the concept of religion (Alston, 1967). He noted a number of elements that are common to religions, belief in a supreme being, a distinction between sacred and profane, rituals and morality based in the sacred, religious feelings, prayer, a worldview, an organization of life based on the worldview, and a social group based on elements of the cluster. As he noted, for each these elements, there are actual and possible religions that lack that element. For example, some forms of Buddhism do not involve belief in a supreme being and Quakers have no sacred objects.

My objection to the cluster concept account of perception is that it does not allow for necessary conditions. I have the same objection to Richard Boyd’s homeostatic property cluster view of natural kinds (Boyd, 19891991Taylor, 2020). As I will be arguing, iconicity, nonconceptuality, and nonpropositionality are necessary and fundamental to perception.

The architecture of a computer is a matter of relatively fixed structural organization of computational components. It is commonly said that architecture is a matter of hardware as opposed to software. That isn’t quite right since Macs and PCs have different hardwares but can both be programmed to have modules that only accept one kind of input, say an input with a certain tag.

One architectural division of the mind is that there are properties that are representable by cognition but not perception. For example, cognition but not perception can represent that an action is justified (Green, 2020a). Arguably, there are features of the world that perception can represent, e.g., according to some philosophers but not others, fine-grained colors are not representable by cognition (Evans, 1982McDowell, 1994Raffman, 1995). It isn’t clear how deep these differences are. One can certainly imagine a creature whose cognition and perception are more aligned than ours.

I will be arguing that there are deeper differences in format and type of state that would still persist even if there were no differences in what properties are represented.

Constitutivity vs. explanatory depth

I am claiming that iconicity, nonconceptuality and nonpropositionality are constitutive, or at least explanatorily deep properties of perception. What is the difference between constitutivity and explanatory depth? The difference is not very important to the themes of this book, so I will be brief. There is a case for constitutivity but a stronger case for explanatory depth.

The difference between what is constitutive of perception and what is an explanatorily deep property of it concerns what kind of a concept the concept of perception is. I have tried to steer clear of arguments based solely on “intuitions” in this book, but when the question is what kind of a concept our concept is, intuitions are relevant. I’ll focus on two alternatives: whether perception is a natural kind concept or a functional concept. These are not the only possibilities, but they are the best candidates.

Issues about concepts can be discussed in terms of the words that express those concepts. This issue about the concept of perception can be discussed in terms of the question of whether the word “perception” has a “natural kind” or a “functional kind” semantics. Hilary Putnam famously imagined that there could be a planet, “Twin Earth,” in which a there is a chemical XYZ that has no hydrogen or oxygen but is still “watery” in Chalmers’s (1996) sense of having the observable properties of water, including being colorless and odorless, sustaining life, falling from the sky, and occupying rivers and streams and oceans (Putnam, 1975). Although the denizens of Twin Earth—whom we can imagine to be neural duplicates of us—use their term “water” to refer to the watery XYZ in their oceans, our word “water” and our concept of water do not apply to XYZ. We should say rather that the liquid in their oceans is the functional equivalent of water but is not water. If this is right, the concept of water is said to be a natural kind concept and the semantics of the term “water” involves natural kind semantics.

The natural kind semantics of “water” contrasts with the functional kind semantics of “mousetrap” and “philosopher.” The mousetraps and philosophers of Twin Earth might be constructed differently from and composed of different materials from our mousetraps and philosophers, but if they are the functional equivalent of our mousetraps and philosophers, then they are genuine mousetraps and philosophers. Philosophers who disagree with Putnam’s view that “water” has a natural kind semantics can hold that “water” has a functional semantics—any watery liquid is water.

Metaphysicians distinguish between possible worlds considered as counterfactual and possible worlds considered as actual (Chalmers, 1996Davies & Humberstone, 1980). I will explain—assuming Putnamian views about water as a natural kind. In considering Putnam’s Twin Earth world as counterfactual, we hold fixed the actual composition of water, considering whether the stuff in the counterfactual ocean is water. The stuff in the counterfactual ocean is water if and only if it is H2O. We can say that the term “water” is metaphysically “rigid,” denoting water (i.e., H2O) in every world in which it exists. When we think about a counterfactual world, we use our term “water,” anchored in H2O, to consider whether the world contains water. If the oceans in the counterfactual world are filled with XYZ and not H2O, they are not filled with water.

A world considered as actual is an actual world candidate. Instead of holding the actual composition of water fixed, we ask ourselves what we should say if we discovered that the actual watery stuff in our oceans is XYZ. We should say that water turns out to be XYZ. We can put this by saying that the term “water” is epistemically nonrigid in that it does not denote the same stuff in every world considered as actual. In Putnam’s Twin Earth considered as an actual world candidate instead of as a counterfactual world, “water” would refer to XYZ. See Chalmers (2012a, pp. 238ff, 318ff).

Although XYZ is not water, if Twin Earth is merely distant in space rather than being counterfactual, and if we had regular travel between earth and Twin Earth, we might broaden the concept we attach to the word “water” to cover both H2O and XYZ. Since I am a Putnamian who thinks “water” has a default natural kind semantics, I think this would be a change of meaning and concept. The changed concept would dictate that whatever functions like actual water is water.

In my view, “perception” is like “water” in having a default natural kind semantics. When I discussed my book at Chris Hill’s and Adam Pautz’s seminar at Brown in October 2020, Pautz pressed me on this, both in the seminar and in our discussions afterward, arguing that I had not given reason to prefer the view that perception is a natural kind to the view that it is a functional kind.

How can the reader decide whether “perception” has a natural kind or functional semantics? (I am ignoring the possibility both are wrong.) A good way to proceed is to consider what to say about various candidates for perceivers. Would creatures that process stimulation in a way that is roughly functionally equivalent to our visual perception but via representations that are conceptual, propositional, and discursive be seeing or even perceiving? I say no, but a functionalist about perception should say yes.

Let us focus on the claim I will be making in Chapters 4 and 5 that perception is constitutively nonpropositional and iconic. One consideration that I will be appealing to in arguing for the constitutive nonpropositionality of seeing is that the contents of perception cannot be logically complex. We can see something as a mixture of red and blue or as indeterminate between red and blue but not as simply red or simply blue, i.e., as having the disjunctive property of being simply red or simply blue. We cannot see something as if red then round, that is, as having the conditional property of being round if red. And, I will argue, that although this case is less straightforward, that perceptual contents cannot be negative and they cannot be conjunctive.

Consider the possibility of a robot that has digital cameras whose outputs are pixel array representations that the robot treats as premises in inferences and uses them to reason about the causes in the environment of those arrays. Let us suppose that the robot can “visually” represent something as having the conditional property of being if red, then round. The robot can also “visually” represent something as having the disjunctive property of being red or blue. And let us suppose that it “visually” represents the disjunctive property using discursive representations rather than iconic representations. For concreteness, we could suppose the representations are strings of words. Would that robot have visual perception? I would describe the robot as having light transducers that directly produce discursive thought, skipping the stage of perception.

To be clear, the issue concerns seeing-as, not seeing-that. Seeing-that includes the products of inference from seeing-as. As Fred Dretske once commented, I can see that the gas tank is empty by looking at the gas gauge, even though the gas tank is not visible. Seeing-that can have logically complex contents, inferred from seeing-as. I can see that either the tank is empty or the gas gauge is broken.

I am not talking about seeing-that in that sense. I am talking about perceptual attribution, i.e., seeing-as; and in particular, the putative perceptual attribution of the conditional property of being if red, then round. It is in that sense that I am saying it is wrong to suppose anyone or anything could literally see something as if red, then round, or as having the disjunctive property of being simply red or simply blue.

I am appealing here to my own intuitions about the application of the concept of seeing. Imagine a conversation with the robot who claims to see the moon as if gray, then round. You say “You mean you see it as gray and then expect it to be round? Or do you mean you see it as gray and also as round?” It says “No neither of those is what I mean. I see it as having a conditional property.”

I haven’t said anything about whether the robot is conscious, but if one brings in consciousness, the point seems more convincing. If we ask the robot, “You mean it looks conditional?” And it answers, “Yes, it looks this way: if gray, then round,” one might naturally conclude that the robot is misapplying the concept of seeing.

It is hard to imagine what it would be like to see something as if gray, then round. It may be said, though, that this point about visual imagination can be explained by appeal to the deep explanatory nature of perception rather than its constitutive properties. I can’t imagine seeing with conditional contents, because in trying to imagine it I must use my visual system, which is not capable of conditional contents. So, according to the objection, my failure of imagination is due to a deep explanatory property of human perception, not a constitutive property of all perception.

In reply I say that although I am incapable of bat-style sonar, I can in some reasonable sense understand it. So I doubt that it is restrictions on imagining that make the robot just described seem to be a nonperceiver.

How can we distinguish between the actual properties of perception that are constitutive and those that are not? One method is to use thought experiments like the one described above involving robots that process stimuli in a way that deviates from actual processing of stimuli. We see via light, electromagnetic radiation with wavelengths of 400–700 nm (a nanometer being a billionth of a meter). But one can imagine a robot with a visual system much like ours that sees or at least perceives via electromagnetic radiation of much shorter (e.g., X-rays) or longer wavelengths (e.g., microwave radiation). And one can imagine perception using many other propagating signals.

Compare the case of our building robots whose vision uses X-Rays instead of light with our chemists producing small quantities of XYZ. In the former case, one would be imagining robots that see but in the latter case we would be imagining watery stuff that is not water.

Just as water being identical to H2O precludes a watery substance that is not H2O from being water, so the fact that perception is a form of iconic and nonpropositional representation precludes the robot described as having visual perception. Deciding to call the robot camera system “perception” would be deciding on a change of meaning.

I agree that it may be that the language would develop so as to give “perception” a functional kind semantics, but I believe that if it does, the term “perception” will have changed its meaning so as to denote a functional kind, just as “water” would change its meaning if we were to merge language communities with the residents of Twin Earth. Whether the language does develop in that way I think will depend on the social role of robots. If their social role involves our conceiving of them as not at all like us, then we may think of their cameras as writing directly into thought, skipping the stage of seeing. But if they are regarded as just a different kind of person, then it will be natural to conceptually engineer the term “visual perception” to encompass the conversion of light stimulation into cognitive representations.

Sometimes the issue of whether we are referring to a natural kind or a functional kind is framed in terms of the intentions of normal speakers of the language in question. David Lewis held that natural kinds are “reference magnets” and this has been explained in terms of the intentions of speakers. If speakers intend to refer to natural properties, then speakers must have the concept of naturalness, which seems most unlikely. Even less likely is the supposition that ordinary speakers who talk about seeing have the concepts of format, concept, or proposition. It cannot be that ordinary users of words like “see” and “perceive” intend to refer to a capacity with a certain format or a capacity that is nonconceptual.

What makes much more sense is that speakers’ implicit intentions drive them in the direction of treating natural properties as the reference of kind-terms, even though subjects do not have the requisite concepts. Thus, naturalness would have a theory-external role, as Lewis thought (Chalmers, 2012bLewis, 1984Sider, 2011).

Although I am defending the idea that “perception” and in particular “vision” has a natural kind semantics, I don’t have a high level of certainty about this. As Gareth Evans emphasized (1982), there is a good deal of indeterminacy in the referential intentions of common-sense reasoning. For this reason, I want to emphasize my fall-back position that the properties of perception that I am delineating are explanatorily deep rather than constitutive.4

The contents of perception

Views of perception differ in whether they acknowledge that perception has representational content. Many “naïve realists” think of perception as a direct awareness relation to objects and properties in the world (Brewer, 2011Campbell, 2002). Some naïve realists allow representational content in unconscious subpersonal states (Travis, 2004). Naïve realists struggle with perceptual illusions in which we perceptually misrepresent the world. Naïve realists also struggle with the effects of attention on perception, in which attention makes stimuli look higher in contrast, faster, stripier (higher in “spatial frequency”), or higher in color saturation (Block, 2019aBrewer, 2019).

I have discussed naïve realism elsewhere (Block, 20102019a) and will not be repeating those discussions here. Except for a short discussion in Chapter 13, I won’t be discussing naïve realism here since I am presupposing an orientation shaped by the science of perception. The science of perception is deeply committed to perceptual representation (though see French and Phillips, 2020). In the next section, I will give a few examples of how practice in vision science presupposes representation and representational content.

One argument for perception having content appeals to perceptual experience. Susanna Siegel notes that there is a powerful argument from the fact that something looks red to the conclusion that the perception is accurate if and only if what is seen is red (Siegel, 2010), in which case the accuracy condition can just be taken to be the content of the perception.

The content of a representation, as I am using the term, is the way it represents the world to be. Contents are indexed by accuracy conditions. My perception as of a red object at a certain location represents a red object at a certain location and is accurate just in case there is a red object in that location (Siegel, 20062016). Note that the category of accuracy conditions is wider than that of truth conditions. For example, noun phrases have accuracy conditions but not truth conditions. The noun phrase “that red object in that location” is accurate just in case the “that” singles out a red object in that location. Although I reject the claim that perceptual content is constitutively singular, I do agree with Crane and Burge that the content of perception is more like the content of a noun phrase than like the content of a sentence (Burge, 2010aCrane, 2009).

The reader may wonder what the cash value is of the difference between “That red object” and “That is a red object.” Tim Crane says what is important is that accuracy admits of degrees but truth does not (Crane, 2009, p. 458: “Accuracy is not truth, since accuracy admits of degrees and truth does not”). This is not my view. Degree-related talk is somewhat more natural for accuracy than for truth, but this is not a principled difference. One could also use degree-related talk for truth. For example, people describe one theory as a closer approximation to the truth than another.

What is important from my point of view is that propositional representations can be used in content-based transitions in which the representations can serve as premises or conclusions in reasoning. The use of “is” and the term “true” as opposed to “accurate” are markers for that deeper difference. (By reasoning I mean a certain kind of content-based rational transition in thought in which propositional representation can serve as premises or conclusions.)

Of course, one can move from a perception that something is red to the judgment that it is red. Why isn’t that “move” the right kind of content-based rational transition in thought? That issue is the burden of Chapter 4 and to some extent Chapter 6.

It is often assumed that the content of perception is propositional content, that is, that in perception one bears the perception relation to a proposition (Byrne, 2005Pautz, 2009). This view was famously endorsed by John McDowell, who claimed that the propositional and conceptual nature of perceptual content is required in order to understand how perception can justify perceptual belief when a perceptual experience is taken at face value (McDowell, 1994). McDowell says (p. 26), “That things are thus and so is the content of the experience, and it can also be the content of a judgement: it becomes the content of a judgement if the subject decides to take the experience at face value. So it is conceptual content.” (See McDowell, 2019, for a somewhat different view.) However, epistemologists should want a conception of perceptual content that is not contradicted by the science of perception.

A distinction is often made between intrinsic and derived intentionality (Haugeland, 1980Searle, 1980). Words on paper have derived intentionality in that their representational contents depend on the minds that use those representations. The representations that the minds use do not themselves depend for their contents on representations other than other mental representations, so they have intrinsic intentionality.5 If there were no minds, marks on paper that happened to look just like words would have no content. Indeed, there was no doubt a time in the history of the human race when there was mental representation without external representation. In these terms, the content of perception is intrinsic, not derived.

The treatment of perceptual contents in terms of accuracy conditions contrasts with treatments of perceptual contents that are grounded in the phenomenology of conscious perception, for example the “appears/looks” conception and the view that the content of a perception consists in the properties that the object presents to us (Chalmers, 2006Pautz, 2009). When an unconscious perception becomes conscious, a content of an unconscious perception becomes a content or part of a content of a conscious perception, so the most basic account of the content of perception will apply both to conscious and unconscious perception. Many accounts of perceptual content do not even purport to apply to all perception, conscious and unconscious. For example, Adam Pautz’s “identity” conception of content says that experiential properties are relations to contents, thereby giving a theory that does not apply to both conscious and unconscious perception (Pautz, 2009).

Fred Dretske (20072010) argued that there was a kind of seeing that is nonconceptual and requires no noticing, attending, or classification of the item seen. He contrasted this “simple” seeing with fact-seeing or seeing-that. I do not recognize any such kind of seeing, though I think we can speak of seeing in a way that makes no commitment to classification. I think all seeing is seeing-as (Block, 2014c) in that vision always attributes properties. (This issue is taken up in Chapter 3.) I do agree with Dretske that seeing is nonconceptual and requires no noticing. (The jury is out on attending.) But I also think that seeing often involves something that can be called categorization, something over and above the mere attribution of properties. This thesis will be spelled out and evidence provided in Chapter 6 in discussing categorical perception of color without concepts of color.

Realism about perceptual and cognitive representation

I will be assuming that there are genuine perceptual representations. I will give some arguments for representations here and there will be a brief discussion of opposed views such as naïve realism and enactivism in the context of debates about consciousness in Chapter 13.

I am a realist about representation and representational content—in contrast to J. J. Gibson (1979) on representation generally and Frances Egan (2014) on representational content. I do not, however, assume that these mental representations have their contents essentially—that is, that a mental representation could not change its content without becoming a different representation or that if a mental representation had had a different content it would have been a different representation. And I will not be assuming that there is a reductive account of representation. So I am not assuming hyperrepresentationalism about mental content in Egan’s terms (2014, 2018).

In my view, realism about mental representation and representational content is baked into the practice of perceptual psychology and cognitive science, as is especially clear in discussions of illusory perception. Where there is misrepresentation, there is representation because perceptual misrepresentation shows that perceptual representation cannot be thought of along the lines of Gricean “natural meaning.” As Grice (1957) noted, we can say that the spots mean measles, but it would be contradictory to say that the spots mean measles but the spotted person doesn’t have measles. Natural meaning in Grice’s terminology contrasts with nonnatural meaning. We can say that the three bells mean the bus is full, and that allows for error. There is no contradiction in saying the three rings mean that the bus is full but it isn’t. Perceptual illusions are cases of nonnatural meaning: They are cases of error.

It may be said that illusion can be glossed without representation: An antirepresentation advocate may say that an illusory perception is one in which the visual system gives its usual response to a certain feature in a situation in which the feature is not instantiated. But that antirepresentational gloss cannot explain the fact that the practice of vision science does not assume that perception is usually veridical.

It may be said that representation is reducible to a concoction of responsivity, functional role, and teleology. If so, that does not show there are no representations. Water is reducible to H2O, but no one will conclude that there is no water. Reduction as applied to water, heat, temperature, light, etc., is distinct from elimination as applied to caloric, phlogiston, and witches.

I will describe a controversy in vision science that does not assume that perception is usually veridical. It is particularly useful because neither the problem nor its solution can be understood without assuming that perception is representational.

The controversy involves “peripheral inflation,” a phenomenon to be discussed in more detail later in Chapter 13. Acuity, contrast sensitivity, and color sensitivity decline exponentially in the peripheral visual field, although most people do not report a decline in colorfulness or sharpness across the visual field. Galvin and colleagues (Galvin, O’Shea, Squire, & Govan, 1997) showed that when subjects matched the appearance of peripheral with foveally viewed edges, an blurrier edge in the periphery tended to be matched with a sharper edge in the fovea. Galvin et al. concluded that there is an illusion of sharpness in the peripheral visual field and there have been similar claims for colorfulness.

It should be obvious that we cannot understand the idea of ubiquitous illusions of sharpness and colorfulness in the peripheral visual field without assuming that perception is representational.

But are there really such illusions? Galvin et al. were implicitly adopting foveal appearance as veridical and anything that deviates from it as illusory. But neither foveal nor peripheral vision should be regarded as “the” standard of veridical perception. A perceptual ability cannot tell us about properties to which it is not sensitive (Anstis, 1998Haun, 2021). In foveal vision, we can’t see spatial frequencies above 50 cycles per degree. (See the caption and text surrounding Figure 1.1 for an explanation of spatial frequency.) In peripheral vision, our sensitivity is lower, but that is not a defect any more than it is a defect of foveal vision to not be sensitive to 70 cycles per degree. Nor is it a defect of color vision that it is not sensitive to ultraviolet or infrared light.

Colorfulness and sharpness are perceptual qualities that arise in different portions of the visual field in a way that is dependent on what kinds of sensitivities exist in that part of the visual field. An edge that would look blurry in foveal vision would look sharp in peripheral vision if the spatial frequency sensitivity required to see the blur was above the sensitivity of that part of peripheral vision. Then that edge would be matched in subjective sharpness to a sharper edge in foveal vision, explaining the Galvin et al. result.

Blurriness and sharpness have to be understood as relative to grain of representation. And that fact shows that blurriness and sharpness cannot adequately be discussed in a framework that eschews representation.

We have to distinguish physical properties like reflectance spectra from subjective properties like perceived colorfulness and sharpness (Haun, 2021). Subjective representation of sharpness and colorfulness depends on what the perceptual channel is sensitive to, and there are differences in that regard between peripheral and foveal vision.

This point does not make vision always veridical. Visual representation of a colorless display as colorful would of course be illusory. Representation of a sharp edge would be illusory if there is no edge. The point rather is that the veridicality of a representation of an edge in the periphery depends only on the low spatial frequencies that can be detected in the periphery and not on the high-spatial frequencies that can be detected in the fovea.

In sum, at least one aspect of the phenomenon of so-called peripheral inflation cannot be understood without appeal to visual representation.

Still another representational debate concerns the “tilted coin.” Do we see the tilted circular coin as circular, as elliptical, or as both? It is hard to get a handle on this debate except via consideration of what is represented in vision. The controversy concerns the question of whether we have two representations of the shape of the tilted coin, of both the circular and elliptical shapes, or whether we have only one, of the circular shape. The naïve realist would have to say that the controversy concerns whether we are directly aware of both the circular and elliptical shapes. But that construal is inadequate to understanding the experimental reasoning.

For example, Morales, Bax, and Firestone (2020) did experiments that suggest that a representation of a tilted circular coin interferes with a representation of an elliptically shaped coin seen head-on, the upshot being that the representation of a tilted circular coin also involves a representation of its circular shape. The explanation here concerns interference of representations and is not adequately construed in terms of direct awareness of shapes.

I don’t mean to endorse the reasoning of Morales, et al. The fact that children find it difficult to master the ability to report perspectival shapes and that even professional artists take longer to report perspectival than objective shapes suggests that the perspectival shape may be only postperceptually represented [Perdreau & Cavanagh, 2011].) (See Burge and Burge, 2022, for a reply to Morales, et al. and the reply by Morales, et al. in the same issue.)

Another type of finding that supplies direct evidence of visual representation comes from single case deficit studies. Michael McCloskey and his colleagues examined a subject who made bizarre errors of mislocation of visual targets. By analyzing her errors, McCloskey was able to show that she separately represented distance and direction from an “origin” whose location was dictated by the location of spatial attention (2009). She was much more accurate on distance from the origin than on direction from the origin, suggesting separate representations of these quantities. The reasoning involved here and its support of representational realism is very nicely spelled out in Chapter 2 of Karen Neander’s (2017) book. Her treatment explains in detail why the subject’s perceptual representations are intensional.

It may be said by those who reject realism that the practice of the science of perception embeds pragmatic decisions that have to be accepted by anyone who wants to join the field so that what the field takes as representation is not objectively representation. But what the field accepts is the task of understanding perception, what it is, how it works, and how it fits into the rest of the mind. That should be enough for real representation.

Three-layer methodology

Here is the methodology used in this book for assessing whether there is a joint in nature between perception and cognition and what constitutes it if it exists. Recall that minimal immediate direct perceptual judgment is the kind of cognitive state that is hardest to distinguish from perception, so my focus will be on distinguishing perception from that kind of perceptual judgment.

1.

Use prescientific ways of thinking of the perception/cognition border to make a preliminary classification of representations, states, and processes as definitely perceptual, as definitely cognitive, and as not definitely either. As explained below, one armchair approach appeals to the idea that perception has the function of being stimulus-dependent in a way that cognition is not; others focus on immediate warrant, and others focus on observable properties.

2.

Look for scientific indicators that make sense of the pretheoretic classifications while being aware that the scientific indicators may not always agree with the prescientific classifications. The main scientific indicator to be used below is perceptual adaptation, but also rivalry, pop-out, illusory contours, and speed of processing. Many other scientific indicators could have been chosen. Each of these indicators is dependent on the others, in a benign circular dependence I will argue that the scientific indicators give a better picture of what is perceptual and what cognitive than the armchair methods.

3.

Consider whether the scientific indicators of perception/cognition are constitutive of perception/cognition or rather symptoms of constitutive features. I will suggest that the use of adaptation, rivalry, pop-out, illusory contours, and speed of processing have more to do with the functions of perception and cognition rather than their underlying natures.

4.

Using the scientific indicators as indexes of what states and representations are perceptual and which cognitive, try to isolate the underlying constitutive features.

The prescientific indicators will be discussed in this chapter, the scientific indicators in Chapter 2, and the constitutive features in Chapters 456, and 7. Of course, this process can derail at any stage and the possibility that there are no explanatorily important differences to be found has to be kept in mind.

I am going to start in Chapter 2 with a first pass at the scientific indicators of perception and cognition and the problem of circularity in so doing. Focusing on perception, the problem of circularity is that the rationale for any indicator as an indicator of perception rather than cognition depends on the verdicts of other indicators. I will argue that the circularity is benign so long as there is a set of indicators that converge on the cases we are most sure of, classifying some cases as perceptual, others as cognitive, and none as both perceptual and cognitive.

The picture that I am assuming is that reality has an objective structure and one of the roles of science is to lay that structure bare. We have a pretheoretic grip on a distinction between perception and cognition, but the methods of Chapter 2 help us to refine those categories, putting some borderline cases on one side or the other and none on both sides. Once we have refined these categories, we can examine whether they have deeper natures and, if so, what they are. That is the topic of Chapters 48. Of course, the terms of refinement that I will be using—“nonpropositional,” “nonconceptual,” “iconic”—will no doubt themselves have borderline cases. I do not propose to characterize the border between perception and cognition in a way that eliminates all borderline cases.

Higher “capacity” in perception (whether conscious or not) than cognition

As I mentioned, the program of this book is to start with a discussion of the nature of perception and how it differs from cognition, putting aside issues of consciousness and phenomenology until the penultimate chapter of the book, Chapter 13. One of the advantages of this approach is that I can adapt arguments first used—ineffectively— with respect to consciousness and apply them with much greater effect as applied to perception. In this section I will describe one such approach.

Fragile visual short-term memory

Victor Lamme’s laboratory at the University of Amsterdam demonstrated fragile visual short-term memory in a series of articles starting with Landman, Spekreijse, and Lamme (2003). In many of these experiments, the subject is shown briefly a circle of rectangles that can either be vertical or horizontal. There is a dot in the middle of the screen which the subject is supposed to fixate in the sense of pointing the eyes at it. The array is replaced by a blank screen for a variable period. Then another array appears in which one of the rectangles may or may not have changed orientation. A line pointing to one of the locations—the cue— can appear at any one of three times; the subject’s task is to say if the rectangle that the line points to changes orientation between the first array and the second array. In Figure 1.3, the rightmost box of the top sequence shows the cue at the last stage, the leftmost box of the middle sequence shows the cue at the beginning, and the third shows the crucial case in which the cue comes in the blank period after the first display has disappeared but the last display has not yet appeared.

 The paradigm of the Amsterdam group led by Victor Lamme. A circular array of eight rectangles is presented for 1 second. There is a blank period of up to 1.5 seconds, followed by a second array of eight rectangles. The second array may or may not contain a rectangle at a different orientation from the first array. One rectangle is cued by a line either in the last array, in the first array, or in the blank period in the middle. There is a free pdf on the Oxford University Press web site that has the color version of this and all the other figures. Thanks to Victor Lamme for this figure.

Figure 1.3

 The paradigm of the Amsterdam group led by Victor Lamme. A circular array of eight rectangles is presented for 1 second. There is a blank period of up to 1.5 seconds, followed by a second array of eight rectangles. The second array may or may not contain a rectangle at a different orientation from the first array. One rectangle is cued by a line either in the last array, in the first array, or in the blank period in the middle. There is a free pdf on the Oxford University Press web site that has the color version of this and all the other figures. Thanks to Victor Lamme for this figure.

Open in new tabDownload slide

Using statistical procedures that correct for guessing, Landman et al. (2003) computed a standard capacity measure showing how many rectangles the subject is able to track. When the cue comes at the beginning, subjects unsurprisingly are close to perfect. In the bottom part of Figure 1.3 there is a graph with three bars: The middle bar indicates that subjects have a capacity of nearly all eight of the rectangle orientations when the cue is at the beginning. (Similar results are obtained if the cue appears within 10 ms from the offset of the first array.) When the cue comes at the end, with the second array, subjects show the classic working memory capacity of roughly four items. (See the next subsection below and Chapters 5 and 6 on working memory.) Lamme and his colleagues argue that the second array has obliterated the ongoing perceptual representation of the first array. Thus, the subjects are able to deploy working memory so as to access only half of the rectangles despite the fact that subjects reported seeing all or almost all of the rectangles. This is a classic “change blindness” result.

The crucial manipulation is when the indicator comes on during the blank period after the original rectangles have gone off but before the new array has appeared. If the subjects are continuing to maintain a visual representation of the whole array and reading their answers off of it—as subjects say they are doing, the capacity measure should be higher than four items. The finding is that the capacity is between six and seven for up to about 4 seconds after the first stimulus has been turned off, suggesting that subjects are able to maintain a visual representation of all or most of the rectangles. This result backs up what the subjects say. Note that the rightmost and middle bars of the graph at the bottom of Figure 1.3 are nearly the same. Similar results using different types of stimuli were obtained by Jeremy Freeman and Denis Pelli at NYU (Freeman & Pelli, 2007).

Lamme has argued that a conscious memory image of the first array persists in the blank period before it is wiped out by the appearance of the second array. According to that point of view, what is both conscious and accessible during the blank period is the impression of a circle of rectangles with their tilts specified clearly enough to distinguish vertical from horizontal. In the blank period the orientations and shapes of the specific items are conscious but there is a limitation on access: Necessarily, not all are accessed. None are inaccessible. Necessarily, most lottery tickets lose but there are no tickets that cannot win. Similarly, necessarily, many specific shapes are not accessed though none are inaccessible. The upshot is “overflow”: The capacity of conscious perception overflows cognition.

Lamme’s argument has been rejected by many critics on the ground that he has not provided any direct evidence that the fragile visual short-term memory representation is actually conscious (Byrne, Hilbert, & Siegel, 2007Cohen & Dennett, 2011de Gardelle, Sackur, & Kouider, 2009Grush, 2007Kouider, de Gardelle, & Dupoux, 2007Kouider, de Gardelle, Sackur, & Dupoux, 2010; Phillips, 2011a2011bVan Gulick, 2007). The critics say that what is conscious is that there is a circle of rectangles, but without the specification of the tilts. The idea is that the tilts are dredged up from unconscious memory when there is a cue. (See Sergent, et al., 2013.)

Advocates of Lamme’s argument (Block, 2007a20082011c) have argued that unconscious memory is too weak to support the high capacity revealed in the experiments by Lamme and colleagues. The critics’ retort is that experiments on unconscious short-term memory require very degraded stimuli in order to make the perceptions unconscious, whereas the stimuli used by Lamme and colleagues are not degraded. The reply to that is that unconscious perception in normal subjects may require such degraded stimuli. Still, I think it is fair to say that this debate is at something of an impasse.

However, the use I am making of Lamme’s results in this chapter is not vulnerable to this criticism. My point is that what the work by Lamme and his colleagues does show is that perception has a higher capacity than cognition. This shows that perception is fundamentally different from cognition, independently of issues of consciousness. This point is another illustration of the utility of my methodology of discussing perception independently of consciousness. In that way we can divide the question of greater capacity in conscious perception than in cognition into two questions: Is there a greater capacity in perception than cognition? Here the answer is clearly yes. And that answer strongly supports the thesis of this book that there is a joint in nature between perception and cognition.6 The second question is, Is the excess capacity conscious? I will take up that question in Chapter 13.

A further point is that the limit of 3–4 items in working memory are, as Jake Quilty-Dunn noted, an “item effect,” in which each item is encoded by a distinct vehicle, requiring more resources to represent more items (Quilty-Dunn, 2019a). Jerry Fodor, who coined the term “item effect” puts the point this way: “It is a rule of thumb that, all else being equal, the ‘psychological complexity’ of a discursive representation (for example, the amount of memory it takes to store it or to process it) is a function of the number of individuals whose properties it independently specifies. I shall call this the ‘item effect’ ” (Fodor, 2007, p. 111). Thus the 3–4 item limit in working memory suggests the discrete constituents typical of language of thought models.

There is one problem that has been raised concerning Lamme’s argument that also applies to mine, and I will discuss that argument in the next subsection.

Slot vs. pool models

These experiments have been criticized for presupposing a “slot” model of working memory in which working memory has a limited capacity—roughly four items for many of the standard stimuli used in experiments like those described above (Gross, 20172018Gross & Flombaum, 2017). The calculation that allows capacities to be computed assumes limited capacity of working memory.

It is true that slots in working memory may not be part of the most basic level of explanation. At the most basic level, working memory allows for many representations at reduced levels of precision. But in certain conditions, slot-like behavior can emerge. When unfamiliar items that don’t fit into a smallish set of categories are used, subjects do not show any clear limit in working memory. Instead, they show decreasing memory precision for larger sets. For example, the pictures (e.g., boat on a beach, narrow street, children holding hands) used in an experiment by Potter and her colleagues (2014) to be described in Chapter 8 show decreased memory precision for 12 items compared to 6 items but there is no sign in her data of any capacity limit. Fougnie, Cormiea, Kanabar, and Alvarez (2016) have shown that if given incentives to remember more items, subjects remember more items at decreased precision. By contrast, familiar closed class items that are easy to discriminate from one another, like digits, letters of the alphabet, and rectangles that can take a small set of cardinal orientations show working memory limits of up to four items.

Whether or not the representations of perception and working memory are probabilistic, there are notable types of slot-like behavior in working memory experiments (Adam & Serences, 2019Adam, Vogel, & Awh, 2017Bouchacourt & Buschman, 2019Donkin, Kary, Tahir, & Taylor, 2016; Pratte, Park, Rademaker, & Tong; Xie & Zhang, 2017Xu, Adam, Fang, & Vogel, 2018). Susan Carey has explored extensive slot-like working memory systems—the “parallel individuation” system—in infants who tend to have three rather than four slots (Carey, 2009).

One explanation of slot-like behavior has to do with the role of inhibition in suppressing less probable representations. Endress and Potter (2015; Endress & Siddique, 2016; Endress & Szabo, 2017) explain slot-like behavior in terms of an underlying variable precision model in which interference from long-term memory representations imposes slot-like processing. This slot-like behavior is often neglected in current controversies, for example by Gross and Flombaum (2017).

Hilary Putnam famously noted (1974) that a square peg of slightly under 1 inch on a side will fit through a 1-inch square hole in a board but not a 1-inch diameter round hole. The explanation of why the peg will fit through one hole but not the other need appeal only to the rigidity of the peg and the board and the geometry of the situation. Descending further to the level of the clouds of particles that make up the peg and the board will yield an explanation that is more general in one respect, since it can explain the cases in which the peg or the board deforms. But the explanation in terms of geometry and rigidity is more powerful in that it gives a simple explanation of why the square peg won’t fit through the round hole, an explanation that would only be obscured by describing the peg and board as clouds of elementary particles. The geometry and rigidity explanation is more general in a way since it applies to rigid pegs and boards made of cellular substances like wood, lattice structures like ice or diamond and amorphous structures like glass.

Explanation in terms of “slots” in working memory is like explanation in terms of rigidity and geometry. The slots in working memory are as real as the rigidity of the board and pegs. The pool of resources model is more general in a respect, since it can explain both the cases in which there is slot-like behavior and those in which there is not. But the slot model gives a simple and elegant explanation of the special cases in which slot-like behavior emerges.

Although the slot vs. pool debate has been ongoing for some time, the pool model is often simply ignored in articles comparing different forms of memory, even in the best journals. For example, a recent article in Current Biology on iconic memory says, “The most stable form of short-term visual memory is working memory. Working memory is resistant to masking . . . , but it has a limited capacity of only a few items” (Teeuwen, Wacongne, Schnabel, Self, & Roelfsema, 2021).

Gross and his colleagues (Gross, 20172018Gross & Flombaum, 2017) see the difference between iconic memory, fragile short-term memory and working memory not in terms of a difference in capacity but in terms of how “flat” the probability representations of the items are. They think iconic memory and fragile memory have relatively flat curves, representing many things at decreased precision, whereas working memory tends to represent a few things at high precision and many other things at lower precision. This picture ignores the role of suppression in slot-like behavior as described in the Endress and Potter model described above. Gross et al. see these kinds of memory as basically the same, whereas because of the role of suppression, there are qualitative differences.7

The most fundamental point about the difference between working memory and iconic/fragile memory, however, is that they are different in format. As will be argued in detail in Chapter 5, iconic and fragile memory have the format of perception, namely iconic format. Working memory, by contrast, is discursive.

A recent experiment provides fairly direct evidence for a difference in format between iconic memory and working memory. Michael Pratte (2018) presented subjects with 10 colored squares, and after a retention interval that could be as short as 33 ms or as long as 1 s, one location was cued. Subjects were asked to indicate the color of the square at that location by moving a cursor to the appropriate point on a smoothly varying color wheel. As Pratte notes, many previous experiments used stimuli that do not allow for measurement of the precision of what is retained. But, following (Zhang & Luck, 2008) Pratte used color stimuli and color wheel responses that do allow for measurement of the precision of memory. See Figure 1.4.

 In Pratte’s experiment, subjects were presented with 10 colored squares, then just a fixation point for a period between 33 ms and 1 second, then a cue indicating one of the 10 locations. They were tasked with identifying the color at that location by moving a cursor to the appropriate part of a color circle. Then they got feedback, showing what they had picked and what the correct color was. There is a free pdf on the Oxford University Press that has the color version of this and all the other figures. Thanks to Michael Pratte for this figure.

Figure 1.4

 In Pratte’s experiment, subjects were presented with 10 colored squares, then just a fixation point for a period between 33 ms and 1 second, then a cue indicating one of the 10 locations. They were tasked with identifying the color at that location by moving a cursor to the appropriate part of a color circle. Then they got feedback, showing what they had picked and what the correct color was. There is a free pdf on the Oxford University Press that has the color version of this and all the other figures. Thanks to Michael Pratte for this figure.

Open in new tabDownload slide

Pratte was interested in comparing a “sudden death” model of decay of the icon with a gradual decay model. He found that the sudden death model predicted the results much better than the gradual decay model. As time went on, subjects remembered fewer of the colors, but the ones they did remember were remembered with as much precision at the end of the delay period as at the beginning. Over time, the guessing rate increased as the memory capacity decreased, but the precision stayed about the same. See Figure 1.5. A similar result was found in Experiment 2 of Pratte (2019), except in this experiment, subjects were given the color and gave a graded response as to the location that color had in the first display. Again, the decay of memory capacity was due to an increase in guessing, rather than a change in precision.

 Two models of decay, gradual decay in the top row and sudden death in the bottom row. Items d and f have been omitted from this diagram. There is a free pdf on the Oxford University Press that has the color version of this and all the other figures. Thanks to Michael Pratte for this figure.

Figure 1.5

 Two models of decay, gradual decay in the top row and sudden death in the bottom row. Items d and f have been omitted from this diagram. There is a free pdf on the Oxford University Press that has the color version of this and all the other figures. Thanks to Michael Pratte for this figure.

Open in new tabDownload slide

The upshot is that iconic memory decays differently than working memory. Iconic memory decay is not compatible with the “pool” model, in which decay is loss of precision, suggesting that iconic memory and working memory have different formats, given that the pool models do appear to fit working memory.8

Armchair approaches to the perception/cognition border

It is commonly said that perception is more dependent on stimuli than are thoughts, beliefs, and judgments (Beck, 2014Camp, 2009Phillips, 2019Prinz, 2002Shea, 2014). However, the cognitive state that is most directly dependent on perception, minimal immediate direct perceptual judgment, is also dependent on stimuli. Perceptions are causally sustained by current proximal stimulation, but so are perceptual judgments. How can we distinguish perception from this kind of perceptual judgment?

Jake Beck (2018) says the key is that all representational elements of a perception have the function of being stimulus-dependent, whereas in a perceptual judgment, that is not necessary. I can see the spots on a bird as it flies by but still formulate the perceptual judgment that it has spots even when it is a dot on the horizon with no visible spots. And that judgment is fulfilling its function.

Recall that a minimal perceptual judgment conceptualizes each representational aspect of a perception and no more. An immediate perceptual judgment conceptualizes a perception with no inferential step. A direct perceptual judgment is based on the simultaneous perception with no intermediary. On the face of it, minimal immediate direct perceptual judgment would seem to be as stimulus-dependent as the perception it conceptualizes. Beck formulates an amended criterion that dictates that if the concepts in perceptual judgment can be applied outside of activation of transducers, they are not individually stimulus dependent and the judgment isn’t stimulus dependent. However, as Jake Quilty-Dunn notes (2020), that rules out conceptual perception by fiat. As we will see in Chapter 6, conceptual perception is an empirical issue and there are some empirical reasons to take it seriously—although ultimately, as I will argue, the best empirical case goes against conceptual perception.

It is often said that perception is fast, automatic and noneffortful and cognition has none of these properties. It is true that there are many cognitive states that are slow, effortful and require a decision and this can often be explained by the fact that cognition often requires global broadcasting, and that its formation takes at least 270 ms. But the minimal immediate direct perceptual judgments I have been talking about are plausibly cognitive and are automatic, seemingly noneffortful and perhaps faster than much of cognition. Further there are some perceptions that are slow, decisional, and effortful, for example perceptions that are the result of “free fusing” two images to make a single stereo image. (For instructions on how to do this, see the Wikipedia article on Stereoscopy or http://www.starosta.com/3dshowcase/ihelp.html.)

Perhaps the right armchair approach is via the epistemological role of perception. Perceptions are often regarded as justifiers that do not themselves require justification. Perceptions provide immediate prima facie warrant for de re beliefs about particulars where immediate warrant is warrant that is not mediated by warrant for something else. (Thanks to Ram Neta for conversation on this topic.) What is a de re belief? Oedipus has a de re belief, of his mother, that he is married to her. That is, he has a belief concerning a certain person (who—unknown to him— happens to be his mother) that he is married to her. The contrast is with a de dicto belief: Oedipus does not believe (de dicto) that he is married to his own mother. For example, he does not have a belief that he could express as “I am married to my mother.” The epistemic view may be right, but it won’t help much in deciding which representations are perceptions since in any real case in which there is a doubt about whether a state is a perception or a perceptual judgment, there will be a corresponding doubt about warrant.

Here is an example of a real dispute for which the epistemological approach does nothing to help us. Experimental results purport to demonstrate that desirable objects appear nearer (Balcetis & Dunning, 2010). After eating salty food, subjects’ judgments about the distance between them and a bottle of water were lower than after drinking water. Did they really perceive the distance differently or was the difference only in a postperceptual cognitive judgment?

Frank Durgin and his colleagues (Durgin, DeWald, Lechich, Li, & Ontiveros, 2011) provide evidence that these results really concern perceptual judgment rather than perception. If desirable objects really do appear nearer, then those distance perceptions provide immediate prima facie warrant for belief in the shorter distance; but if desirability affects perceptual judgment without affecting perception, then the distance perceptions do not. No appeal to immediate prima facie warrant will help adjudicate between Balcetis and Dunning (2010) and Durgin et al. (2011), since we can only decide the warrant question by deciding the perception question.

Another approach to distinguishing perception from cognition would be to appeal to the phenomenology of perception (Montague, 2018) as compared with the phenomenology of cognition—or the phenomenal properties ascribed by perception as compared with the phenomenal properties ascribed by cognition (Glüer, 2009; Kriegel, 2019Nes, Sundberg, & Watzl, 2021).

Perception could be said to be particularly phenomenally fine-grained or rich or to ascribe fine-grained or rich phenomenal properties. However, minimal immediate direct perceptual judgments may share the fine-grainedness of the perceptions they are based on. Further, it is not of the essence of the phenomenology of perception to be fine-grained. Larry Weiskrantz noted that blindsight patient DB had greater acuity in portions of his blind field in some circumstances than in the patient’s sighted field (Weiskrantz, 1986/2009 ). (Here I assume, controversially, that blindsight is truly blind.)

It may be said that there are properties that can be phenomenally represented in thought but not in perception. But that observation won’t help us with distinguishing a perceptual and a cognitive representation of the same property (Kriegel, 2019).

I am not going to pursue the phenomenological issue further, in part because I am looking for a characterization of the natural kind common to conscious and unconscious perception.

Some have argued that there is no one privileged way of delineating a border between perception and cognition (Beck, 2014Phillips, 2019). I will be arguing for a way of drawing the border that has fundamental explanatory significance. It is always open to others to try to find some other way of drawing the border that has equal explanatory significance.

Conceptual engineering

Joints are discovered, not stipulated. Fundamental structural divides are not a matter of convention. However, there is a role for what might be called conceptual engineering in the discovery of joints. What I have in mind is that to the extent that there is vagueness in the concepts of perception and cognition, they should be understood so as to home in on a joint if there is one. I’ll discuss four cases, perceptual learning, the superimposition of imagery on perception, the use of perceptual materials in cognition, and dual component views. (See Cappelen, 2018.)

Perceptual learning

There are many types of perceptual learning ranging from associative learning to more sophisticated forms of Bayesian updating. It is commonly said that perceptual systems are shaped by perceptual experience. For example, chess players often say that they can see patterns of strength and weakness. What is controversial is whether there is a direct effect of cognition on the shaping of perceptual systems or whether the shaping of perceptual systems is a matter of exposure, subsequent familiarity, and sensorimotor training.

The reason this issue is controversial is that it is difficult to separate the effects of training and exposure from the effects of conceptualization of that exposure. As Ellen Fridland (2014, p. 4) puts it, “In fact, proponents of cognitive penetrability often appeal to cases of perceptual learning and expertise in order to support their position. It is in cases of, e.g., expert radiologists, chess players, chicken sexers, artists, musicians, and athletes that changes in perception seem plausibly to occur. But in these cases, the change in perception results from regular, long-term exposure to and training with a certain class of perceptual stimuli.”

Perceptual categories can be learned, even in an hour of training. Ester, Sprague, and Serences (2020) trained subjects in categorizing tilted lines as on one side or another of a standard orientation (chosen arbitrarily for each subject; see Figure 1.6). Subjects became adept at this categorization and then two forms of brain scanning showed that representations in early vision of the orientations were repelled by the boundaries. That is, orientations near the boundaries were represented as comfortably on one side or the other. These categorical biases emerged at the earliest stages of visual processing.

 Tilt categorization task from (Ester et al., 2020). An arbitrary tilt was selected for each subject, indicated by the category boundary in the figure. Orientations on the clockwise side of the boundary were classified as category 2 and categories on the other side were 1. There is a free pdf on the Oxford University Press web site that has the color version of this and all the other figures. This figure is from the Journal of Neuroscience, which does not require permission.

Figure 1.6

 Tilt categorization task from (Ester et al., 2020). An arbitrary tilt was selected for each subject, indicated by the category boundary in the figure. Orientations on the clockwise side of the boundary were classified as category 2 and categories on the other side were 1. There is a free pdf on the Oxford University Press web site that has the color version of this and all the other figures. This figure is from the Journal of Neuroscience, which does not require permission.

Open in new tabDownload slide

Note that there is no evidence in this experiment for a cognitive effect on perceptual categorization. The mechanism by which these perceptual categories emerge may be entirely sensorimotor.9

I know of only one published study that makes a serious attempt to separate the effects of exposure from a direct cognitive effect (Emberson, 2016Emberson & Amso, 2012). They argue for a cognitive effect that is distinct from exposure, but in my view the case is underwhelming.

Some argue from categorical perception to cognitive penetration of perception (Emberson, 2016Gerbino & Fantoni, 2017). In categorical perception, discrimination is faster and more accurate when stimuli are on different sides of a category boundary: that is, better across than within categories. For example, children are born with equal discriminatory capacities for all the world’s languages, but as they learn their own languages, they lose sensitivity to differences within—and gain sensitivity across—phonological categories of the home language. It is well established that perceptual categories can be influenced by training. For example, training with new color categories can reshape color category boundaries (Özgen & Davies, 2002).

If categorical perception can be acquired in part on the basis of knowledge in addition to mere exposure, then there is “diachronic” (over time) cognitive penetration. (Categorical perception will be discussed in detail in Chapters 6 and 12.)

Pylyshyn (1999, p. 360) discusses some cases of perceptual learning, arguing that “none of these results is in conflict with the independence or impenetrability thesis as we have been developing it here.” Pylyshyn, following Fodor (1983), is taking the cognitive impenetrability thesis to hold only synchronically (at a time), not diachronically.10

This dispute can seem rather verbal, but there is an approach to a principled response: draw the borders between kinds that will reveal a joint if there is one. Suppose there are direct diachronic cognitive effects that change the structure of perceptual systems and, even more doubtfully, that those effects challenge a joint. Both assumptions are highly controversial (and I don’t subscribe to either), but let’s suppose they are true for the sake of the example. My proposal is that if excluding cognitively driven structural changes in perception homes in on a joint, then we should understand the concept of perception so as to exclude those diachronic effects (Block, 2016b). To be clear, I am not claiming that perception is only synchronic. Perception takes time and there are diachronic perceptual effects such as adaptation. Perceptual tracking is constitutively temporal. What I am suggesting might be excluded (on the basis of highly controversial premises) from perception is cognitively driven diachronic structural effects on perception if there are any.

One might call this a clarification of the concept, but not a clarification in a sense that counts as changing the concept, any more than realizing that whales are not fish was part of a change in the concept of fish. Many types of concepts contain within them a pressure toward natural kinds. By age 5, children understand that if you paint a stripe on a racoon and implant stink glands the result is still a racoon and not a skunk (despite being shown a “before” picture depicting a racoon and an “after” picture that looks like a skunk) (Gelman, 2003Keil, 1989). And they do not reason in the same way for artifactual kinds. If shown a “before” picture depicting a coffee pot and an “after” figure depicting a bird feeder, 5-year-olds tend to regard the operation as having changed a coffee pot into a bird feeder. Five-year-olds know that something can look more like coal than gold but be gold nonetheless. Even if they are unsure of how to classify something, they judge that there is a correct answer that an expert would know. And they take internal constitution as a better inductive base than external properties. These results have been replicated cross-culturally including with Native Americans (the Menominee), the Vezo in Madagascar, Yucatec Mayans, Yoruba in Nigeria, and the Torguud of Mongolia. (For a more skeptical spin on this issue, see Leslie, 2013.)

Just as realizing that whales are not fish clarifies what a fish is without changing the antecedent concept of a fish into a different concept, so the decision to exclude perceptual learning from perception clarifies what perception is without changing the concept of perception. We could call it clarification of the concept of perception but understanding “clarification” so as not to require conceptual change. The scientific concept of fish, excluding as it does, marine mammals, better captures the natural kind intent of the concept than a description based on shape. The “fish-shape” concept doesn’t even help with generalizations about modes of swimming, since true fish swim differently than warm-blooded marine mammals. The fact that it is so natural to use the phrase “true fish” expresses that implicit commitment to a natural kind element in the concept. The hope of clarification of the concepts of perception and cognition to home in on the natural kinds is to achieve a success of the same sort as the natural kind concept of fish.

It is worth noting that this kind of conceptual clarification can result in giving up part of what had earlier seemed essential. People sometimes think of the property of water of sustaining life as part of some kind of a definition of “water.” However, water is a mixture of “light” and “heavy” water, where heavy water involves an isotope of the more familiar form of hydrogen. This isotope (deuterium) has a nucleus with a neutron and a proton instead of the more usual lone proton. Heavy water is poisonous. So, the property of being able to sustain life is not a necessary condition of being water.

Superimposition of imagery on perception

Another example: Mental imagery can be superimposed on perception. Is the resulting state a perception? An example will make this issue more concrete.

Brockmole, Wang, and Irwin (2002) used a “locate the missing dot” task in which the subject’s task is to move a cursor to a missing dot in a 5 by 5 array. A partial grid of 12 dots appears briefly and then disappears followed soon after by another partial grid of 12 different dots in the same location that stays on the screen until the response. If the time between the two stimuli is short enough, subjects can fuse the 2 partial grids and move the cursor to the missing dot, remembering nearly 100% of the dots on the first array. However, if the second array is delayed to 100–200 ms, subjects’ ability to remember the first array falls precipitously (from nearly 12 dots down to 5 dots). Brockmole et al. explored extended delays—up to 5 seconds before the second array appears. The amazing result is that if the second array of 12 dots comes more than 1.5 seconds after the first array has disappeared, subjects become very accurate on the remembered dots. Instructions encourage them to create a mental image of the first array and superimpose it on the array that remains on the screen. This experiment will be described in more detail later in the section on mental imagery (p. 339).

We can call the resulting superimposition of imagery on perception, a “quasi-perceptual state.” If quasi-perceptual states are counted as perception, would this be cognitive penetration of perception? And if it is cognitive penetration, would it challenge a joint? I don’t think cognitive penetration challenges a joint (see Chapter 9 for many cases of cognitive penetration that do not challenge a joint). But those who think it does have a reason to exclude the quasi-perceptual states that result from superimposition from the category of perception. There is also an “ordinary language” reason: We do not normally count perceptual imagery as perception, since imagery does not function to be stimulus-dependent.

I believe that many quasi-perceptual states of this sort have routinely been counted as perception. Effects of this sort are involved in many of the phenomena used to criticize a joint in nature between cognition and perception. Many of the effects of language and expectation on perception cited by opponents of a joint in nature seem very likely to involve superimposition of imagery on perception. For example, Gary Lupyan notes that in a visual task of identifying the direction of moving dots, performance suffers when the subject hears verbs that suggest the opposite direction (Lupyan, 2015). The plausibility that this result involves some sort of combination of imagery and perception is enhanced by the fact that hearing a story about motion can cause motion aftereffects (Dils & Boroditsky, 2010).

Another example: Delk and Fillenbaum (1965) showed that when subjects are presented with a heart shape and asked to adjust a background to match it, the background they choose is redder than if the shape is a circle or a square. As Fiona Macpherson (2012) points out, there is evidence that subjects are forming a mental image of a red heart and superimposing it on the cut-out. Macpherson regards this as a case of cognitive penetration. But whether it is cognitive penetration depends on whether the resulting quasi-perceptual state is a genuine perceptual state (see the discussion in the section on mental imagery in Chapter 9).

In a recent article in Nature Human Behaviour, Chunye Teng and Dwight Kravitz (2019) showed that holding a color or orientation in mind while doing a perceptual task biased the perceptual classification toward a distractor if the mental image was similar to the distractor. Again, whether this is cognitive penetration depends on a prior decision as to whether quasi-perception is perception.

There is also evidence that visual imagery is involved in ordinary perception where subjects are not asked to imagine anything. Tarr and Pinker (1989) taught subjects to recognize line drawings and then examined recognition of the objects depicted at unfamiliar angles. They showed that the time it took to recognize an object at an unfamiliar angle depended on the angular distance that would be required to rotate the object to the familiar view. This suggests that perceptual object representation can involve coordinated representations from different vantage points, at least some of which can be characterized as mental images.

Fodor’s and Pylyshyn’s notion of “cognitive penetration” requires a direct effect of cognition on perception in the sense of no intermediate causal link that is a person-level mental state. (See Chapter 9 for further discussion.) By that criterion, these imagery effects will count as cognitive penetration only if quasi-perception is perception.

Cognitive states that use perceptual materials

The issue of drawing borders to home in on a joint if there is one arose earlier in connection with the question of why I was counting perceptual simulations used in cognition (e.g., to determine whether the couch would fit through the doorway if rotated) as cognition. In such process, there are iconic, nonconceptual, and nonpropositional representational elements and these elements are deployed in reasoning. I said that because they are deployed in reasoning, these elements are enclosed in a cognitive envelope. I argued that what makes a perceptual representation perceptual is not just being iconic, nonpropositional, and nonconceptual but also whether those properties are constitutive of the state.

Earlier in this chapter, I mentioned three important differences between perceptual representations in perception and similar representations in working memory. First, perceptual simulations used in cognition may not use the fine-grained representations of true perception, at least with regard to color representations. I mentioned minimal shades of colors that may require being driven by bottom-up world-driven information flow. Second, there are computational differences between perceptual representation in perception and perceptual representation in working memory having to do with “divisive normalization,” a notion that will be explained in the next chapter. Finally, working memory representations do not have the phenomenology of perception.

Still, one should have an open mind about whether the best way of thinking about perceptual simulations is by treating them as part of the same natural kind as perception, that natural kind being what one might call perceptual representation. This would handle perceptual simulations and superimpositions of imagery on perception in a uniform manner. And if perceptual memories, perceptual anticipations, and mental maps constitutively involve iconic, nonconceptual, and nonpropositional representations, the same proposal may classify these representations as perceptual representations. (However, see the section below on whether map-like structures are actually conceptual, and indeed are the basis of thought.)

The constitutive iconic format and nonconceptual and nonpropositional nature of perceptual representation provide necessary conditions of perceptual representation but more is required to provide a sufficient condition. As mentioned earlier, to distinguish between perception and sensation, we may require that the perceptual representations involve the constancies of perception (Burge, 2010a).

Dual component views

Some theorists hold “dual component” theories of perception (Smith, 2002) in which a perceptual experience is a complex state that has a nonconceptual nonpropositional component and, at least sometimes, a conceptual/propositional component (Peacocke, 1992b). The conceptual/propositional component is usually supposed to be caused by the nonconceptual component. According to some, the conceptual/propositional component is a belief (Quilty-Dunn, 2015), according to others it is a “seeming,” a propositional attitude that is formed automatically and persists despite not being endorsed by the subject (Brogaard, 2014Reiland, 2014Tucker, 2010). Is there any substantive difference between such views and the one that I will be advocating?

Someone might argue that the disagreement is just a matter of how expansive we wish to be about what to include in perception, and in particular whether to include the propositional component in perception. I reject such views and not just on the ground of conceptual engineering that I have been talking about. In Chapter 4, I will be arguing that perceptual representations do not have the logical properties required of propositional representations.

If there is a fundamental difference between perception and cognition, why don’t we see the border in the brain?

Here is part of a recent interview by Jordana Cepelewicz of Lisa Feldman Barrett and other psychologists and neuroscientists (Cepelewicz, 2021b):

“Scientists for over 100 years have searched fruitlessly for brain boundaries between thinking, feeling, deciding, remembering, moving and other everyday experiences,” Barrett said. A host of recent neurological studies further confirm that these mental categories “are poor guides for understanding how brains are structured or how they work.” Barrett is giving voice to a widespread view that the real mental categories are the ones that neuroscientists discover.”

This is a view famously put forward by Paul and Patricia Churchland (P. Churchland, 1986P. M. Churchland, 1981).

I certainly agree with the Churchlands that the categories of folk psychology will be refined and in some cases eliminated by neuroscience, but importantly—and this is often left out—many of our folk categories can be validated by neuroscience, although perhaps not current neuroscience. Further, as Cepelewicz goes on to say, “However, often neuroscientists can only discover crucial mental categories once they are identified by psychology.” Indeed the categories neuroscientists validate are often categories that are provided by psychology with no help from folk psychology. The opponent process theory of color perception was discovered in the nineteenth century by Ewald Hering and further elaborated by Dorothea Jameson and Leo Hurvich in the 1950s, all on the basis of behavioral data and introspection. Then the theory was validated by finding opponent cells in the lateral geniculate nucleus and later refined using both neural and behavioral data.

The absence of a clean spatially specific border in the brain is illustrated by a study in which subjects were asked to make similarity judgments of many pairs of pictures (Bracci & Op de Beeck, 2016). The pictures fit into six categories: animals, vegetables/fruit, minerals, musical instruments, sports items, and tools. They also differed along nine shape dimensions. For example, some instruments, vegetables, and minerals were long and thin, others roughly circular. Early visual areas were dominated by shape-based similarity judgments, prefrontal areas were dominated by category-based judgments, but most of the brain showed a mix of responses to shape and category. Albert Newen recently argued that studies of this sort show that neuroscience dictates a third category in between perception and cognition and hence that there is no joint in nature between perception and cognition (Newen, 2021).

It would be natural to interpret both Newen and Barrett as saying that no one has discovered boundaries in the brain between perception and cognition so there is only a superficial reality to the difference.

But the problem with this argument is that there can be a basic difference that is not realized spatially. There is a basic difference between data and program representations in computers, and early computers stored data and program in separate registers, but modern computers often use distributed representations. In distributed representations, the separate parts are defined by functional relations rather than physical areas. Two data registers are linked not by adjacency but by pointers.

Further, a bottom-up analysis of the chips in a computer would not easily lead to an understanding of basic principles of organization. This point was famously made in an article in PLOS Computational Biology titled “Could a Neuroscientist Understand a Microprocessor?” applying standard neuroscience techniques to a primitive Atari chip (also used in the first Apple computer) (Jonas & Kording, 2017). The authors conclude (p. 14), “However, in the case of the processor, we know its function and structure and our results stayed well short of what we would call a satisfying understanding.” Of course, given that the difference between perception and cognition is real and fundamental, it will be in principle possible to find its neural implementation. But no one should expect that it will be easy to find using current techniques.

Interface of perception with cognition

The main interface of perception with cognition is when perceptual materials are retained in working memory, as described earlier in this chapter and in greater detail in Chapters 5 and 6. But there are two other cognitive systems that are not closely associated with working memory, to be described briefly in this section.

I’ll start with the system that is most unlike perception, the language of thought system, long discussed in the philosophical literature, including in the late medieval period by William of Ockham (Rescorla, 2019), and revived more recently by Jerry Fodor and Gilbert Harman (Fodor, 1975Harman, 1973). Perhaps the leading property of language of thought systems is the independence of syntactic roles and the lexical items that are fillers of those roles (Frankland & Greene, 2020a, 2020b). This kind of independence does not apply to iconic representations where a smudge that represents a hand in one part of one picture could represent a foot in another part, a claw in another part or a flipper in another part.

A somewhat different emphasis involves tree structures. Tecumseh Fitch proposed the “dendrophilia hypothesis that ‘humans love trees,’ ” more specifically, “that ‘humans have a multi-domain capacity and proclivity to infer tree structures from strings’ even in the simplest cases, to a degree that is difficult or impossible for most nonhuman animal species” (2014, p. 352).

Stanislas Dehaene and his colleagues have provided strong evidence for the first part, that humans love trees, and more specifically that humans have a proclivity to code sequences into recursive tree structures (Planton et al., 2021). What is meant here by “recursive”? If you concatenate an adjective (e.g., “big”) with a noun phrase (e.g., “green egg”) you get a new noun phrase (“big green egg”), and that new noun phrase can itself be concatenated with an adjective to form still another noun phrase (e.g., “expensive big green egg”). This is an example of one kind of recursion. Applied to procedures, a recursive procedure is one whose implementation requires that very procedure.

As Dehaene notes (2000; Planton et al., 2021), current neural networks cannot represent truths that involve recursion, such as “Every number has a successor.” Humans learn such rules easily. Four-year-old humans can learn to reverse the sequence ABCD to DCBA in five trials, but nonhuman primates take tens of thousands of trials. Thus, human cognition seems importantly different from the cognition of nonhuman primates and the computations of current neural networks.

George Miller famously proposed that working memory has 7, plus or minus 2, “slots.” (Recall as explained earlier in this chapter, slots in working memory are real but not fundamental.) Later work suggested that the evidence for Miller’s estimate did not control for “chunking.” For example, as will be noted in Chapter 2, the series of letters “FBI CIA KGB” is much better recalled than “KBA GFI BFC” (Rosenberg & Feigenson, 2012). Chunking is a data compression mechanism. As Planton et al. propose, the complexity of a sequence of stimuli can be indexed by the length of its compressed form using an internal language that allows for nesting. The internal language relevant to the binary stimuli involves a term for “stay” and a way of representing change. For example, AAAA would be represented as four stays, whereas ABAB would be an item plus three repetitions of a change.

Dehaene and his colleagues explored how subjects process sequences of stimuli. They gave subjects sequences of auditory and visual stimuli that were binary in the sense of being composed of two types of items that repeated in sequences. The items could be high and low tones or red and green dots, for example. Assuming that subjects coded the stimuli using two instructions, “same” and “change,” they tested subjects’ ability to detect sequence violations in five different experiments, finding that “data compression” coding schemes using recursive tree structures explained subjects’ behavior for all but the shortest sequences. The psychological complexity of the sequences were indexed by the size of the most compressed mental representations of them as shown by the subjects’ detections of violations of the sequence rules, but also by subjects’ complexity ratings.

One surprising prediction that was borne out concerned AnBn patterns. For sequences of 16 items, these sequences could be two chunks of 8 (i.e., AAAAAAAABBBBBBBB), four chunks of 4 (i.e., AAAABBBBAAAABBBB), eight chunks of 2, or 16 chunks of 1. These sequences all have the same model complexity (the log number of repetitions) and were found to have the same psychological complexity.

A similar model also predicted subjects’ behavior in a task involving spatial locations on a regular octagon (Amalric et al., 2017). The experimenters found that a language of nested sequences of geometrical primitives of rotation and symmetry explained subjects’ behavior in judgments of regularity, in completion of sequences, and in an eye-tracking task. Sequence complexity also predicted brain activity in the inferior frontal cortex. Results were similar among three very different groups, French adults who had been through school, French kindergarteners, and the Munduruku, indigenous Amazonians who had no schooling and very limited language concerning numbers and geometrical terms, suggesting that the “language of thought” as applied to geometry is a basic human ability that does not depend on culture.

Later experiments by the same group (Sablé-Meyer et al., 2021) used a methodology that required subjects to detect an “odd one out” among six quadrilaterals. This task is suitable both for nonlinguistic subjects and subjects with language. The behavior of unschooled members of the Himba tribe, French kindergarteners, and French adults were predicted by the language of thought model, but that model did not predict the behavior of baboons in this task. Baboon behavior was best modeled by a convolutional neural net model, not by the language of thought model. As the authors note, this result suggests that symbolic abstraction with nested structures is a basic human capacity that distinguishes humans from other primates.

Steven Frankland and Joshua Greene have identified different brain regions connected with different cognitive systems (Frankland & Greene, 2020a, 2020b). They take the language of thought system grounded in the dorsolateral prefrontal cortex to code instructions for thoughts, and “grid-like” representations in the default mode network (DMN) to serve as the “canvas” for these thoughts. The dorsolateral prefrontal cortex, the top/side of the prefrontal cortex, is widely thought to be the center of thought, the control of working memory, and executive function. The DMN involves the “inside” of the prefrontal cortex, where the hemispheres face each other; the posterior cingulate cortex, also on the midline of the brain; and the angular gyrus, on the border of the parietal and temporal lobes. The DMN was originally identified as a wakeful rest area implicated in mind-wandering, but Frankland and Greene review very different functions in mental maps involving both physical space and conceptual spaces.

The DMN is implicated in conceptual combination both in fMRI results and in lesion studies. For example, patients with low gray matter density in the angular gyri, parts of the DMN, show impairment in conceptual combination tasks but not in single word tasks. According to Frankland and Greene (2020a, p. 295), “Research on the timing of semantic processing indicates that the DMN is where semantic production begins and semantic comprehension ends.”

Grid cells, centered in the DMN, are involved in spatial navigation but also play a role in the use of spatial abilities in conceptual combination. Grid cells as used in spatial navigation represent space via representation of equilateral triangles combined to form hexagons, six triangles to a hexagon. Each of the six triangles assembled in the hexagon occupies one-sixth of the hexagon, spanning 60o. This 60o structure is revealed in greater grid-cell activity for spatial changes of 60o than other changes. Remarkably, this difference can be observed in fMRI recordings at various parts of the DMN during spatial navigation (Doeller, Barry, & Burgess, 2010Frankland & Greene, 2020a).

Constantinescu, O’Reilly, and Behrens (2016) used the Doeller procedure on tasks involving a two-dimensional space in which one axis was the length of a bird’s neck and the other was the length of the bird’s legs. They observed the same 60o signature. Frankland and Greene suggest that when conceptual representations involve magnitudes, vector displacement representations in the grid-cell network may be used in conceptual combination. These grid-cell representations are iconic in the sense to be introduced in chapter 6 of analog mirroring, in which relations among represented environmental properties are mirrored by instantiations of brain-analogs of those relations.

Frankland and Greene review a great deal of literature on the uses and limitations of the grid-cell network in cognition. What is relevant for the purposes of this book are first that the uses they review are conceptual, involving conceptual combination in thought. The grid cell system is not a perceptual system. Unlike “place cells” that remap according to perceptual input, grid cells are relatively insensitive to perceptual input. Second, given that grid-cells involve iconic representation it is natural to suppose that they interface with the iconic representations of perception. Third, as Frankland and Greene make clear, this work is in its infancy and much of what they say is framed in the language of speculation. As we will see, the speculative nature of this work contrasts with what we know about perception, as explored in later chapters of this book.

Why should philosophers be interested in this book?

Here are some reasons why philosophers should be interested in this book.

1.

The relevance to epistemology is significant, given that perception but not perceptually based cognition is often supposed to provide unjustified justifiers. If perception is conceptual, propositional, and discursive, then a compelling view of how perception justifies belief is just that we believe what we see—or hear, feel, etc. But if I am right that perception is none of those things, then a different model of perceptual justification is required. Traditionally, the philosophy of perception has been geared toward illuminating the epistemology of perceptual judgment—what justifies the judgments about the world that we base on perception (Stoljar, 2009). This project has often ignored the science of perception. The presumption of this book is that epistemologists would do well to find out what perception is from the science of perception and to base the epistemology of perception on that scientific answer.

2.

If perception is nonconceptual, nonpropositional, and iconic, then certain kinds of robots will not be perceivers. Specifically, if a camera output writes directly into cognitive representations, the robot would have data-driven cognitive states that are not perceptual.

3.

The conclusions of this book are relevant to issues concerning the synthetic a priori. I’ll give an example from a recent controversy.

4.

Most importantly, the conclusions of this book concern the nature of minds.

Paul Boghossian and Timothy Williamson have debated whether there are synthetic a priori truths that are justifiable by intuition, more specifically, justified by intellectual seemings (Boghossian & Williamson, 2020). The kind of intellectual seeming at issue would include, for example, the appreciation of the truth of the proposition that it is morally wrong to inflict pain merely for one’s own amusement. Boghossian argues that such intellectual seemings are similar to perceptual seemings in that they are “predoxastic” in the sense of prior to actual belief and also that these intellectual seemings dispose us to believe. He argues further that the considerations that show that perceptual seemings justify perceptual belief apply also to the claim that intellectual seemings justify intellectual beliefs. Williamson opposes predoxastic seemings in both cases. According to Williamson, we have a visual seeming that the Müller-Lyer lines are the same length, but it is not predoxastic because it is constitutively tied to the “felt visually-based inclination to judge that one line is longer” (pp. 232–233).

Similarly, according to Williamson, an inclination to judge that it is morally wrong to inflict pain for one’s own amusement does not present itself to him as based on its seeming true. If asked why he is inclined to judge that p, an appeal to p seeming true “sounds forced and feeble” (p. 233) because the explanans is too close to the explanandum.

In my view, both Boghossian and Williamson are mistaken. Boghossian is mistaken because intellectual seemings are not predoxastic and Williamson is mistaken because perceptual seemings are predoxastic. More specifically:

1.

Perception is plausibly nonconceptual and nonpropositional, and so are perceptual seemings, but intellectual seemings have to be conceptual and propositional since the contents cannot be appreciated without thought.

2.

Perception is iconic, while cognition, including intellectual seemings, is largely discursive, the exception being map-like representations.

3.

Perception is subject to large adaptation effects. If I look at a red square for more than a few seconds, it will look slightly less red, as the perception shifts toward the green end of the red/green opponent process channel. (See Chapter 2.)

Because of adaptation, perception of ambiguous stimuli results in rivalry, as explained in detail in Chapter 9. If I look at a Necker cube, one face will appear to come toward me. That perception will then weaken due to adaptation, and then the other way of perceiving it will win out and another face will come forward. This can continue indefinitely. An ambiguous figure/ground display yields comparable oscillations in how we see it because of adaptation. (Again, see Chapter 2.) But there is no comparable oscillation in intellectual seemings. People disagree as to whether XYZ is water or not, but we do not experience oscillating views of the sort we do with perception.

4.

Perception is to a large extent architecturally separate from cognition and so to a large extent functions autonomously of the subject’s theories. The main exception is for ambiguous stimuli. See Chapters 911. We cannot say the same however of intellectual seemings. They are part of the cognitive system and so not architecturally distinct from it. There is every reason to think that they are highly influenced by the subject’s theories. Cognitive penetration of perception is limited, but cognitive penetration of intellectual seemings is likely to be relatively unconstrained.

This last item is by far the most significant of the four points for the Boghossian/Williamson debate. The epistemic value of intellectual seemings is likely to be greatly reduced compared to the epistemic value of perceptual seemings. Susanna Siegel (2017) argues that the epistemic status of a perceptual seeming is affected by how it is formed. For example, wishful seeing or fearful seeing weaken the epistemic force of the perception. But a similar point applies to intellectual seemings that are influenced by one’s theoretical views. To allow intellectual seemings to support conclusions that play a role in producing the intellectual seeming in the first place would be a kind of “double counting” and so the intellectual seeming should be epistemically downgraded.

Roadmap

Chapter 2 is concerned with markers of the perceptual.

Chapter 3 is concerned with whether the content of perception is singular, whether perception is attributional, and whether there are two kinds of seeing-as. It ends with a brief discussion of racially biased perceptual responses by way of illustrating how we can distinguish between perception and perceptual judgment.

Chapters 47 make the positive case that perception is constitutively iconic, nonconceptual, and nonpropositional. Then Chapter 8 makes the negative case—that arguments to the contrary are mistaken.

Chapter 6 describes a special kind of perceptual representation, a perceptual category representation. These representations are often conflated with concepts—wrongly, I will argue. This is the central chapter for my argument that perception is nonconceptual and the basis for my new argument for “overflow”.

Chapter 7 discusses evidence from neuroscience that perception is nonconceptual.

Chapter 8 discusses evidence that is wrongly taken to show that perception is conceptual.

Chapter 9 describes fundamental machinery of perception that determines direct content-appropriate effects of the content of cognition on the content of perception—i.e., cognitive penetrations (by many common standards). The idea here is that once one sees what the joint between perception and cognition is, we can see that feature-based attention, imagery, and other ubiquitous phenomena involve cognitive penetration. Then I will observe that from what we can tell so far, the mechanisms of cognitive penetration (and the representations produced by these mechanisms) divide into the perceptual and the cognitive; so, there is no reason to believe that interpenetration of perception and cognition show any problem with the joint.

Chapter 10 discusses top-down effects that have been mistakenly supposed to be effects of cognition on perception.

Chapter 11 discusses modularity. I will argue against modularity in the sense of Fodor and Pylyshyn, but also that there is substantial truth in the modularity thesis.

Chapter 12 discusses core cognition, arguing against the view that representations of causation and numerosity form a third category intermediate between perception and cognition.

Chapter 13 discusses the consequences of the joint for cognitivist and conceptualist theories of consciousness.

This book presents a certain conception of perception, of cognition, and of the difference between them. I give evidence and argument for some but not all of the details. I am hoping that the plausibility and coherence of the picture presented will carry some of the burden of argument. I take it to be generally agreed that cognition is paradigmatically conceptual, propositional, and discursive (noniconic), though I will say a bit more in what follows in contrasting perception with cognition.

As the reader will see, I focus much more on perception than on cognition. The reason for that is that the psychology and neuroscience of perception is vastly better developed than the psychology and neuroscience of cognition. The perceptual systems all have fairly similar tasks—of making the output of sense organs useful to the organism. And phenomena discovered in one sensory modality often appear in others. For example, “change blindness,” first discovered in vision, also appears in auditory and haptic perception. By contrast, the aspects of the mind that use conceptual propositional discursive representations are a disparate lot with little uniformity. The best developed of the sciences of cognition are those, as with the psychology of language, that are most like perception.

Notes

1

There is, however, evidence that spatial resolution is not decreased in working memory (Tamber-Rosenau, Fintzi, & Marois, 2015). At the neural level, Zhao et al. used an orientation working memory task while the subjects were undergoing fMRI scanning. They found that the earliest cortical representations (in V1) on the opposite side from the stimulus lost precision over time, but oddly the representations on the same side as the stimulus did not lose precision. Thus it may be that there are two different kinds of spatial representations involved in spatial working memory (Zhao, Kay, Tian, & Ku, 2021).

2

 Keijzer (2013) argues that Wooldridge exaggerates. In the most systematic study he describes, using 31 wasps, 10 repeated until the end of the experiment, 10 broke the loop, and the others seem a grab-bag of cases. However, no single stereotyped action patterns should be expected to operate exceptionlessly. Perhaps different stereotyped action patterns are mixed, probabilistically. Keijzer also argues that in certain cases, fixed action patterns may be useful from an evolutionary point of view. However, stereotyped behavior that is evolutionarily selected is still stereotyped behavior.

3

Hallucinations and visual imaginings are iconic, nonconceptual, and nonpropositional without being perception (because of the absence of the normal causal relation to objects perceived). That normal causal relation could be included in an analysis of the ordinary concept of perception, but I do not regard it as part of the fundamental scientific nature of perception. See the section on conceptual engineering in this chapter. Another kind of failure of the sufficient condition has to do with the distinction between sensation and perception, the difference being that sensation lacks the objective import of perception, argued by Tyler Burge to involve perceptual constancies (Burge, 2010a). The representation-like states of sensation are iconic, nonconceptual, and nonpropositional without being perceptual. So, if perception is a natural kind, an additional requirement of constancies would have to be imposed.

Thought can perhaps use mental maps, perceptual memories, and perceptual anticipations without conceptualization of them (Burge, 2010aFridland, 2014), though see the discussion later in this chapter for evidence that maplike structures may be the foundation of conceptual thinking. In Chapter 5, there will be a discussion of iconic representations used in cognition, such as iconic representation of number. Perceptual simulations can be used in cognition though perhaps only by being conceptualized. Standard imagistic cognition tasks include deciding if the tip of a horse’s tail goes below the horse’s “knees,” whether there is a letter formed by rotating a capital ‘N’ 90 degrees (either clockwise or counterclockwise). One can use perceptual simulations in all sorts of cognitive tasks, for example, imagining the layout of one’s apartment in order to decide how many paint shades are needed.

One could perhaps fashion an ungainly necessary and sufficient characterization of perception in terms of being constitutively nonconceptual, nonpropositional, and iconic, while having objective import (unlike sensation) and involving actual objects in an appropriate causal relation to the perceptual state (unlike hallucination, mental maps, perceptual memories, perceptual anticipations, and perceptual simulations). I am offering this characterization only in a footnote instead of the text because I am not in the business of offering necessary and sufficient conditions. The conditions just sketched depend on a list of problematic kinds of cases that may not be complete.

4

I am indebted to conversation with Adam Pautz concerning this section.

5

As far as I know, this distinction was first introduced into the literature in John Haugeland’s comment on Searle’s “Minds, Brains and Programs.” In his response to Haugeland (the references in the text are to these two publications), Searle gives the distinction his own terminology, which I have used here. Haugeland used “original” and “derivative.”

6

That the overflow argument can be given in a form that does not mention consciousness was also noted by Peter Carruthers (Carruthers, 20152017).

7

Interested readers may want to consult the increasingly baroque controversy between advocates of models of working memory that are more partial to slot-like aspects and models that emphasize a pool of resources (Adam & Serences, 2019Adam, Vogel, & Awh, 2017Bays, 2018Brady, Konkle, & Alvarez, 2011Donkin et al., 2016Ma, Husain, & Bays, 2014Pratte, 2019Suchow, Fougnie, Brady, & Alvarez, 2014Xie & Zhang, 2017Z. Xu et al., 2018).

I argued in the previous section that perception (whether conscious or unconscious) has a higher capacity than cognition. There is a result however that may be thought to undermine that conclusion. Wu and Wolfe did an experiment involving multiple object tracking (Cohen, 2019Wu & Wolfe, 2018). The multiple object tracking paradigm is described in Chapter 4. In Wu and Wolfe’s experiment, a number of animal pictures move around the screen. (They used from 6 to 32 animals at a time.) At a randomly chosen time, the animals are replaced by gray disks and the subject is asked to move a cursor to a specified animal, e.g., the horse. Earlier experiments calculated a capacity to track of about 2.7 items, but Wu and Wolfe collected multiple guesses, reasoning that later guesses might reveal approximate knowledge of the locations. And that is what they found: approximate knowledge of the locations of up to 9.9 items. This result may seem to challenge my argument, because if the capacity of working memory is much larger than we had thought, then the capacity of perception may not be larger than the capacity of cognition.

However, there is a flaw in this objection. Slot-like behavior only emerges with closed-class items, such as letters or rectangles, that can have only a few orientations. This experiment involves continuous values of locations, so the slot reasoning does not get a grip.

8

This point is also made in (Quilty-Dunn, 2019a). As we will see in Chapter 5, Quilty-Dunn holds that object representations in perception and working memory are discursive, so if the representations of the colored squares in this experiment are object representations, Quilty-Dunn owes us an explanation of why two discursive representations have such different properties.

10

Fodor isn’t that clear about the matter, but in a discussion of the diachronic modification of associative connections, he says, “to put the matter somewhat metaphysically, the formation of interlexical connections buys the synchronic encapsulation of the language processor at the price of its cognitive penetrability across time” (Fodor, 1983, p. 82). What I am calling the diachronic/synchronic distinction is sometimes referred to as the off-line effect/on-line effect distinction (Lupyan, Rahman, Boroditsky, & Clark, 2020).

9

There are category repulsion effects in other perceptual paradigms that are partly perceptual and partly postperceptual. See Fritsche and de Lange (2019).

The Border Between Seeing and Thinking. Ned Block, Oxford University Press. © Oxford University Press 2023. DOI: 10.1093/oso/9780197622223.003.0001

This is an open access publication, available online and distributed under the terms of a Creative Commons Attribution-Non Commercial-No Derivatives 4.0 International licence (CC BY-NC-ND 4.0), a copy of which is available at https://creativecommons.org/licenses/by-nc-nd/4.0/. Subject to this license, all rights are reserved.

Metrics

View Metrics

Email alerts

New books

New reference articles

New articlesActivity related to this book

Sign up for marketing

Arrow

Arrow

Recommended

Powered by

More from Oxford Academic

Arts and Humanities

Philosophy

Philosophy of Mind

Books

Journals

Leave a comment