Visual Perception  Complete Notes

by Chiqui C. Patio

Audio version created with Paper2Audio.

Listen on Paper2Audio

Visual Perception  Complete Notes

Chiqui C. Patio
Audio by Paper2Audio.

Brain Trivia

• Your brain is a random thought generator.
- In 2005, the National Science Foundation published an article on research about human thoughts per day.
- The average person has about 12,000 to 60,000 thoughts per day. Of those, 95% are exactly the same repetitive thoughts as the day before, and about 80% are negative.
• Lecture add-ons:
- Humans have many thoughts every day, and a lot of them are repetitive — the same thoughts tend to resurface day after day.
- A large proportion of our daily thoughts are negative.
- Imagining scenarios, making "what if" situations, and creating stories in your head (e.g., imagining yourself as the lead in a K-drama you're watching) all add to the sheer number of thoughts we have per day.

Learning Outcomes

• We explore vision — humans' dominant sensory modality.
- We discuss the mechanisms through which the visual system detects patterns in incoming light, and also how it interprets and shapes that information.
- Perception of one aspect of input is shaped by perception of other aspects — detection of simple features depends on how the overall form is organized, and perception of size depends on the perceived distance of the target object.
- Interpretation of visual input is usually accurate — but the same mechanisms can lead to illusions, and studying those illusions illuminates how perception functions.
• Lecture add-ons:
- Visual perception feels fast, easy, and automatic — we just open our eyes and immediately see things — but the actual process is complex.
- Our ability to perceive the world depends on many separate, complicated processes working together.
- Visual perception isn't just about receiving information; the visual system also processes, interprets, and organizes what we receive.
Image summary: A composite photograph capturing a sequence of a girl jumping from a stone ledge into a swimming pool. The sequence shows her mid-air progression in several stages, from the initial leap to the final splash in the water. The setting is a tropical resort with blue skies, bright sunlight, palm trees, and a view of the ocean in the background.
Illustration used in the slide to represent continuous motion — the kind of motion an akinetopsia patient cannot perceive as fluid.
- A rare neuropsychological condition where a person cannot perceive motion — also called “motion blindness.”
- People with akinetopsia see objects clearly when still, but moving objects appear frozen in place or as a series of still frames (like a strobe-light effect).
- Usually caused by damage to area V5/M.T in the occipital lobe, which is responsible for processing motion.
• Can happen due to stroke, brain injury, or certain neurodegenerative diseases.
• Lecture add-ons:
- The professor discussed Patient L.M, one of the main patients through whom akinetopsia was studied — she developed the condition at age 43 after a blood clot in the brain. Her other visual abilities (recognizing objects, seeing color, discerning detail) stayed normal — only motion perception was affected.
- A moving object looks like a “series of still pictures” instead of continuous movement — for example, if a person walks from your left to your right, you don't see them move; you just perceive them on the left, then suddenly on the right, as if they “teleported.”
- Playing games like patintero becomes very hard — you can't see the other player moving, so you can't tag them in time.
- Akinetopsia makes crossing the street dangerous, because the person can't tell if a car is approaching or how fast it's moving.
- Activities involving motion, like riding a roller coaster, are also difficult since the continuous motion isn't perceived properly.
- Patient L.M had to rely more on hearing to know what was happening around her.
- It also affects the following conversations, since lip movement and changing facial expressions are hard to perceive — you just see a static expression, not the movement of the mouth forming words (this also breaks lip-reading).
- Games that require watching quick movement (e.g., guessing a word from someone silently mouthing it) become very difficult too.
- Damage is usually to V5/M.T, and can result from stroke, brain injury, or neurodegenerative disease.
- Everyday example given: a car that looks like it's stopping and starting instead of moving smoothly, or coffee/juice being poured without the motion of the liquid being perceived smoothly.

The Visual System

- You receive information about the world through multiple sensory modalities: hearing the sound of an approaching train, smelling freshly baked bread, feeling a tap on your shoulder.
- For humans, vision is the dominant sense — reflected in how much brain area is devoted to vision compared to other senses.
- This is the basis for ventriloquism: you see the dummy's mouth moving while the sound actually comes from the dummy's master.
• Lecture add-ons:
- We usually trust what we see more than what we merely hear — this is the idea behind “seeing is believing.” Example given: if someone tells you they saw your partner out with someone else, you're likely to deny it until you see it with your own eyes.
- In ventriloquism, even though the voice is really coming from the person controlling the puppet, because you see the dummy's mouth move, your brain perceives the voice as coming from the dummy — the eyes override the ears, creating the illusion.
- A large part of the brain is devoted to vision, spread across multiple lobes (parietal, temporal, occipital), consistent with the idea of distributed processing discussed in an earlier lecture.
Ventriloquism — vision “overriding” hearing.

The Photoreceptors

Image summary: A cross-sectional medical diagram of a human eye showing the path of light entering from the left through the pupil and cornea, passing through the lens, and focusing on the fovea of the retina. Labeled anatomical parts include the pupil, cornea, iris, lens, retina, fovea, and the optic nerve, which is noted as the path to the brain.
Basic anatomy of the eye: cornea, pupil, iris, lens, retina, fovea, optic nerve.
How does vision operate?
- The process begins with light. Light is produced by objects like the sun, lamps, and candles, then reflects off other objects.
- Light hits the front surface of the eyeball, passes through the cornea and lens, then hits the retina — the light-sensitive tissue lining the back of the eyeball.
- The cornea and lens refract light rays to produce a sharply focused image on the retina.
- The iris opens or closes to control how much light reaches the retina.
• Lecture add-ons:
- Vision begins with light — sources include the sun, lamps, candles, fluorescent lights, and light from electronic devices; light bounces off objects in the environment and that reflected light is what lets us see them.
- Basic pathway emphasized in class: Light goes to Cornea goes to Lens goes to Retina.
- Babies in the womb can't see because there's no light there, even though hearing, taste, and touch develop earlier — vision is the last sense to develop because it needs light. The first light most of us see is right after birth (though almost no one actually remembers it).
- The cornea and lens work like a camera lens, focusing light so a clear image forms on the retina.
- When looking at something near, the muscles around the lens tighten, making the lens more rounded; when looking at something far, the muscles relax and the lens flattens — this lets us focus clearly whether an object is near or far.
:Image summary: An anatomical diagram of a human eye in profile, illustrating the path of light from the pupil to the optic nerve. Labeled components include the cornea, iris, pupil, lens, retina, fovea, and the optic nerve leading to the brain.
The retina is made up of three main layers
- Rods and cones — the light-sensitive cells at the back of the retina that launch the neural process of vision.
- Bipolar cells — interneurons connecting photoreceptors (rods and cones) to ganglion cells; they convey signals from photoreceptors to ganglion cells, which then transmit visual information to the brain.
- Ganglion cells — the output neurons of the retina, sending visual information from the retina to the brain; they integrate signals from bipolar cells and other retinal neurons and transmit this along the optic nerve.
• Lecture add-ons:
- The professor described bipolar cells like "delivery riders" (e.g., Shopee/TikTok riders) — they don't generate the message, they just carry it from the photoreceptors to the next layer (the ganglion cells).
- Basic visual pathway emphasized: Rods & Cones to Bipolar Cells to Ganglion Cells to Optic Nerve to L.G.N (in the thalamus) to Primary Visual Cortex to Occipital Lobe.
- The axons of ganglion cells come together to form the optic nerve, which carries information to the lateral geniculate nucleus (L.G.N), located in the thalamus.
- Nearly all sensory information passes through the thalamus first — the one exception is smell, which goes almost directly to the brain for processing (which is part of why smell can feel so immediate/powerful).
- A fuller version of the pathway given in class: light  cornea ("window of the eye")  anterior chamber  pupil (opening in the iris)  iris dilates/constricts to control light  lens  vitreous  retina, which converts light energy into electrical impulses  optic nerve  primary visual cortex in the occipital lobe.
Image summary: A colorized scanning electron micrograph of the retina showing photoreceptor cells. Labeled in the image are a Cone, which is one of the tapered green structures, and a Rod, which is one of the more cylindrical brown structures.
A. Rods: Sensitive to very low levels of light, so they play an essential role when moving around in semidarkness.
- Rods are color-blind — they distinguish different intensities of light (contributing to brightness perception) but cannot discriminate one hue from another.
- B. Cones: Less sensitive than rods, so they need more light to operate at all.
- Cones are sensitive to color differences — there are three types, each with its own pattern of sensitivity to different wavelengths.
- Rods are useful in dim/semi-dark environments (helping you move around), but because they're color-blind, in low light you mostly see shades of gray rather than the true colors of your surroundings — for example, in a dim room you might not notice how colorful it actually is until the lights are turned on.
- Cones need more light/stimulation to function well and work best in bright light — but they're what let you see color.
Colorized microscope image showing rod and cone photoreceptors.

Color Perception and Acuity

Image summary: A line graph showing the normalized absorbance of different photoreceptors in the eye across a range of wavelengths from 400 to 700 nanometers. Four curves illustrate the spectral sensitivity of S cones, Rods, M cones, and L cones. The S cone curve peaks earliest at approximately 440 nm, followed by the Rod curve peaking around 500 nm, the M cone curve peaking around 530 nm, and the L cone curve peaking around 560 nm.
• Color perception occurs by comparing the outputs from the three types of cones in the retina.
- Each cone type is sensitive to a specific range of wavelengths: Short wavelength (S-cones), Medium wavelength (M-cones), Long wavelength (L-cones).
- Cones also let you perceive fine detail — the ability to see detail is called acuity, and acuity is much higher for cones than for rods.
- This is why you point your eyes toward a target to learn more about it — you're positioning your eyes so the image falls on the fovea, the very center of the retina.
• Lecture add-ons:
- Color perception depends on the firing pattern across the three cone types (similar to how firing patterns of neurons combine to build up a percept, as discussed in an earlier lecture). Examples given:
- Purple: strong firing from S-cones, with weak/no firing from M-and L-cones.
- Blue: equally strong firing from S-cones and M-cones, with only modest firing from L-cones.
- Bright red: mainly involves L-cones. Yellow-to-green: mainly involves M-and L-cones (not S-cones).
- A student explained (and the professor agreed) that cones give better acuity than rods because cones are each connected to fewer neurons individually devoted to processing them, while rods share connections more broadly (geared toward brightness/gray-scale) — so more neural resources go toward processing fine, color-rich detail from cones.
- The fovea is the center of the retina, packed with cones and containing no rods at all — this is why focusing your gaze on something ("squinting" your attention onto it) gives you the clearest, most detailed vision, because the image falls onto the fovea.
- Outside the fovea, the retina is dominated by rods.
Normalized absorbance of S-, M-, L-cones and rods across wavelengths (nm).

Lateral Inhibition (Edge Enhancement)

• Rods and cones do not send signals directly to the cortex.
- Process: Photoreceptors (rods and cones) stimulate bipolar cells leads to bipolar cells excite ganglion cells leads to ganglion cells are spread across the retina leads to the axons of all ganglion cells converge to form the optic nerve.
- The optic nerve leaves the eyeball, carries visual information to the L.G.N in the thalamus, and from the L.G.N information is sent to the primary visual cortex in the occipital lobe.
- The optic nerve is not just a passive cable — it also participates in early visual processing.
- Lateral inhibition helps sharpen visual contrast and improves edge detection in the visual field — known as edge enhancement.
• Lecture add-ons:
- Lateral inhibition means that when a retinal cell is active, it can reduce the activity of its neighboring cells, which makes edges and contrast clearer.
- Worked example from class: three cells, A (edge), B (middle), C (edge), representing a box/shape. Cell B (in the middle) receives the most light and is strongly stimulated, but it also gets inhibited by both of its neighbors (A and C), so its net activity ends up only moderate. Cells A and C are each inhibited by only one neighbor (B), so they end up firing more strongly relative to B.
- The result: the edges of a shape (cells A and C) fire more strongly than the interior (cell B) — this is exactly why we perceive edges clearly and can tell where one object ends and another begins, even when two shapes are placed right next to each other (e.g., two adjoining rectangles are still seen as two separate shapes).
- This edge-detection process happens at the level of the eye/retina itself, before the information even reaches the brain — the visual system starts analyzing and interpreting the scene immediately.
Image summary: A labeled medical illustration of a human eye showing a beam of white light entering from the left and hitting the cornea, which is explicitly labeled with text.

Multiple Types of Receptive Fields

Image summary: A diagram labeled B and C depicting neural firing frequency as a purple circle bisected by a vertical white line.
Image summary: A line graph representing neural firing frequency, showing a single horizontal trace with vertical spikes indicating action potentials. The frequency of these spikes increases significantly within a highlighted light purple rectangular region, demonstrating a change in neural activity.
Image summary: A diagram of neural firing frequency over time. A horizontal axis shows time moving from left to right, with a signal that begins as a low-amplitude waveform and transitions into a high-frequency, high-amplitude series of spikes within a shaded purple rectangular region.
: Image summary: A diagram representing neural firing frequency, consisting of a purple circle with a horizontal white line through the center and the letter A positioned below it.
: A line graph representing neural firing frequency over time. The graph shows a relatively flat baseline frequency with three vertical markers indicating specific firing events, two of which fall within a shaded purple rectangular region.
• In 1981, David Hubel and Torsten Wiesel won the Nobel Prize for their research on the brain's visual system.
- They discovered specialized neurons in the brain, each with its own receptive field — the specific area in the visual world that a neuron responds to.
- Some neurons act like "dot detectors" — they fire most strongly when light appears in a small, circular area in a specific position within the field of view; if light is shown just outside that area, the neuron fires less.
• In this way, each neuron responds to a very specific type of visual input.
- Other specialized neurons act as edge detectors — they fire at their maximum rate only when an edge of a specific orientation appears in their receptive field (horizontal, vertical, or in-between angles).
- I nese cells have a preferred orientation: the closer the stimulus is to that orientation, the stronger the firing; the further away, the weaker the response; a sharply different orientation elicits little to no response.
- Hubel and Wiesel's experiments were done on cats — studying the mammalian visual system.
- Different neurons are dedicated to different triggers: some fire for dots, some for vertical lines, some for horizontal lines, and so on, each answering only to its own specific stimulus.
Dot-, line-, and orientation-selective receptive fields, with corresponding neural firing frequency traces.

Center-Surround Cells

Image summary: A four-part diagram illustrating the effect of light placement on a neuron's firing frequency across a receptive field consisting of a center and a surround. In scenario A, with no light, there is a low baseline firing frequency. In scenario B, light hitting only the center increases the firing frequency. In scenario C, light hitting only the surround decreases the firing frequency below baseline. In scenario D, light covering both the center and surround results in a firing frequency similar to the baseline.
- These neurons are often called center-surround cells.
- Light presented to the center of the receptive field has one effect; light presented to the surrounding area has the opposite effect.
- If both center and surround are strongly stimulated, the cell fires at its baseline rate — no increase or decrease.
• For these cells, a strong uniform stimulus is equivalent to no stimulus at all.
- Instead of simply reporting how bright an entire scene is, center-surround cells tell the brain where brightness changes occur — which is exactly what's needed to detect edges between
- objects, enhance contrast, recognize shapes and patterns, and see objects clearly under different lighting conditions.
- If a room is pure, uniform white light with no edges or objects, there's essentially no stimulus for lateral inhibition or center-surround cells to respond to — nothing fires, because there's nothing to differentiate.
- Edge detectors specifically fire most when they detect an edge/boundary between light and dark in a particular direction (e.g., the outline of a person against the background).

Parallel Processing in the Visual System

Image summary: Two anatomical diagrams of the human brain, one in sagittal section and one in lateral view, both highlighting Area V1, the primary visual projection area, in yellow at the posterior end of the occipital lobe.
- The visual system uses a "divide and conquer" strategy: different types of cells specialize in analyzing specific features of visual input, located in different areas of the cortex.
- Area V.1 (occipital lobe): the first cortical area to receive input from the L.G.N. Contains cells that respond to horizontal lines, vertical lines, and other orientations in specific positions.
- Together, these cells cover the entire visual field and all possible orientations, ensuring some cell responds to any visual stimulus.
- This connects to “distributed representation/processing” discussed in a previous lecture — different brain parts (parietal, temporal, occipital) each contribute a piece of the overall percept, and most of that earlier discussion actually drew on findings from visual system research.
- Beyond V.1, the occipital lobe also contains V.2, V.3, V.4, P.O, and M.T/V5 (the motion area discussed earlier under akinetopsia).
- The parietal cortex is involved in spatial and motion information; the temporal cortex handles object recognition and form.
- If there's damage to V.1, a person may not consciously see objects in part of their visual field; damage to V.4 causes color blindness; damage to V5/M.T causes akinetopsia.
Area V1 in the occipital lobe — the first cortical area to receive input from the LGN.

Specialized Brain Areas in Vision

Image summary: A flow diagram illustrating the neural pathways of the visual system, beginning with the retina and LGN, which then feed into the occipital cortex areas V1 through V4, PO, and MT. From there, the pathways branch into the parietal cortex, including VIP, MST, LIP, and 7a, and the inferotemporal cortex, including TEO and TE, with arrows indicating the complex, interconnected flow of information between these brain regions.
- Visual processing involves multiple brain areas, each with specific functions: Occipital cortex (V.1, V.2, V.3, V.4, P.O, M.T), Parietal cortex (spatial and motion information), Temporal cortex (object recognition and form).
- Area M.T: sensitive to direction and speed of movement; damage here can cause akinetopsia (motion blindness).
- Area V.4: sensitive to color and shape; cells fire most strongly when both a specific color and shape are present.
- Each region hands off information to the next, like passing messages along, until the full percept of what you're looking at is assembled.

Advantages of Parallel Processing

Image summary: A conceptual digital illustration of a human head in profile with a visible pink brain. The head is composed of metallic, circuit-like patterns, and the brain is connected to a network of glowing nodes and lines that merge into a background of blue circuit board traces, representing the intersection of human intelligence and artificial intelligence.
- All specialized visual areas (e.g., Area M.T, Area V.4) are active simultaneously — Area M.T detects motion, Area V.4 detects color and shape, and these processes happen at the same time, not one after another.
- This simultaneous activity is called parallel processing. By contrast, serial processing does steps one at a time, in sequence (e.g., retina leads to L.G.N leads to V 1 leads to V 2 leads to V 3 leads to V 4 leads to V 5 slash M.T, one after another).
- 1. Speed: different brain areas don't need to wait for one another — for example, shape analysis can begin immediately without waiting for motion or color processing to finish.
- 2. Mutual Influence: systems can inform and influence each other in real time — motion interpretation may depend on perceived 3D shape, and shape perception may depend on observed motion; the brain allows concurrent processing and negotiation between systems.
• Lecture add-ons:
- Most researchers favor parallel processing over strictly serial processing because it's faster (the brain doesn't have to wait for one stage to fully finish before starting the next) and because it allows teamwork between systems — sometimes you need to know an object's shape to understand its motion, and other times you need to see the motion first to understand the shape; parallel processing lets these systems help each other out in real time.

Two Major Visual Pathways in the Brain

The experiment by Leslie Ungerleider and Mortimer Mishkin (1982) is a landmark study showing that different parts of the brain handle different kinds of visual information.
• Method: Brain ablation — surgically removing or destroying a specific brain area to study its function.
The experiment used monkeys as subjects; brain ablation basically means temporarily/surgically removing a brain part to see what function is lost — a method that, before modern technology, was one of the only ways to figure out what different brain regions actually did.

Object Discrimination Problem (the "What" Task)

Image summary: An anatomical diagram of a human brain in profile, identifying key regions. Labels pinpoint the Parietal lobe at the top, the Occipital lobe at the back (highlighted in yellow), and the Temporal lobe along the bottom. Two blue arrows designate specific functional areas: the Posterior parietal cortex in the upper back region and the Inferotemporal cortex in the lower back region.
- Step 1: Monkey sees a sample object (e.g., a rectangular solid).
- Step 2: Later shown two objects — the sample object and a different one (e.g., a triangular solid). Goal: test the monkey's ability to identify and remember an object.
- Temporal lobe ablation to monkeys could no longer solve the object discrimination task. This means the temporal lobe is crucial for identifying objects to the ventral stream, or “what pathway.”
• Lecture add-ons:
- Before ablation, the monkey could correctly identify which object was the rectangle and which side it was on. After the temporal lobe was removed, the monkey could no longer tell the rectangle from the triangle at all.

Landmark Discrimination Problem (the "Where" Task)

- Monkey sees two identical food wells, but one is always closer to a tall cylinder landmark. Goal: test the monkey's ability to locate something in space relative to a cue.
- Parietal lobe ablation to monkeys could no longer solve the landmark discrimination task. This means the parietal lobe is crucial for locating objects in space to the dorsal stream, or “where pathway.”
- Before ablation, the monkey could correctly choose the food well closer to the cylinder landmark. After the parietal lobe was removed, it could no longer tell which food was closer to the landmark.

Both Streams Work Together

Image summary: A diagram of a human brain highlighting two visual processing pathways. The dorsal stream, colored red and located in the parietal lobe, is labeled vision-for-action. The ventral stream, colored blue and located in the temporal lobe, is labeled vision-for-perception. Both pathways originate from the occipital lobe at the back of the brain.
Both systems function simultaneously: you recognize an object while also understanding its position in space — an example of parallel processing in higher-level visual cognition.
• Two-streams hypothesis: Ventral stream (temporal lobe) leads to what something is (object recognition); Dorsal stream (parietal lobe) leads to where something is (spatial location).
• Lecture add-ons:
Example: in a coffee shop, the ventral/"what" pathway helps you recognize that what's in front of you is a cup of coffee, while the dorsal/"where" pathway helps you know exactly where it is so you can reach for it — recognizing isn't enough on its own; you also need to locate it to actually pick it up.

Evidence from Brain Damage: “What” versus “Where” Systems

: Image summary: A photograph of a man against a black background holding a bunch of bananas. A white thought bubble containing a question mark is positioned above his head, suggesting confusion or a question regarding the fruit.
1. Damage to the "What" System (temporal lobe) leads to visual agnosia: the inability to recognize objects visually. Patients can't recognize objects but can still reach for them and know where they are.
- 2. Damage to the "Where" System (parietal lobe) leads to problems reaching or locating objects in space.
- With “where” system damage, you might see the coffee cup and know it's there, but when you try to reach for it, you misjudge its location and grab the wrong spot (e.g., reaching to the right when the cup is actually on the left) because you can no longer properly locate the object in space.

Damage to Other Visual Areas

Image summary: An educational graphic illustrating total color blindness, also known as monochromacy or achromatopsia. The image is split vertically to compare two perspectives of the same mountain and lake landscape. The left side, labeled Normal Vision, shows the scene in full color with a blue sky and green trees. The right side, labeled Monochromatic Vision, shows the exact same scene in grayscale, depicting the world as it is perceived by someone with total color blindness.
- Damage to the motion area (Area M.T) to Akinetopsia: the patient sees the world as a series of frozen images and cannot judge the speed or direction of moving objects.
- Damage to the color area (Area V.4) to Cerebral achromatopsia: the patient loses color vision and sees the world in "dirty shades of gray."

Visual Maps and Firing Synchrony

There is an ongoing debate about how the visual system solves the "binding problem" — how it combines separate information about color, shape, and motion into a single, coherent perception of one object.
• Researchers have identified three main elements that contribute to this process: spatial position, neural synchrony, and attention.

1. Spatial Position

- Location information provides a frame of reference used to solve the binding problem.
- Spatial position is a major organizing theme in all the brain areas concerned with vision, with each area providing its own map of the visual world.
- The brain makes sure the neurons/action potentials for a given object are encoded at one specific spatial location, which is how the object is later remembered and how the brain tells objects apart. Spatial position also helps us judge how near or far something is, and whether something nearby might be dangerous — important for survival.

2. Neural Synchrony and Feature Binding

- Spatial position isn't the whole story. Evidence suggests the brain uses special rhythms to identify which sensory elements belong together.
- Example: one group of neurons fires maximally for vertical lines, another fires maximally for downward motion. If a vertical line is moving downward and both groups fire in synchrony, these features are registered as part of the same object. If not synchronized, the features are not bound together.

3. The Role of Attention in Binding

Synesthesia 0123456789
- A crucial factor causing this synchrony is attention — it plays a key role in binding together the separate features of a stimulus.
- Evidence: when attention is overloaded, people often make conjunction errors (e.g., seeing a blue H and a red T but reporting a blue T and a red H).
- Brain recordings show synchronized neural firing when a stimulus is attended, but not for unattended stimuli.
• Lecture add-ons:
- Attention is the key ingredient that makes neural synchrony happen — when we focus on something, the brain enhances the synchronization of neurons responding to its different features, which is how it "labels" those action potentials as belonging to one object.
- Conjunction errors happen when attention is overloaded — for example, looking at a page full of differently colored letters and then being asked what color a specific letter was; with too much competing for attention, it's easy to mix up which color went with which item.

Form Perception

- Detection is just the start of the process, because the visual system still has to assemble these features into recognizable wholes.
- When we look at something, our eyes detect basic features like color, edges, motion, and light. But recognizing an object — say, seeing a cube or a face — requires more than just noticing those features; the brain has to combine and organize the simple features into a clear, meaningful form. This process is called form perception.
- Two possible processes:
- Bottom-up processing: starts at the "bottom" or beginning of the system, when environmental energy stimulates the receptors.
- Top-down processing: originates in the brain, at the "top" of the perceptual system. This knowledge lets people rapidly identify objects and scenes, and go beyond mere identification to determining the story behind a scene.
• Lecture add-ons:
- Form perception doesn't require observing every single feature of something one by one (e.g., every feature of a face) each time — because of bottom-up versus top-down processing.
- Bottom-up processing is literal, ground-up encoding of every detail — used the first time you encounter someone or something completely new, since the brain has no prior experience to draw on yet.
- Top-down processing means the brain already has relevant knowledge/experience and unconsciously fills in the gaps — for example, recognizing someone in a crowded room just from the back of their head or the way they walk, without needing to see their face, because you already know them so well.
- Example given: the first time you encounter a pothole on a street, you have to consciously process it (bottom-up); after experiencing it repeatedly, you start avoiding it automatically because your brain has learned/represented the gap (top-down).
- Another example: going to Baguio for the first time (bottom-up — absorbing everything, what it looks like, where to go) versus having been there many times already (top-down — you already know the shortcuts, the food, the places to visit).

A. Unconscious Inference

- Hermann von Helmholtz offered a hypothesis based on the simple inverse relationship between distance and retinal image size: if an object doubles its distance from the viewer, the size of its image is reduced by half; if it triples its distance, the image size is reduced to a third of its initial size.
- Helmholtz knew we don't consciously calculate this every time we perceive an object's size, but believed we're calculating it nonetheless — hence he called the process an unconscious inference.

B. The Gestalt Principles

Image summary: A photograph labeled "Darkness" consisting of a solid black rectangle.
]')Image summary: A graphic illustration titled "Darkness" featuring a bright, white and yellow glowing sun with sharp, radiating spikes positioned on the right side of a solid black background.
(c) The second light flashes
Image summary: A diagram showing two radiant, sun-like spheres set against a solid black background, with a yellow arrow pointing from the left sphere to the right sphere, labeled as Darkness.
(d) Flash—dark—flash - The Gestalt approach to perception originated, in part, as a reaction to Wilhelm Wundt's structuralism.
Image summary: A stippled line drawing of a person's face, created using a dense collection of small black dots on a white background to form the features and shading of the portrait.
- Wilhelm Wundt's idea (Structuralism): our overall perception is made by adding up many tiny sensations.
- The Gestalt Reaction: Gestalt psychologists disagreed — they believed perception can't be explained just by adding up small parts. The whole is different from the sum of its parts.
• Lecture add-ons:
- Structuralism (Wundt) is compared to identifying every individual note that makes up a chord on a piano — the Gestalt view rejects that we perceive things this piece-by-piece; the brain instead has a tendency to simplify and make sense of things as a whole.
- Classic example: flip-book/motion-picture-style drawings — individual still frames that, when flipped quickly, are perceived as continuous motion, even though each frame is really just a separate static picture. This illustrates the Gestalt idea that we perceive the whole (the motion), not just the isolated individual pictures.
(a) One light flashes

Good Continuation

Image summary: A photograph of a tangled coil of thick, light-colored rope. An orange line highlights a specific section of the rope, tracing a path from the outer edge of the coil into its center.
Image summary: A graphic design presentation depicting a carousel of five social media slides for a brand called INFALE Fashion. A smartphone in the center displays Slide 3, which features a collage of female models and the text VISUAL Merchandise. The other four slides are arranged linearly behind the phone, with Slide 1 on the far left and Slide 5 on the far right, showcasing a consistent aesthetic of earth tones, fashion photography, and minimalist typography.
Points that, when connected, form straight or smoothly curving lines are seen as belonging together. We naturally follow the smoothest path. If one object overlaps another, we still see it as continuing behind the overlap.
• Lecture add-ons:
Example: a coiled rope is perceived as one single, long, continuous rope, even though it loops over itself — we don't see it as separate broken segments.
Another example: Instagram carousel posts — even though each photo is technically a separate image, the brain tends to mentally stitch them into one continuous picture rather than isolated parts.
(a)
(b)

Pragnanz (Good Figure / Simplicity)

Image summary: A line drawing comparing two geometric configurations of five circles. In configuration (a), the circles overlap in a staggered arrangement of three on top and two on bottom. In configuration (b), the circles are modified such that the overlapping areas are replaced by concave cutouts, creating a series of interconnected, crescent-like gaps between the circles.
Image summary: A graphic of the Olympic rings, consisting of five interlocking circles of equal size arranged in two rows. The top row contains blue, black, and red rings, while the bottom row consists of yellow and green rings.
From the German word for "good figure," this principle means we see patterns in the simplest way possible. Our brains prefer organizing what we see into the most straightforward and stable form.
• Lecture add-ons:
Example: overlapping circles (like the Olympic rings) are perceived simply as “overlapping circles,” not as a complicated set of unique, oddly-shaped overlapping regions — the brain defaults to the simplest possible interpretation.

Similarity

Image summary: A screenshot of a social media profile page for cocacola, featuring a red-themed image gallery. The grid consists of various graphic designs and typography, including phrases such as "REFRESH YOUR PERSPECTIVE", "EVERYONE Welcome", "ALL FOR LOVE. LOVE FOR ALL.", "LOVE HARDER", "HAPPINESS LOOKS GOOD ON YOU", and "LOVE IS THE STANDARD". Some images feature stylized Coca-Cola bottles and heart motifs. Above the profile, there are two diagrams of dots: one showing a 5x5 grid of red dots and another showing a 5x5 grid with a mix of blue and red dots.
• We group objects together when they share similar features — such as color, size, shape, or orientation.
Items that look alike (same color, size, shape, orientation) catch our attention as belonging together and are easier to perceive as a group — for example, a very aesthetic, color-coordinated Instagram feed reads as a cohesive set precisely because of this similarity effect.
Closure
According to the law of closure, we perceive elements as belonging to the same group if they seem to complete some entity. Our brains often ignore contradictory information and fill in gaps.
Example: a zebra rendered only in disconnected black stripes (no solid outline) is still immediately recognized as a zebra, because the way the stripes are arranged and grouped together lets the mind complete the shape, even though the pieces aren't literally touching.
Figure-Ground
Image summary: A two-panel composition featuring a visual illusion and a row of physical objects. The top panel is a black-and-white silhouette illustration of the Rubin vase, creating an optical illusion where the white center can be seen as a vase or the black sides as two facing profiles. The bottom panel is a photograph of five black vases of varying shapes and heights standing in a row against a wooden background.
People instinctively perceive objects as either being in the foreground or the background. They either stand out prominently in front (the figure) or recede into the back (the ground).
• Lecture add-ons:
Classic example: the face/vase illusion, where either two faces or a vase can be seen depending on which part you perceive as figure versus ground. Another example given: a set of pillar/candlestick shapes that can also be seen as a row of human silhouettes in the negative space, depending on what you focus on as the "figure."

C. Taking Regularities of the Environment into Account

- Modern perceptual psychologists take experience into account by noting that certain characteristics of the environment occur frequently.
• There are two types: Physical Regularities and Semantic Regularities.

A. Physical Regularities

Image summary: A vertical composite of two photographs. The top image is a forest scene featuring tall, thin trees with prominent, sprawling surface roots spreading across a dark, leaf-covered forest floor. The bottom image shows a landscape of towering, light-colored sandstone rock pillars and monoliths rising above a dense forest of green evergreen trees.
- Physical Regularities are regularly occurring physical properties of the environment. For example, there are more vertical and horizontal orientations in the environment than oblique (angled) orientations.
- This occurs in human-made environments (e.g., buildings contain lots of horizontals and verticals) and also in natural environments (trees and plants are more likely to be vertical or horizontal than slanted).
• Lecture add-ons:
- Because these vertical/horizontal patterns show up so often — in buildings, roads, institutions, trees, and plants — the brain processes them faster, since they're already "encoded" as familiar; this is essentially top-down processing applied to physical regularities.

B. Semantic Regularities

- In language, semantics refers to the meanings of words or sentences. Applied to perceiving scenes, semantics refers to the meaning of a scene.
- Example: food preparation, cooking, and eating occur in a kitchen; waiting around, buying tickets, checking luggage, and going through security happen in airports.
- Semantic regularities are the characteristics associated with the functions carried out in different types of scenes.
- Our visualizations contain information based on our knowledge of different kinds of scenes. This knowledge of what a given scene typically contains is called a scene schema, and the expectations it creates contribute to our ability to perceive objects and scenes.
- Examples given in class: a bathroom is associated with bathing, grooming, and personal hygiene; a kitchen is associated with food prep, cooking, and eating; an airport is associated with buying tickets, checking luggage, and going through security.
- Seeing a microscope immediately brings to mind a laboratory setting; seeing a zebra brings to mind a zoo or savanna — you don't just see the object, you imagine the whole setting it typically belongs to. That association is scene schema.

Perception and Action: Behavior

Movement Facilitates Perception
Image summary: A photograph of a contemporary sculpture representing a human head and neck in profile. The sculpture is constructed from fragmented, interlocking, and floating bronze-colored metallic pieces, creating a shattered or disassembled effect with gaps between the forms. Small blue accents are integrated into the upper and rear sections of the head.
- Although movement adds complexity to perception (compared to sitting still), movement also helps us perceive objects in the environment more accurately.
- One reason: moving reveals aspects of objects that are not apparent from a single viewpoint.
- Seeing an object from different viewpoints provides added information that results in more accurate perception, especially for unusual objects, such as a distorted/anamorphic image that only resolves correctly from a specific angle.
• Lecture add-ons:
- Physical and semantic regularities are things we can perceive without needing to move (e.g., just standing in a room, you already know what's going on there). But some visual perception requires movement — as you change position, what you perceive changes and updates.
- A common real-world example: museum/forced-perspective art installations where a person appears taller or shorter, or an image only "resolves" correctly, depending on the viewer's position/angle — moving around changes your understanding of the scene.
- Similarly, some illusions look like disconnected, broken-up pieces from one angle, but as you move and change your position, the pieces resolve into one coherent image — movement literally builds up the correct percept.
You have reached the end of the document.