Perception & Design

PSY 410: Data Science for Psychology

Dr. Sara Weston

2026-04-27

How we see data

What’s wrong with this?

Deliberately poorly designed bar chart with red and green bars, tiny text, yellow background, heavy gridlines, and an uninformative title — used as a 'what's wrong' warm-up exercise.

Red-green palette (colorblind-unfriendly). No informative title. Tiny text. Gridline noise. Bars hiding the data.

Today we learn why your brain rejects bad figures — and how to design ones that work.

Your brain processes some visual features before you even think

Some visual features are processed almost instantly — before conscious attention kicks in:

Attribute Example
Color A red dot among blue dots
Size A large circle among small ones
Position A point far from the others
Shape A triangle among circles
Orientation A tilted bar among vertical bars

These are your tools for drawing the viewer’s eye.

Preattentive in action

Three panel demonstration of preattentive processing: the left panel uses color to make red points pop among gray ones, the middle panel uses size to make two large points stand out, and the right panel uses position to isolate one outlier far from a cluster.

Your eye goes to the red points, the big points, and the outlier — instantly.

Working memory limits what you can encode

Research suggests working memory holds roughly 4 items at a time (Cowan, 2001).

In data visualization, this means:

  • Limit categorical groups in a legend to 4–5 maximum
  • Each additional aesthetic (color, shape, size) uses up a “slot”
  • If you have 8 conditions, consider faceting rather than overlapping
Two side-by-side bar charts: the left has 4 clearly distinguishable groups using a Set2 palette; the right has 8 groups where colors become hard to tell apart, illustrating how working memory limits categorical encoding in figures.

Gestalt principles: how the brain groups things

The brain automatically groups visual elements — whether you intend it or not.

Proximity

Six gray dots arranged in two tight clusters separated by empty space — demonstrating the Gestalt principle of proximity.

Nearby items feel like a group

Similarity

Six dots in a row, alternating blue and orange — demonstrating Gestalt similarity grouping by color.

Same appearance = same group

Enclosure

Two clusters of gray dots, each enclosed by a shaded rectangle — demonstrating the Gestalt principle of enclosure.

Inside a border = same group

Gestalt principles show up everywhere in ggplot2

Principle How it shows up in your figures
Proximity facet_wrap() puts related panels side by side
Similarity Color-coding makes same-group points feel connected
Enclosure A shaded background or border sets a group apart
Continuity Connecting points across time implies a single trend

These groupings happen whether you design them or not — use them intentionally.

Color theory

Three types of color palettes

Type When to use Example
Sequential One continuous variable (low → high) Blues, viridis
Diverging Values relative to a midpoint Red–white–blue
Qualitative Categorical groups (no order) Set1, tab10

Using the wrong type is one of the most common visualization mistakes.

First question: fill or color?

Which aesthetic did you map in aes()? That tells you which scale family to use.

# Geoms with `fill=`: bars, boxplots, tiles, density
ggplot(data, aes(x, y, fill = group)) + geom_boxplot() +
  scale_fill_manual(values = c(...))     # ← scale_fill_*

# Geoms with `color=`: points, lines
ggplot(data, aes(x, y, color = group)) + geom_point() +
  scale_color_manual(values = c(...))    # ← scale_color_*

Same pattern, just swap fill for color throughout.

Second question: which palette?

Match the palette to your data:

Your data Function
Continuous numeric (e.g., reaction time) scale_fill_viridis_c()
Ordered categories (e.g., dosage levels) scale_fill_viridis_d()
Qualitative groups (e.g., therapy type) scale_fill_brewer(palette = "Set2")
Diverging from a midpoint (e.g., change scores) scale_fill_distiller(palette = "RdBu")
Hand-picked colors scale_fill_manual(values = c(...))

_c for numeric, _d for character/factor. Same options work for scale_color_*.

Sequential palettes

A row of ten tiles colored with the viridis sequential palette, transitioning smoothly from dark purple on the left to bright yellow on the right, illustrating how sequential palettes encode low-to-high values.

Diverging palettes

Use when values diverge from a meaningful midpoint — often 0 or a baseline.

A row of ten tiles colored with the RdBu diverging palette, transitioning from dark blue on the left through white in the middle to dark red on the right, illustrating how diverging palettes encode values relative to a midpoint.

Qualitative palettes

A row of eight tiles in distinct pastel colors from the ColorBrewer Set2 qualitative palette, showing how each category gets a visually distinguishable hue with no implied ordering.

Colorblind-friendly choices

About 8% of men have some form of color vision deficiency. Red-green is the most common.

# viridis is colorblind-safe AND sequential
scale_fill_viridis_d()   # discrete
scale_fill_viridis_c()   # continuous

# ColorBrewer palettes designed for colorblindness
scale_fill_brewer(palette = "Set2")

# Or set colors manually with safe choices
scale_fill_manual(values = c("#0072B2", "#E69F00", "#009E73"))
# (blue, orange, green — distinguishable for most color vision types)

A bad color choice vs a good one

Side-by-side comparison of two boxplots: the left uses red and green fills that are indistinguishable to colorblind viewers, while the right uses blue and orange fills that are accessible to virtually all viewers.

Why cluttered figures fail: cognitive load theory

Cognitive load theory (Sweller, 1988) says mental effort has three components:

Intrinsic load

The complexity of the data itself — you can’t reduce it

Extraneous load

Clutter, noise, redundant elements — this is what we control

Germane load

The insight the reader builds — what we’re trying to maximize

Every gray background, redundant legend, and unlabeled variable is extraneous load — it taxes working memory without adding information.

Design goal: minimize extraneous, maximize germane.

Decluttering

The default is cluttered

ggplot2’s default theme adds a lot of visual noise. Compare:

Side-by-side comparison of the same boxplot using ggplot2's default gray theme on the left versus the cleaner minimal theme on the right, showing how removing the gray background, heavy gridlines, and redundant legend reduces visual clutter.

The declutter checklist

Ask: does this element help the reader understand the data?

Remove

Element Why
Gray background Noise, no info
Gridlines (most) Distraction
Redundant legend X-axis says it
Generic axis labels Add units instead

Keep

Element Why
Title + subtitle Orients the reader
Caption Source, N, error bar type
Meaningful color Highlights comparisons

Meet the element_* family

Every non-data piece of a ggplot is one of four kinds of thing:

Function Key arguments Used for
element_text() size, face, color, hjust Titles, axis text, legend text
element_line() color, linewidth Axis lines, gridlines, ticks
element_rect() fill, color Panel background, legend box
element_blank() — Removes the element entirely

theme() puts elements to use

Inside theme(), name the piece of the plot, then assign an element_* function to style it:

theme(
  plot.title      = element_text(face = "bold", size = 16),  # text → bold, larger
  panel.grid      = element_line(color = "gray90"),           # lines → subtle gray
  plot.background = element_rect(fill = "white"),             # rect → white background
  axis.ticks      = element_blank(),                          # blank → remove ticks
  legend.position = "none"                                    # special: just "none"/"top"/etc.
)

Not sure what an argument is called? Run ?theme — there are ~90 options.

Pair coding break

Your turn: 10 minutes

I’ll put a cluttered graph on screen. With a partner, rewrite it to follow the design principles we just covered.

Remove at least 5 unnecessary elements. Make it tell a clear story.

Tip

Think about: colors, legend, labels, theme, size mapping, and whether every aesthetic is adding information.

Your turn: 10 minutes

First, run this to create the dataset:

Data setup (click to expand)
set.seed(123)
reaction_data <- tibble(
  participant = rep(1:40, each = 2),
  condition   = rep(c("Control", "Treatment"), 40),
  rt          = c(rnorm(40, mean = 520, sd = 60), rnorm(40, mean = 480, sd = 55)),
  accuracy    = c(rbinom(40, 1, 0.82), rbinom(40, 1, 0.88)),
  age_group   = rep(c("Young", "Young", "Older", "Older"), 20)
) |>
  mutate(rt = round(rt, 1))

Now fix this:

# The cluttered version — fix this!
reaction_data |>
  ggplot(aes(x = condition, y = rt, fill = condition, color = condition, size = accuracy)) +
  geom_point() +
  geom_boxplot(alpha = 0.3) +
  scale_fill_manual(values = c("Control" = "red", "Treatment" = "green")) +
  scale_color_manual(values = c("Control" = "red", "Treatment" = "green")) +
  labs(x = "condition", y = "rt", title = "data") +
  theme_gray()
Deliberately cluttered scatterplot with red and green colors, points sized by a binary variable, redundant color and fill mappings, uninformative title, and the default gray theme — students are asked to redesign it.

Solution: Redesigned reaction time figure

reaction_data |>
  ggplot(aes(x = condition, y = rt,
             fill = condition,
             color = condition)) +
  geom_jitter(width = .1) +
  geom_boxplot(alpha = 0.3) +
  scale_fill_viridis_d() +
  scale_color_viridis_d() +
  scale_y_continuous(
    labels = scales::label_number(suffix = " ms")) +
  labs(
    x = NULL,
    y = NULL,
    title = "Reaction time by condition and accuracy",
    caption = "Reaction time in miliseconds; 0 = inaccurate, 1 = accurate"
  ) +
  facet_wrap(~accuracy,
             labeller = as_labeller(c("0" = "incorrect", "1" = "correct"))) +
  theme_bw() +
  theme(legend.position = "none",
        plot.title.position = "plot")

Solution: output

Boxplots with jittered points showing reaction time by condition (Control vs Treatment), faceted by accuracy ('incorrect' vs 'correct'). Viridis fill and color, ms-suffixed y-axis, minimal black-and-white theme, no legend.

Critiquing bad graphs

What’s wrong here? (1 of 3)

Bar chart comparing mean reaction times for Control and Treatment conditions where the y-axis starts at 478 instead of 0, making a small ~40 ms difference appear dramatically large.

The problem: The y-axis starts at 478, not 0. The difference looks massive — but it’s only ~40 ms. A truncated axis exaggerates the effect.

What’s wrong here? (2 of 3)

Boxplot with red and green fills, a yellow background, heavy gridlines, variable names as axis labels, and the uninformative title 'Boxplot' — demonstrating multiple common design mistakes at once.

The problems: Red-green palette (colorblind-unfriendly). Title says “Boxplot” (a label, not a finding). Variable names as labels. Distracting background color and gridlines.

What’s wrong here? (3 of 3)

Pie chart showing reaction times arbitrarily binned into Fast, Medium, and Slow categories, hiding the continuous distribution, individual observations, and any comparison between conditions.

The problems: Continuous data was binned into arbitrary categories, then displayed as a pie chart. We lost the actual reaction times, the condition comparison, and the ability to see distributions. A histogram or density plot would show far more.

Get a head start

Assignment 4 preview

Assignment 4 will ask you to:

  1. Create a “bad” version of a figure — deliberately violate design principles
  2. Create a “good” version following what we covered today
  3. Create a colorblind-accessible version

Assignment 4 preview

Start experimenting now:

  • Take the reaction_data dataset
  • Make the worst possible version of a figure
  • Then make it great
  • What did you change?

Wrapping up

Design principles checklist

Before next class

📖 Read:

✅ Practice:

  • Try the “bad vs good” exercise on your own
  • Find a graph online and identify ways to improve it

Key takeaways

  1. Preattentive processing — your brain detects color, size, and position before conscious thought
  2. Working memory — limit categories to 4–5; every aesthetic uses a slot
  3. Gestalt principles — proximity, similarity, and enclosure create grouping automatically
  4. Cognitive load — minimize extraneous clutter so the insight can land
  5. Palette choice — match sequential, diverging, or qualitative to your data type
  6. Design for colorblindness — always; 1 in 12 men has color vision deficiency

These are not design opinions — they are perceptual psychology.

The one thing to remember

You’re not designing for yourself. You’re designing for a reader who will look at your figure for five seconds.

Next time: Exploratory Data Analysis