Layers & Aesthetics

PSY 410: Data Science for Psychology

Dr. Sara Weston

2026-04-22

Setup

library(tidyverse)
library(nycflights13)
library(showtext)

showtext_auto()  # lets ggplot render Unicode symbols like ◀ ▶

set.seed(422)
some_flights <- flights |>
  filter(carrier %in% c("9E", "B6", "HA")) |>
  filter(!is.na(air_time)) |>
  group_by(carrier) |>
  slice_sample(n = 100)

min_time <- min(some_flights$air_time)
max_time <- max(some_flights$air_time)

Base figure

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  stat_summary(fun = "mean", geom = "bar")
Bar chart of mean air_time by carrier. Three bars of different heights — 9E short, B6 medium, HA much taller.

Remove legend and turn sideways

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  stat_summary(fun = "mean", geom = "bar") +
  guides(fill = "none") +
  coord_flip()
The same bar chart rotated so carriers appear on the left and air_time extends horizontally. Legend removed.

Boxplot

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_boxplot() +
  guides(fill = "none") +
  coord_flip()
Horizontal boxplots of air_time for each of three carriers. 9E and B6 have tight boxes at the low end; HA's box sits far to the right.

Points

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time)) +
  geom_point() +
  coord_flip()
Scatter plot showing every flight on a single horizontal line per carrier — points overlap heavily because the carrier axis is categorical.

Jitter

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time)) +
  geom_jitter(width = .2) +
  coord_flip()
Same plot, but each point now jittered vertically so overlapping observations are visible. Three horizontal clouds, one per carrier.

Colored

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  coord_flip()
The jittered dot plot with each point filled by carrier and outlined in black using shape 21.

Cleaner theme

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  coord_flip() +
  theme_minimal()
Same plot with theme_minimal() applied — grey background replaced with white, gridlines subtler.

Better colors, remove legend

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  scale_fill_manual(values = c("9E" = "#15616d", "B6" = "#ff7d00", "HA" = "#78290f")) +
  guides(fill = "none") +
  coord_flip() +
  theme_minimal()
Same dot plot with a custom palette — teal for 9E, orange for B6, deep red-brown for HA. Legend removed.

Clean up labels

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  scale_y_continuous(labels = scales::label_number(suffix = " min")) +
  scale_fill_manual(values = c("9E" = "#15616d", "B6" = "#ff7d00", "HA" = "#78290f")) +
  guides(fill = "none") +
  labs(x = NULL, y = NULL, title = "Air time by carrier") +
  coord_flip() +
  theme_minimal()
Dot plot with a title and the horizontal-axis values suffixed with ' min'. Default x and y labels removed.

Reference lines at min and max

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  geom_hline(yintercept = min_time, linetype = "dashed") +
  geom_hline(yintercept = max_time, linetype = "dashed") +
  scale_y_continuous(labels = scales::label_number(suffix = " min")) +
  scale_fill_manual(values = c("9E" = "#15616d", "B6" = "#ff7d00", "HA" = "#78290f")) +
  guides(fill = "none") +
  labs(x = NULL, y = NULL, title = "Air time by carrier") +
  coord_flip() +
  theme_minimal()
Same dot plot with two vertical dashed reference lines marking the overall shortest and longest flights in the sample.

Annotate the extremes

Code
some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  geom_hline(yintercept = min_time, linetype = "dashed", alpha = .4) +
  geom_hline(yintercept = max_time, linetype = "dashed", alpha = .4) +
  annotate("text", x = 3.5, y = min_time, hjust = 0,
           label = paste0("◀ shortest: ", min_time, " min"),
           color = "grey30", size = 4) +
  annotate("text", x = 3.5, y = max_time, hjust = 1,
           label = paste0("longest: ", max_time, " min ▶"),
           color = "grey30", size = 4) +
  scale_x_discrete(expand = expansion(add = c(0.6, 1))) +
  scale_y_continuous(labels = scales::label_number(suffix = " min")) +
  scale_fill_manual(values = c("9E" = "#15616d", "B6" = "#ff7d00", "HA" = "#78290f")) +
  guides(fill = "none") +
  labs(x = NULL, y = NULL, title = "Air time by carrier") +
  coord_flip(clip = "off") +
  theme_minimal()
Same plot with two small arrows and labels at the top — one pointing left to the shortest reference line, one pointing right to the longest.

Carrier labels inline

Code
carrier_labels <- some_flights |>
  group_by(carrier) |>
  summarize(mean_time = mean(air_time)) |>
  mutate(name = c("9E • Endeavor Air", "B6 • JetBlue", "HA • Hawaiian Airlines"))

some_flights |>
  ggplot(aes(x = carrier, y = air_time, fill = carrier)) +
  geom_jitter(width = .2, shape = 21, color = "black", stroke = .4) +
  geom_hline(yintercept = min_time, linetype = "dashed", alpha = .4) +
  geom_hline(yintercept = max_time, linetype = "dashed", alpha = .4) +
  annotate("text", x = 3.5, y = min_time, hjust = 0,
           label = paste0("◀ shortest: ", min_time, " min"),
           color = "grey30", size = 4) +
  annotate("text", x = 3.5, y = max_time, hjust = 1,
           label = paste0("longest: ", max_time, " min ▶"),
           color = "grey30", size = 4) +
  geom_text(data = carrier_labels,
            aes(x = carrier, y = mean_time, label = name),
            inherit.aes = FALSE,
            hjust = 0.5, nudge_x = -0.35,
            fontface = "bold", size = 4.2) +
  geom_text(data = carrier_labels,
            aes(x = carrier, y = mean_time,
                label = paste0("avg: ", round(mean_time), " min")),
            inherit.aes = FALSE,
            hjust = 0.5, nudge_x = -0.6,
            size = 3.5, color = "grey30") +
  scale_x_discrete(expand = expansion(add = c(0.8, 1)), labels = NULL) +
  scale_y_continuous(labels = scales::label_number(suffix = " min")) +
  scale_fill_manual(values = c("9E" = "#15616d", "B6" = "#ff7d00", "HA" = "#78290f")) +
  guides(fill = "none") +
  labs(x = NULL, y = NULL, title = "Air time by carrier") +
  coord_flip(clip = "off") +
  theme_minimal()
Same plot with the carrier tick labels on the left removed and replaced by inline labels sitting just below each row of dots: '9E • Endeavor Air', 'B6 • JetBlue', 'HA • Hawaiian Airlines'.

Choosing the right geom

Geom selection guide

Your data Good geom Avoid
One continuous variable geom_histogram(), geom_density() geom_bar()
One categorical variable geom_bar() geom_histogram()
Two continuous geom_point(), geom_smooth() —
One continuous + one categorical geom_boxplot(), geom_violin() pie charts
Two categorical geom_count(), geom_tile() —
Change over time geom_line() geom_point() alone

One variable: distributions

flights_sample |>
  ggplot(aes(x = dep_delay)) +
  geom_histogram(
    binwidth = 10,
    fill = "steelblue",
    color = "white") +
  labs(
    title = "Distribution of Departure Delays",
    x = "Departure delay (min)"
  ) +
  theme_minimal(base_size = 14)

One variable: distributions

Histogram of departure delays in minutes with 10-minute bins, showing a right-skewed distribution: most flights depart close to on time, with a long tail of delayed flights.

histogram vs density

Side-by-side comparison of a histogram and density plot of departure delays. The histogram shows discrete count bins while the density plot shows a smooth continuous curve, both revealing the same right-skewed shape.

Histogram = actual counts. Density = smoothed estimate of the shape.

One categorical variable

flights_sample |>
  ggplot(aes(x = origin)) +
  geom_bar(fill = "steelblue") +
  labs(
    title = "Flights per NYC Airport",
    x = "Origin airport",
    y = "Count"
  ) +
  theme_minimal(base_size = 14)

One categorical variable

Bar chart showing the count of sampled flights by NYC airport of origin. EWR, JFK, and LGA each have a few hundred flights in the sample.

One continuous + one categorical

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay, fill = origin)) +
  geom_boxplot(alpha = 0.7) +
  labs(
    title = "Departure Delay by Airport",
    x = "Origin airport",
    y = "Departure delay (min)"
  ) +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

One continuous + one categorical

Boxplots comparing departure delay distributions across the three NYC airports. Medians are close to zero at all three, with long upper whiskers reflecting the right-skewed tail.

boxplot vs violin

Side-by-side comparison of boxplot and violin plot for departure delay by airport. The boxplot shows median, quartiles, and whiskers; the violin reveals the full right-skewed shape, with most density concentrated near zero and a long tail.

Violin shows the full distribution shape. Boxplot shows summary stats. Both have their place.

Showing individual points: geom_jitter()

When your x-axis is categorical, geom_point() would stack every observation on the same vertical line. geom_jitter() adds a small random horizontal wiggle so each point is visible.

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay, color = origin)) +
  geom_jitter(width = 0.2, alpha = 0.4) +
  labs(
    title = "Every Flight, Plotted",
    x = "Origin airport",
    y = "Departure delay (min)"
  ) +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

Showing individual points: geom_jitter()

Jittered dot plot of departure delays by NYC airport. Each point is a single flight, with horizontal jitter so overlapping points are visible. All three airports show a dense cluster near zero and a sparse tail of long delays.

stat_summary()

The psychology staple: means + error bars

In psych papers, you’ll see this constantly: a bar chart showing group means with error bars. stat_summary() is how you make it.

stat_summary() basics

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay)) +
  stat_summary(
    fun = mean,
    geom = "bar",
    fill = "steelblue", width = 0.5) +
  stat_summary(
    fun.data = mean_cl_normal,
    geom = "errorbar",
    width = 0.3) +
  labs(
    title = "Mean Departure Delay by Airport",
    x = "Origin airport",
    y = "Departure delay (min)"
  ) +
  theme_minimal(base_size = 14)

stat_summary() basics

Bar chart showing mean departure delay by airport with 95% confidence interval error bars. Means are close to each other across the three airports, with overlapping error bars.

Breaking down stat_summary()

# The bar (mean)
stat_summary(fun = mean, geom = "bar")

# The error bars (95% confidence interval)
stat_summary(fun.data = mean_cl_normal, geom = "errorbar")
  • fun = the function to calculate (mean, median, etc.)
  • fun.data = a function that returns ymin, y, ymax (like mean_cl_normal)
  • geom = what shape to use to show the result

What are error bars showing?

This matters! Always state it in your figure caption.

Error bar type What it means How to get it
SE (standard error) Precision of the mean mean_se
SD (standard deviation) Spread of the data Write your own
95% CI Confidence interval mean_cl_normal
# SE
stat_summary(fun.data = mean_se, geom = "errorbar")

# 95% CI (most common in psych)
stat_summary(fun.data = mean_cl_normal, geom = "errorbar")

Why does this matter?

The APA Publication Manual (7th ed.) recommends showing individual data points alongside summary statistics whenever possible.

Bar charts with error bars are ubiquitous in psychology — but they hide the shape of the data. Outliers, skew, and bimodality all disappear behind a rectangle.

Adding individual data points

Error bars alone hide the data. Show the points too:

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay, color = origin)) +
  geom_jitter(width = 0.15, alpha = 0.4, size = 2) +
  stat_summary(fun = mean, geom = "point", size = 4, color = "black") +
  stat_summary(fun.data = mean_cl_normal, geom = "errorbar", width = 0.2, color = "black") +
  labs(
    title = "Departure Delay by Airport",
    subtitle = "Error bars = 95% CI. Individual flights shown.",
    x = "Origin airport",
    y = "Departure delay (min)"
  ) +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

Adding individual data points

Dot plot showing individual flight delays as jittered points for each airport, with black dots marking airport means and error bars showing 95% confidence intervals. Individual variation is clearly visible behind the summary statistics.

Pair coding break

Your turn: 10 minutes

Using the flights dataset (filter to !is.na(dep_delay) first):

  1. Create a bar chart with error bars showing mean dep_delay by origin
  2. Add individual data points (jittered) behind the bars
  3. Color the bars by origin
  4. Add a caption noting that error bars show the 95% CI

Tip

You’ll need stat_summary() twice — once for the bar, once for the error bars. Look at the examples from the last few slides.

Solution: Mean departure delay by airport

flights |>
  filter(!is.na(dep_delay)) |>
  ggplot(aes(x = origin, y = dep_delay, fill = origin)) +
  geom_jitter(width = 0.15, alpha = 0.2) +
  stat_summary(fun = mean, geom = "bar", alpha = 0.6, width = 0.5) +
  stat_summary(fun.data = mean_cl_normal, geom = "errorbar", width = 0.2) +
  labs(
    title = "Mean Departure Delay by NYC Airport",
    x = "Origin airport",
    y = "Departure delay (min)",
    caption = "Error bars represent 95% confidence intervals."
  ) +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

Solution: output

Bar chart of mean departure delay by NYC airport (EWR, JFK, LGA) with 95% confidence interval error bars and jittered individual flight points behind the bars. The three airports have similar means with overlapping confidence intervals.

Position & scales

Position adjustments

When geoms overlap, position adjustments fix it:

Position What it does When to use
"dodge" Side by side Grouped bar charts
"stack" Stacked on top Stacked bars
"fill" Stacked to 100% Comparing proportions
"jitter" Random wiggle Overplotted points

dodge: grouped bar charts

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay, fill = time_of_day)) +
  stat_summary(fun = mean, geom = "bar", position = "dodge", width = 0.6) +
  stat_summary(fun.data = mean_cl_normal, geom = "errorbar",
               position = position_dodge(0.6), width = 0.2) +
  labs(
    title = "Departure Delay by Airport and Time of Day",
    x = "Origin airport",
    y = "Mean delay (min)",
    fill = "Time of day"
  ) +
  theme_minimal(base_size = 14)

dodge: grouped bar charts

Grouped bar chart showing mean departure delay by airport and time of day with dodged bars and error bars. Morning and Afternoon/Evening bars sit side by side within each airport, showing afternoon delays are consistently higher.

stack and fill

Side-by-side comparison of stacked and filled bar charts. The stacked chart shows raw counts of Morning and Afternoon/Evening flights at each airport; the filled chart normalizes each bar to 100% to compare proportions directly.

Scales: controlling axes and colors

Scales translate data values into visual properties:

# Control axis breaks (tick marks)
scale_x_continuous(breaks = seq(0, 180, by = 30))

# Set specific colors
scale_fill_manual(values = c("EWR" = "lightblue", "JFK" = "coral", "LGA" = "lightgreen"))

scale_fill_manual()

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay, fill = origin)) +
  geom_boxplot(alpha = 0.7) +
  scale_fill_manual(values = c("EWR" = "#3498db", "JFK" = "#e74c3c", "LGA" = "#f1c40f")) +
  labs(title = "Custom colors", x = "Origin airport", y = "Delay (min)") +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

Use a website like HTML color codes to find the appropriate HEX code for your color.

scale_fill_manual()

Boxplots of departure delay by airport using custom colors: blue for EWR, red for JFK, and yellow for LGA, demonstrating how scale_fill_manual() applies specific hex color codes to groups.

Coordinate systems

# Flip x and y (great for long category labels)
coord_flip()

# Zoom in without dropping data
coord_cartesian(ylim = c(-10, 30))
# vs
scale_y_continuous(limits = c(-10, 30))  # This DROPS data outside the range!

coord_flip() in action

flights_sample |>
  ggplot(aes(y = origin, x = dep_delay, fill = origin)) +
  geom_boxplot(alpha = 0.7) +
  labs(title = "Flipped axes — better for category labels", y = "Origin airport", x = "Delay (min)") +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

coord_flip() in action

Horizontal boxplots of departure delay by airport, with airport codes on the y-axis and delay on the x-axis, demonstrating how flipped axes improve readability for categorical labels.

Putting it together

A polished figure

flights_sample |>
  ggplot(aes(x = origin, y = dep_delay, fill = origin, color = origin)) +
  geom_jitter(width = 0.2, alpha = 0.35, size = 1.5) +
  # outlier.shape = NA hides boxplot outliers because geom_jitter already shows every point
  geom_boxplot(alpha = 0.5, width = 0.4, outlier.shape = NA) +
  # shape = 18 is a solid diamond — draws attention to the mean
  stat_summary(fun = mean, geom = "point", shape = 18, size = 5, color = "black") +
  scale_fill_manual(values = c("EWR" = "#3498db", "JFK" = "#e74c3c", "LGA" = "#f1c40f")) +
  scale_color_manual(values = c("EWR" = "#2980b9", "JFK" = "#c0392b", "LGA" = "#b7950b")) +
  labs(
    title = "Departure Delays Are Similar Across NYC Airports",
    subtitle = "Individual flights, box summaries, and mean (diamond) shown",
    x = "Origin airport",
    y = "Departure delay (min)",
    caption = "N = 500 sampled flights. Diamond = mean."
  ) +
  theme_minimal(base_size = 14) +
  theme(legend.position = "none")

A polished figure

Polished figure combining jittered individual flight delays, boxplots, and diamond-shaped mean markers for each NYC airport. Uses custom colors, informative title and subtitle, and a caption noting sample size and mean indicator.

Get a head start

Assignment 4 preview

Assignment 4 will ask you to create “bad” and “good” versions of figures. Start experimenting:

  1. Take any plot from today and make it deliberately bad — wrong geom, missing labels, confusing colors
  2. Then fix it. What did you change and why?
  3. Try creating the same data with three different geoms. Which one communicates best?

Wrapping up

Today’s toolkit

Tool What it does
geom_histogram() Distribution of one continuous variable
geom_density() Smooth distribution estimate
geom_boxplot() Summary of continuous by categorical
geom_violin() Full distribution by categorical
geom_jitter() Individual points with horizontal wiggle
stat_summary() Calculate and display summary stats
position = "dodge" Side-by-side grouped plots
scale_fill_manual() Custom colors
coord_flip() Swap axes

Before next class

📖 Read:

  • Supplementary: Visual perception principles (will be posted)
  • Optional: Knaflic, Storytelling with Data, Ch 1–3

✅ Practice:

  • Create a bar chart with error bars for a dataset of your choice
  • Try boxplot vs violin on the same data
  • Experiment with coord_flip() and scale_fill_manual()

Key takeaways

  1. Match your geom to your data — the wrong choice misleads
  2. stat_summary() is powerful — means, error bars, custom functions
  3. Position adjustments handle overlap (dodge, stack, fill, jitter)
  4. Scales control axes, colors, and legends
  5. coord_cartesian() zooms; scale limits drop data — know the difference

The one thing to remember

The difference between a chart and a figure is intention — every layer should earn its place.

Next time: Perception & Design