
PSY 410: Data Science for Psychology
2026-05-27
You’ve learned to:
But technical skills ≠ communication skills

Stories are memorable:
Stories are persuasive:
Decisions are emotional:
You have one finding. You can present it to:
Is the figure the same for all of them?
No. The data is. The figure isn’t.
A simulated dataset of 200 college students:
sleep_hours — average nightly sleepgpa — cumulative GPAyear, study_hours, caffeine_cupsThe finding:
Each extra hour of nightly sleep is associated with about +0.21 higher GPA. Students who sleep 7+ hours average 3.39; students who sleep less than 6 average 2.88.
Paste this in your console to get the dataset.
Note
The data are simulated — 200 fake students with a known sleep–GPA relationship built in. Code that made them: scripts/generate-sleep-study.R.
You’re mid-analysis. You need to know what’s there.
A reader of a published paper.
annotate(..., parse = TRUE)annotate() adds text without needing a data frameparse = TRUE reads the label as an R expression — math notation, Greek letters== (not =) for equality in plotmathggplot(sleep_study, aes(x = sleep_hours, y = gpa)) +
geom_point(alpha = 0.5, color = "gray40") +
geom_smooth(method = "lm", color = "black", fill = "gray80") +
annotate("text", x = 4.6, y = 3.95,
label = "italic(r) == 0.56", parse = TRUE, hjust = 0, size = 4.2) +
annotate("text", x = 4.6, y = 3.80,
label = "beta == 0.21", parse = TRUE, hjust = 0, size = 4.2) +
annotate("text", x = 4.6, y = 3.65,
label = "italic(N) == 200", parse = TRUE, hjust = 0, size = 4.2) +
labs(
title = "Figure 1",
subtitle = "Cumulative GPA by average nightly sleep duration in college students",
x = "Average nightly sleep (hours)",
y = "Cumulative GPA",
caption = "Note. Shaded region represents 95% CI. N = 200."
) +
scale_x_continuous(breaks = 4:10) +
coord_cartesian(ylim = c(1.5, 4)) +
theme_classic(base_size = 12) +
theme(plot.title = element_text(face = "bold"),
plot.subtitle = element_text(face = "italic"))An academic advisor making a flyer or running a wellness workshop.
geom_text() direct labels + highlight + assertion titlegeom_text() puts one label per row of dataaes(label = ...) says which column to displaysprintf("%.2f", x) formats the number (2 decimal places)vjust = -0.6 nudges the label above the barWhy direct labels? The reader shouldn’t need to look at the axis to know the value.
sprintf()sprintf() formats numbers (and other values) into strings.
sprintf() — reading the format string| Piece | What it does |
|---|---|
% |
“format a value here” |
.2 |
two decimal places |
f |
floating-point number (decimal) |
d |
integer (no decimal) |
%% |
a literal % sign |
Example: "%.1f%%" → format the value with 1 decimal place, then add a literal %. Applied to 67.5 → "67.5%".
vjustvjust controls how the label sits relative to its anchor point.
vjust value |
What happens |
|---|---|
1 |
label’s top sits at the data point (label extends down) |
0.5 |
label is centered on the data point |
0 |
label’s bottom sits at the data point (label extends up) |
-0.6 |
label floats above the data point with a gap |
The trick: vjust is “how much of the label is below the anchor.” Negative values lift the label up off the data point.
vjust in actiongroup_means <- sleep_study |>
group_by(sleep_group) |>
summarize(mean_gpa = mean(gpa), .groups = "drop") |>
mutate(highlight = sleep_group == "Short (<6)")
ggplot(group_means, aes(x = sleep_group, y = mean_gpa, fill = highlight)) +
geom_col(width = 0.65) +
geom_text(aes(label = sprintf("%.2f", mean_gpa)),
vjust = -0.6, size = 5.5, fontface = "bold", color = "gray20") +
scale_fill_manual(values = c(`TRUE` = "#c0392b", `FALSE` = "gray70")) +
scale_y_continuous(limits = c(0, 4), breaks = 0:4,
expand = expansion(mult = c(0, 0.05))) +
labs(
title = "Short sleepers earn nearly half a letter grade less",
subtitle = "Mean cumulative GPA by average nightly sleep (N = 200)",
x = NULL, y = "Mean GPA"
) +
theme_classic(base_size = 14) +
theme(legend.position = "none",
plot.title = element_text(face = "bold", size = 16))Someone scrolling a feed for two seconds.
ggplot() with no data — the canvas is emptyannotate() places elements at chosen coordinatestheme_void() removes axes and gridlinesplot.background paints the canvasA figure doesn’t have to contain a “chart.”
ggplot() +
annotate("text", x = 0.5, y = 0.78,
label = "+0.5", size = 38, fontface = "bold",
color = "white", hjust = 0.5) +
annotate("text", x = 0.5, y = 0.60,
label = "GPA points", size = 9,
color = "white", hjust = 0.5) +
annotate("text", x = 0.5, y = 0.40,
label = "Students who sleep 7+ hours earn\nhigher grades than those who sleep < 6",
size = 5.5, color = "white", hjust = 0.5, lineheight = 1.2) +
annotate("segment", x = 0.25, xend = 0.75, y = 0.20, yend = 0.20,
color = "#2ecc71", linewidth = 1.2) +
annotate("text", x = 0.5, y = 0.13,
label = "PSY 410 · N = 200", size = 3.8,
color = "gray70", hjust = 0.5, fontface = "italic") +
xlim(0, 1) + ylim(0, 1) +
coord_fixed() +
theme_void() +
theme(
plot.background = element_rect(fill = "#1a2942", color = NA),
panel.background = element_rect(fill = "#1a2942", color = NA),
plot.margin = margin(20, 20, 20, 20)
)| Framing | What it actually means |
|---|---|
| “β = 0.21” | The journal version |
| “0.5 GPA gap” | The clinician version |
| “Higher grades” | The social card |
All true. Each chooses how much detail to drop.
Drop too much and you cross into misleading. We’ll come back to this.
A roomful of people during a 30-second slide in your talk.
geom_vline(), annotate("rect"), annotate("label")annotate("rect", ...) draws a shaded region (use -Inf/Inf to span the panel)geom_vline() adds a reference line at a valueannotate("label", ...) is annotate("text") with a backgroundggplot(sleep_study, aes(x = sleep_hours, y = gpa)) +
annotate("rect", xmin = 7, xmax = 9.5, ymin = -Inf, ymax = Inf,
fill = "#2ecc71", alpha = 0.10) +
geom_point(alpha = 0.4, color = "gray50", size = 2) +
geom_smooth(method = "lm", color = "#2c3e50", linewidth = 1.4, se = FALSE) +
geom_vline(xintercept = 7, linetype = "dashed",
color = "#2ecc71", linewidth = 0.8) +
annotate("label", x = 8.25, y = 2.1,
label = "Recommended\n7+ hours",
color = "#27ae60", fontface = "bold", size = 5,
label.size = 0, fill = "white", lineheight = 1) +
annotate("label", x = 5.4, y = 3.85,
label = "Every extra hour\n= +0.2 GPA",
color = "#2c3e50", fontface = "bold", size = 5.5,
label.size = 0, fill = "#ecf0f1", lineheight = 1) +
labs(title = "More sleep, higher grades",
x = "Nightly sleep (hours)", y = "Cumulative GPA") +
scale_x_continuous(breaks = 4:10) +
coord_cartesian(ylim = c(1.5, 4)) +
theme_minimal(base_size = 18) +
theme(plot.title = element_text(face = "bold", size = 26))| Audience | Annotation skill |
|---|---|
| You | (restraint) |
| Researchers | annotate(..., parse = TRUE) for stats |
| Clinicians | geom_text() direct labels + highlight |
| Public | Text-as-data with theme_void() |
| Live audience | geom_vline() + annotate("rect") + annotate("label") |
Five figures. One finding. One R skill toolkit.
| Tool | When to use |
|---|---|
annotate("text", ...) |
One-off text at coordinates (use parse = TRUE for math) |
annotate("label", ...) |
Same, with a background box |
annotate("rect", ...) |
Shaded region (e.g., a “recommended zone”) |
annotate("segment", ...) |
Line or arrow between two points |
geom_text(aes(label = ...)) |
Label every row of data (use sprintf() to format) |
geom_vline() / geom_hline() |
Vertical / horizontal reference line |
geom_abline() |
Reference line with a slope (e.g., y = x) |
ggrepel::geom_text_repel() |
Labels that auto-arrange to avoid overlap |
Pick one figure from your final project draft. (If you don’t have one, use any figure from a recent assignment.)
Produce two versions:
Use at least one annotation technique from today on each version.
Time: 12 minutes
📤 Submit your code on Canvas for participation credit. Paste what you have — both versions don’t need to be polished.
Every annotation we just learned can also mislead:
The skill is knowing when you’ve crossed the line.
Same data. The left chart will get reposted; the right one won’t.
✅ The baseline is non-zero (body temperature, blood pressure)
✅ You’re showing change over time (line plot)
✅ You explicitly note it in the caption
❌ Never truncate bar charts. Bars encode magnitude — the bar must start at zero.
One figure shows a finding.
A sequence makes an argument:
Your final project is a sequence. Your video presentation is a sequence.
Build the reader up to your finding.
Order: sample → relationship → effect size. Each figure earns the next.
Lead with the punchline.
Order: headline → action. The headline figure carries the evidence; the action slide says what to do about it.

Florence Nightingale (1858)
Showed Parliament that more soldiers died of preventable disease than battle wounds. Changed military medicine.
The blue wedges (disease) dwarf the red (wounds). The visual carries the argument.
Charles Joseph Minard’s “Carte Figurative” (1869) — Napoleon’s 1812 Russian campaign. Six variables on one chart: army size (band width), location (geography), direction (color), distance, temperature (bottom strip), time.
Hans Rosling — 200 Countries in 200 Years (4 min video)
A bubble chart of life expectancy vs. income, animated across two centuries.
He made global development legible to a general audience by adding time as a fourth dimension (animation).
Your final video presentation is a sequence:
Every figure earns the next. Cut anything that doesn’t.
For each:


Before finalizing any figure, ask:
You’ll receive a handout on APA figure formatting.
Note
The handout covers the formal rules. Class time is better spent on clarity than on margin sizes.
Start the reproducible report due next week. Pick one dataset: palmerpenguins::penguins, psych::bfi, a nycflights13::flights subset, or your final project data.
Build a minimal .qmd with:
author, date, toc: true, code-fold: trueglimpse() of your dataNote
A8 is not your final project. It’s a short, focused reproducible report. Your final project is due June 10 and will be much more substantial.
📖 Read:
✅ Do:
Next session (Correlation & Regression) we’ll reveal Fun Challenge 10: The Final Prediction.
A quick one — predict the correlation from a scatterplot. The deadline is Tuesday at 11:59 PM, so you’ll have Monday in class to work on it with your team.
annotate(), geom_text(), geom_vline(), reference lines, shaded regions.A figure is an argument made for a reader.
Ask “who’s reading?” before you ask “what’s the chart?”
See you Monday for correlation and simple regression!
PSY 410 | Session 16