| id | major |
|---|---|
| 1 | Psychology |
| 2 | psychology |
| 3 | Psych |
| 4 | PSYCHOLOGY |
| 5 | Biology |
| 6 | biology |
| 7 | Bio |
PSY 410: Data Science for Psychology
2026-05-11
The top 2 teams on the scoreboard right now get a one-time strategic choice:
Option A — Boost: Add +3 points to your own team’s total
Option B — Sabotage: Remove 2 points from each of two other teams of your choice
Choices are final. First place picks first.
| id | major |
|---|---|
| 1 | Psychology |
| 2 | psychology |
| 3 | Psych |
| 4 | PSYCHOLOGY |
| 5 | Biology |
| 6 | biology |
| 7 | Bio |
R says 7. You meant 2.
The problem is inconsistent text — different cases, abbreviations, trailing spaces. Today we learn to fix that.
Strings are text data — anything in quotes:
The stringr package (part of tidyverse) gives you tools for working with strings. All stringr functions start with str_.
[1] "JaneDoe"
[1] "Jane Doe"
messy_data <- c("PSYCHOLOGY", "Biology", "psychology", "BIOLOGY", "Psychology")
str_to_lower(messy_data)[1] "psychology" "biology" "psychology" "biology" "psychology"
[1] "PSYCHOLOGY" "BIOLOGY" "PSYCHOLOGY" "BIOLOGY" "PSYCHOLOGY"
[1] "Psychology" "Biology" "Psychology" "Biology" "Psychology"
Essential for cleaning survey data!
Survey data often has extra spaces:
survey <- tibble(
major = c(" Psychology", "BIOLOGY ", "psychology", "Biology", " sociology")
)
survey |>
mutate(
major_clean = str_to_lower(str_trim(major))
)# A tibble: 5 × 2
major major_clean
<chr> <chr>
1 " Psychology" psychology
2 "BIOLOGY " biology
3 "psychology" psychology
4 "Biology" biology
5 " sociology" sociology
str_detect() checks if a pattern is present:
filter()str_detect() returns TRUE/FALSE — exactly what filter() needs:
Patterns are case-sensitive by default:
Open-ended responses often have repeated words:
responses <- c(
"very very tired",
"very stressed and very anxious",
"I felt fine"
)
# Replace the first match in each string
str_replace(responses, "very", "extremely")[1] "extremely very tired" "extremely stressed and very anxious"
[3] "I felt fine"
[1] "extremely extremely tired"
[2] "extremely stressed and extremely anxious"
[3] "I felt fine"
Imagine clinical notes with diagnosis abbreviations scattered inside the text:
Pass a named vector to str_replace_all() to expand them wherever they appear:
str_replace_all(notes, c(
"MDD" = "Major Depressive Disorder",
"GAD" = "Generalized Anxiety Disorder",
"OCD" = "Obsessive-Compulsive Disorder"
))[1] "Patient presented with Major Depressive Disorder symptoms"
[2] "Differential: Major Depressive Disorder vs Generalized Anxiety Disorder"
[3] "Family history of Obsessive-Compulsive Disorder and Generalized Anxiety Disorder"
Regular expressions (regex) are powerful pattern matching tools.
Examples:
\\d matches digits\\s matches whitespace. matches any character+ means “one or more”* means “zero or more”Note
Regex is powerful but complex. For this course, stick to simple patterns. When you need more, check R4DS Ch 14 or regex101.com.
You have messy survey responses:
major to lowercase with no extra spacescomment to title case with no extra spacesis_negative that is TRUE if the comment contains “long” or “confusing” (case-insensitive)Time: 10 minutes
messy_survey |>
mutate(
# clean up major variable
major = str_to_lower(major),
major = str_trim(major),
# clean up comment variable
comment = str_trim(comment),
comment = str_to_title(comment),
# make is_neg column
is_negative = str_detect(comment, "Long") | str_detect(comment, "Confusing"),
is_negative = str_detect(comment, regex("long", ignore_case = T)) |
str_detect(comment, regex("confusing", ignore_case = T))
) |>
filter(is_negative)Factors are R’s way of representing categorical data with a fixed set of possible values.
Notice the Levels line — those are the possible categories.


The forcats package (part of tidyverse) provides functions for working with factors.
All forcats functions start with fct_.
Key functions:
fct_relevel() — manually reorder levelsfct_reorder() — reorder by another variablefct_infreq() — order by frequencyfct_recode() — rename levelsfct_collapse() — combine levels[1] High School Bachelor's Master's High School
Levels: Bachelor's High School Master's
Extremely useful for plots!


Useful for grouping rare categories:
diagnosis <- factor(c("MDD", "GAD", "OCD", "PTSD", "Panic Disorder",
"MDD", "GAD", "Social Anxiety"))
fct_collapse(diagnosis,
Depression = "MDD",
Anxiety = c("GAD", "OCD", "PTSD", "Panic Disorder", "Social Anxiety")
)[1] Depression Anxiety Anxiety Anxiety Anxiety Depression Anxiety
[8] Anxiety
Levels: Anxiety Depression
demo_data |>
mutate(
age_group = factor(age_group, levels = c("18-25", "26-35", "36-45", "46+")),
education = fct_recode(factor(education),
"High School" = "HS",
"Bachelor's" = "BA",
"Master's" = "MA",
"Doctorate" = "PhD"
)
)# A tibble: 6 × 2
age_group education
<fct> <fct>
1 18-25 High School
2 26-35 Bachelor's
3 18-25 Bachelor's
4 36-45 Master's
5 26-35 High School
6 46+ Doctorate
After filtering, factors keep old levels:
Problem 1: Factors created from numbers
Problem 2: Factors behave differently than strings
colors <- factor(c("red", "blue"))
# Can't just add new values
colors[3] <- "green" # This creates NA!
colors[1] red blue <NA>
Levels: blue red
Solution: Convert to character first, or use fct_expand() to add levels.
On the class survey, you described your relationship with data in 1–2 words. Let’s see what you said.
unnest_tokens() splits text into individual words:
# A tibble: 77 × 1
word
<chr>
1 frustrating
2 and
3 overwhelming
4 novel
5 novice
6 hopeful
7 work
8 in
9 progress
10 big
# ℹ 67 more rows
Each row is now a single word.
# A tibble: 5 × 2
word n
<chr> <int>
1 complicated 2
2 curious 2
3 love 2
4 adjacent 1
5 barely 1
Most words appeared just once. Counting alone won’t tell us much.
But every word carries a feeling — positive, negative, neutral. Can we measure that?
tidytext ships with the Bing lexicon — ~6,800 English words tagged positive or negative. We can inner_join() to keep only the words our class used and the lexicon recognized.
words |>
inner_join(get_sentiments("bing"), by = "word") |>
count(sentiment) |>
ggplot(aes(x = sentiment, y = n, fill = sentiment)) +
geom_col() +
scale_fill_manual(values = c("negative" = "#c0392b", "positive" = "#27ae60")) +
labs(
title = "How our class feels about data",
x = NULL, y = "Words"
) +
theme_minimal(base_size = 14) +
theme(legend.position = "none")
# A tibble: 26 × 3
word sentiment n
<chr> <chr> <int>
1 good positive 3
2 complicated negative 2
3 interesting positive 2
4 love positive 2
5 brutal negative 1
6 confused negative 1
7 derogatory negative 1
8 difficult negative 1
9 frustrating negative 1
10 fun positive 1
# ℹ 16 more rows
Read carefully — did the lexicon get every word right?
unnest_tokens(..., drop = FALSE) keeps the source phrase so we can spot-check:
class_survey |>
select(data_words) |>
unnest_tokens(word, data_words, drop = FALSE) |>
inner_join(get_sentiments("bing"), by = "word") |>
filter(word %in% c("good", "great", "love")) |>
mutate(data_words = str_trunc(data_words, 35))# A tibble: 6 × 3
data_words word sentiment
<chr> <chr> <chr>
1 unrequited love love positive
2 not great great positive
3 not good good positive
4 i have never coded before and it... good positive
5 good good positive
6 love-hate love positive
Out of 8 “scored” tokens, the lexicon got one right.
You have messy Likert scale data:
response variable to title caseresponse to a factor with logical orderingfacet_wrap() to make separate panels for each questionfct_infreq() to order responses by overall frequencyopen_responses tibble with a response column, tokenize it and find the most common non-stopwords:likert_data_clean <- likert_data |>
mutate(
# Clean to title case
response = str_to_title(response),
# Convert to factor with logical order
response = factor(response, levels = c(
"Strongly Disagree", "Disagree", "Neutral", "Agree", "Strongly Agree"
))
)
# Basic plot with facets
ggplot(likert_data_clean, aes(x = response)) +
geom_bar() +
facet_wrap(~question) +
labs(title = "Survey responses by question",
x = "Response", y = "Count") +
theme_minimal() +
theme(axis.text.x = element_text(angle = 45, hjust = 1))
# Bonus: order by frequency
likert_data_clean |>
mutate(response = fct_infreq(response)) |>
ggplot(aes(x = response)) +
geom_bar() +
facet_wrap(~question) +
theme_minimal() +
theme(axis.text.x = element_text(angle = 45, hjust = 1))| Function | Purpose |
|---|---|
str_c() |
Combine strings |
str_length() |
Get string length |
str_to_lower(), str_to_upper() |
Change case |
str_trim(), str_squish() |
Remove whitespace |
str_detect() |
Find pattern |
str_replace(), str_replace_all() |
Replace pattern |
str_sub() |
Extract substring |
| Function | Purpose |
|---|---|
factor() |
Create a factor |
fct_relevel() |
Manually reorder levels |
fct_reorder() |
Order by another variable |
fct_infreq() |
Order by frequency |
fct_recode() |
Rename levels |
fct_collapse() |
Combine levels |
fct_drop() |
Remove unused levels |
stringr::str_*() functions to manipulatefct_*()) to manipulate factorsclass() or glimpse()📖 Read:
✅ Do:
Messy categories turn into messy results. str_to_lower() and factor() are your first line of defense.
See you Wednesday for missing data!
PSY 410 | Session 13