Putting It All Together

PSY 410: Data Science for Psychology

Dr. Sara Weston

2026-06-03

Looking back

The data science workflow

R4DS data science workflow diagram showing the cycle: Import, Tidy, Transform, Visualize, Model, Communicate

Ten weeks ago, you couldn’t do any of this. Now you can do all of it.

Live demonstration

A real analysis: Start to finish

Let’s analyze a real dataset together — the General Social Survey, included in forcats as gss_cat.

Research question: Does age predict daily TV viewing among Americans — and does it differ by political party?

I’ll demonstrate the full workflow:

  1. Import and explore a real dataset
  2. Clean and recode categories
  3. Look at distributions (univariate & by group)
  4. Quantify the relationship (correlation + regression)
  5. Reshape with pivots and joins
  6. Build a Quarto report

The dataset

gss_cat is a sample of the General Social Survey (NORC at the University of Chicago), one of the longest-running social surveys in the US.

gss_cat
# A tibble: 21,483 × 9
    year marital         age race  rincome        partyid    relig denom tvhours
   <int> <fct>         <int> <fct> <fct>          <fct>      <fct> <fct>   <int>
 1  2000 Never married    26 White $8000 to 9999  Ind,near … Prot… Sout…      12
 2  2000 Divorced         48 White $8000 to 9999  Not str r… Prot… Bapt…      NA
 3  2000 Widowed          67 White Not applicable Independe… Prot… No d…       2
 4  2000 Never married    39 White Not applicable Ind,near … Orth… Not …       4
 5  2000 Divorced         25 White Not applicable Not str d… None  Not …       1
 6  2000 Married          25 White $20000 - 24999 Strong de… Prot… Sout…      NA
 7  2000 Never married    36 White $25000 or more Not str r… Chri… Not …       3
 8  2000 Divorced         44 White $7000 to 7999  Ind,near … Prot… Luth…      NA
 9  2000 Married          44 White $25000 or more Not str d… Prot… Other       0
10  2000 Married          47 White $25000 or more Strong re… Prot… Sout…       3
# ℹ 21,473 more rows

Step 1: Explore the data

glimpse(gss_cat)
Rows: 21,483
Columns: 9
$ year    <int> 2000, 2000, 2000, 2000, 2000, 2000, 2000, 2000, 2000, 2000, 20…
$ marital <fct> Never married, Divorced, Widowed, Never married, Divorced, Mar…
$ age     <int> 26, 48, 67, 39, 25, 25, 36, 44, 44, 47, 53, 52, 52, 51, 52, 40…
$ race    <fct> White, White, White, White, White, White, White, White, White,…
$ rincome <fct> $8000 to 9999, $8000 to 9999, Not applicable, Not applicable, …
$ partyid <fct> "Ind,near rep", "Not str republican", "Independent", "Ind,near…
$ relig   <fct> Protestant, Protestant, Protestant, Orthodox-christian, None, …
$ denom   <fct> "Southern baptist", "Baptist-dk which", "No denomination", "No…
$ tvhours <int> 12, NA, 2, 4, 1, NA, 3, NA, 0, 3, 2, NA, 1, NA, 1, 7, NA, 3, 3…

Step 1 (cont.): Check for missing data

gss_cat |>
  summarize(across(everything(), ~sum(is.na(.x))))
# A tibble: 1 × 9
   year marital   age  race rincome partyid relig denom tvhours
  <int>   <int> <int> <int>   <int>   <int> <int> <int>   <int>
1     0       0    76     0       0       0     0     0   10146

What’s in partyid?

Before we recode, let’s see what’s actually there:

gss_cat |>
  count(partyid)
# A tibble: 10 × 2
   partyid                n
   <fct>              <int>
 1 No answer            154
 2 Don't know             1
 3 Other party          393
 4 Strong republican   2314
 5 Not str republican  3032
 6 Ind,near rep        1791
 7 Independent         4119
 8 Ind,near dem        2499
 9 Not str democrat    3690
10 Strong democrat     3490

Step 2: Clean the data

gss_clean <- gss_cat |>
  drop_na(tvhours, age) |>
  mutate(
    party = case_when(
      partyid %in% c("Strong republican", "Not str republican") ~ "Republican",
      partyid %in% c("Strong democrat", "Not str democrat")     ~ "Democrat",
      partyid %in% c("Ind,near rep", "Independent", "Ind,near dem") ~ "Independent",
      TRUE ~ NA_character_
    ),
    party = factor(party, levels = c("Democrat", "Independent", "Republican"))
  ) |>
  drop_na(party)

Step 3: Descriptive statistics

gss_clean |>
  summarize(
    n = n(),
    mean_age = mean(age),
    mean_tv = mean(tvhours),
    sd_tv = sd(tvhours)
  )
# A tibble: 1 × 4
      n mean_age mean_tv sd_tv
  <int>    <dbl>   <dbl> <dbl>
1 11026     47.2    2.98  2.57

Step 3 (cont.): Group sizes

gss_clean |>
  count(party)
# A tibble: 3 × 2
  party           n
  <fct>       <int>
1 Democrat     3832
2 Independent  4461
3 Republican   2733

Step 4: How much TV do people watch?

Start with univariate EDA — one variable at a time.

ggplot(gss_clean, aes(x = tvhours)) +
  geom_histogram(binwidth = 1, fill = "steelblue", color = "white") +
  labs(
    title = "Most respondents watch 1-3 hours per day",
    subtitle = "But the tail stretches to 24 hours",
    x = "Hours of TV per day",
    y = "Number of respondents"
  ) +
  theme_minimal()

Step 4: How much TV do people watch?

Histogram of daily TV hours peaking at 2 hours per day with a long right tail.

Step 4 (cont.): Does it differ by party?

Now bivariate EDA — compare distributions across groups.

ggplot(gss_clean, aes(x = party, y = tvhours, fill = party)) +
  geom_boxplot(alpha = 0.7) +
  scale_fill_manual(values = c(
    "Democrat"    = "#3498db",
    "Independent" = "#7f8c8d",
    "Republican"  = "#e74c3c"
  )) +
  labs(
    title = "TV viewing distributions look similar across parties",
    x = NULL,
    y = "Hours of TV per day"
  ) +
  theme_minimal() +
  theme(legend.position = "none")

Step 4 (cont.): Does it differ by party?

Side-by-side boxplots of daily TV hours for Democrats, Independents, and Republicans showing similar medians and spreads.

Step 5: Initial visualization

ggplot(gss_clean, aes(x = age, y = tvhours)) +
  geom_jitter(alpha = 0.15, height = 0.3, width = 0) +
  geom_smooth(method = "lm", color = "steelblue") +
  labs(
    title = "Older Americans report more daily TV viewing",
    subtitle = "General Social Survey, 2000-2014",
    x = "Age (years)",
    y = "Hours of TV per day"
  ) +
  theme_minimal()

Step 5: Initial visualization

Scatterplot with a linear trend line showing that older Americans report more daily TV viewing in the General Social Survey.

Step 6: Compute correlation

cor_value <- cor(gss_clean$age, gss_clean$tvhours)
cor_value
[1] 0.1424147

Finding: A small but reliable positive correlation (r = 0.14).

Step 7: Run a regression

A correlation gives one number. A regression gives a slope — how much does TV viewing change per year of age?

tv_model <- lm(tvhours ~ age + party, data = gss_clean)

library(broom)
tidy(tv_model)
# A tibble: 4 × 5
  term             estimate std.error statistic   p.value
  <chr>               <dbl>     <dbl>     <dbl>     <dbl>
1 (Intercept)        2.24     0.0795      28.2  5.22e-169
2 age                0.0212   0.00140     15.2  2.17e- 51
3 partyIndependent  -0.262    0.0562      -4.67 2.99e-  6
4 partyRepublican   -0.618    0.0635      -9.73 2.69e- 22

Step 7 (cont.): How well does it fit?

glance(tv_model) |>
  select(r.squared, adj.r.squared, sigma, nobs)
# A tibble: 1 × 4
  r.squared adj.r.squared sigma  nobs
      <dbl>         <dbl> <dbl> <int>
1    0.0286        0.0284  2.54 11026

R² is tiny — age and party explain only a small slice of TV viewing. The story isn’t wrong, it just isn’t the whole story.

Step 8: Does the pattern differ by party?

ggplot(gss_clean, aes(x = age, y = tvhours, color = party)) +
  geom_smooth(method = "lm") +
  scale_color_manual(values = c(
    "Democrat"    = "#3498db",
    "Independent" = "#7f8c8d",
    "Republican"  = "#e74c3c"
  )) +
  labs(
    title = "Same pattern across the political spectrum",
    x = "Age (years)",
    y = "Hours of TV per day",
    color = "Party"
  ) +
  theme_minimal() +
  theme(legend.position = "top")

Step 8: Does the pattern differ by party?

Scatterplot with separate colored trend lines for Democrats, Independents, and Republicans, showing a similar positive age-TV slope for all three groups.

Step 9: TV viewing categories

gss_clean |>
  mutate(
    tv_category = case_when(
      tvhours == 0 ~ "None",
      tvhours <= 2 ~ "Light (1-2 hr)",
      tvhours <= 4 ~ "Moderate (3-4 hr)",
      tvhours >= 5 ~ "Heavy (5+ hr)"
    ),
    tv_category = factor(tv_category,
                         levels = c("None", "Light (1-2 hr)",
                                    "Moderate (3-4 hr)", "Heavy (5+ hr)"))
  ) |>
  count(tv_category) |>
  ggplot(aes(x = n, y = fct_rev(tv_category), fill = tv_category)) +
  geom_col() +
  scale_fill_manual(values = c(
    "None"              = "#2ecc71",
    "Light (1-2 hr)"    = "#f39c12",
    "Moderate (3-4 hr)" = "#e67e22",
    "Heavy (5+ hr)"     = "#e74c3c"
  )) +
  labs(
    title = "Most Americans watch 1-4 hours of TV per day",
    x = "Number of respondents",
    y = NULL
  ) +
  theme_minimal() +
  theme(legend.position = "none")

Step 9: TV viewing categories

Horizontal bar chart of TV viewing categories showing most Americans watch 1-4 hours per day, with smaller groups at None and Heavy.

Step 10: Has TV viewing changed over time?

gss_clean |>
  group_by(year) |>
  summarize(mean_tv = mean(tvhours), .groups = "drop") |>
  ggplot(aes(x = year, y = mean_tv)) +
  geom_line(color = "steelblue", linewidth = 1.2) +
  geom_point(color = "steelblue", size = 3) +
  labs(
    title = "Mean daily TV viewing across GSS years",
    x = "Survey year",
    y = "Mean hours of TV per day"
  ) +
  theme_minimal()

Step 10: Has TV viewing changed over time?

Line plot of mean daily TV hours from 2000 to 2014, drifting slightly downward across the period.

Step 11: Add political context with a join

What if we want to label each year by who was president? We need a second table, then a join (Session 15).

president_lookup <- tribble(
  ~year, ~president,
   2000, "Clinton",
   2002, "G.W. Bush",
   2004, "G.W. Bush",
   2006, "G.W. Bush",
   2008, "G.W. Bush",
   2010, "Obama",
   2012, "Obama",
   2014, "Obama"
)

president_lookup
# A tibble: 8 × 2
   year president
  <dbl> <chr>    
1  2000 Clinton  
2  2002 G.W. Bush
3  2004 G.W. Bush
4  2006 G.W. Bush
5  2008 G.W. Bush
6  2010 Obama    
7  2012 Obama    
8  2014 Obama    

Step 11 (cont.): The join

gss_with_president <- gss_clean |>
  left_join(president_lookup, by = "year")

gss_with_president |>
  group_by(president) |>
  summarize(
    n = n(),
    mean_tv = mean(tvhours),
    .groups = "drop"
  ) |>
  arrange(desc(mean_tv))
# A tibble: 3 × 3
  president     n mean_tv
  <chr>     <int>   <dbl>
1 Obama      4245    3.02
2 Clinton    1799    2.99
3 G.W. Bush  4982    2.95

Tiny differences — but the technique is what counts. Joins let you bring in any external data (codebooks, demographics, neighborhood stats).

Step 12: Summary table by party

summary_table <- gss_clean |>
  group_by(party) |>
  summarize(
    N        = n(),
    Mean_Age = mean(age),
    Mean_TV  = mean(tvhours),
    SD_TV    = sd(tvhours),
    .groups  = "drop"
  )

knitr::kable(summary_table, digits = 1,
             col.names = c("Party", "N", "M Age", "M TV", "SD TV"))
Party N M Age M TV SD TV
Democrat 3832 48.7 3.3 2.8
Independent 4461 44.8 2.9 2.6
Republican 2733 49.2 2.7 2.2

Step 13: Write it up in Quarto

## Results

We analyzed data from `r nrow(gss_clean)` respondents in
the General Social Survey (2000-2014), with a mean age of
`r round(mean(gss_clean$age), 1)` years. Respondents reported
watching `r round(mean(gss_clean$tvhours), 1)` hours of TV per
day on average (SD = `r round(sd(gss_clean$tvhours), 1)`).

Age was positively correlated with daily TV viewing
(r = `r round(cor_value, 2)`), such that older respondents
tended to report more hours of TV.

Step 13 (cont.): What it renders to

Results

We analyzed data from 11026 respondents in the General Social Survey (2000-2014), with a mean age of 47.2 years. Respondents reported watching 3 hours of TV per day on average (SD = 2.6).

Age was positively correlated with daily TV viewing (r = 0.14), such that older respondents tended to report more hours of TV. The relationship was similar across Democrats, Independents, and Republicans.

Small scatterplot of age and daily TV hours with a steelblue linear trend line.

What this demonstrates

  • ✅ Explore & clean — glimpse, count, case_when, drop_na, factor
  • ✅ Univariate & bivariate EDA — geom_histogram, geom_boxplot
  • ✅ Quantify — cor, lm, broom::tidy, broom::glance
  • ✅ Group & reshape — group_by, summarize, tribble, left_join, arrange
  • ✅ Visualize — scatter, line, bar, box; color, scale, theme, labs
  • ✅ Communicate — kable tables + inline R in Quarto

One pipeline. Sixteen sessions of skills.

When things break

The 6 errors you’ll see most often (1/2)

Error What it means Fix
object 'x' not found Typo, wrong capitalization, or object not created yet Check spelling; run the code that creates it
could not find function Typo in function name or package not loaded Check spelling; run library()
unexpected symbol Missing |>, +, comma, or parenthesis Check the line before the error

The 6 errors you’ll see most often (2/2)

Error What it means Fix
non-numeric argument Math on text — variable is character, not numeric Check type with class() or glimpse()
did you mean '=='? Used = (assignment) instead of == (comparison) in filter() Change = to ==
ggplot + error Missing + between ggplot layers Every line except the last needs +

A 5-step debugging strategy

  1. Read the error message — actually read it. It tells you where and what.
  1. Check the basics — package loaded? Object created? Spelling correct?
  1. Run line by line — pipe chains: run from the top, adding one line at a time. Which line breaks?
  1. Simplify — make a tiny test dataset (tibble(x = 1:3)) and try the same operation.
  1. Google it — include “R”, the package name, and the exact error message.

When you ask for help: make a reprex

Reprex = reproducible example — the smallest code that recreates your error.

Bad: “My code doesn’t work. Help!”

Good:

“I’m trying to filter my data but getting ‘object not found’:

library(tidyverse)
data <- tibble(x = 1:3, y = c("a", "b", "c"))
filter(data, x > 1)
# Error: object 'data' not found

I expected rows where x > 1.”

Tip: Use dput() to share a small slice of your real data so others can recreate it exactly.

Where to get help after this course

Before posting: search first, be specific, show what you tried, and include a reprex.

Pair coding break

Your turn: Debug this code

Tip

There are at least 4 bugs. Read each line carefully and check: names, operators, punctuation.

library(tidyverse)

Study_data <- tibble(
  id = 1:5,
  Score = c(10, 15, 12, 18, 14),
  Group = c("A", "B", "A", "B", "A")
)

study_data |>
  filter(Group = "A") |>
  summarize(
    mean_score = mean(score)
    sd_score = sd(score)
  )

Time: 10 minutes

Before we move on

📤 Submit your code on Canvas for participation credit. Paste what you have — it doesn’t need to work perfectly.

Where to go from here

You have the foundation

This course covered data wrangling and visualization — the foundation of data science.

What’s next?

  • Statistics in R — inference, hypothesis testing
  • Writing functions — automate repetitive work
  • Version control — Git and GitHub
  • Interactive tools — Shiny dashboards
  • The wider R ecosystem — tidymodels, tidytext, and more

Statistics in R

You can now learn inferential statistics:

# t-test (compare two groups)
t.test(tvhours ~ I(party == "Republican"), data = gss_clean)

# ANOVA (compare multiple groups)
aov(tvhours ~ party, data = gss_clean)

# Regression
lm(tvhours ~ age + party, data = gss_clean)

Courses to consider:

Writing functions

Automate repetitive tasks:

# Instead of copying this code multiple times...
compute_scale_score <- function(data, items) {
  data |>
    rowwise() |>
    mutate(score = mean(c_across(all_of(items)), na.rm = TRUE)) |>
    ungroup()
}

# Use it anywhere
phq9_scored <- compute_scale_score(my_data, c("phq1", "phq2", "phq3"))

Version control with Git

Track changes to your code over time:

  • Never lose work — full history of changes
  • Collaborate easily — merge changes from multiple people
  • Professional standard — expected in industry and academia

Interactive visualizations with Shiny

Build web apps for your data — no HTML or JavaScript required:

library(shiny)

ui <- fluidPage(
  selectInput("party", "Choose party:",
              choices = c("Democrat", "Independent", "Republican")),
  plotOutput("plot")
)

server <- function(input, output) {
  output$plot <- renderPlot({
    gss_clean |>
      filter(party == input$party) |>
      ggplot(aes(age, tvhours)) + geom_point()
  })
}

shinyApp(ui, server)

See Shiny in action

Posit’s classic Movie Explorer — try the sliders and dropdowns.

Go deeper into the R ecosystem

The R world is huge. A few doorways worth knowing:

Modeling:

  • tidymodels — predictive modeling with tidy syntax
  • brms — Bayesian regression
  • lme4 — mixed-effects models

Text & language:

  • tidytext — turn text into tidy data (you’ve seen this)
  • ellmer — call LLMs (Claude, GPT) from R

Go deeper (cont.): communication & workflow

Communication:

  • gtsummary — publication-ready summary tables
  • quarto — websites, books, dashboards (you’ve used the basics)

Workflow:

  • targets — reproducible analysis pipelines
  • renv — project-level package management

Learning resources

Free online resources

Books:

Interactive learning:

Community resources

Weekly challenges:

Communities:

Keep practicing

The only way to maintain skills: use them

Ideas:

  • Analyze data from your own research
  • Replicate figures from published papers
  • Join #TidyTuesday
  • Help friends/labmates with their data
  • Create a personal website with Quarto
  • Build a data visualization portfolio

Tip

Aim for 1 hour per week — consistency matters more than intensity

Final reflections

What makes a good data scientist?

It’s not about knowing every function or memorizing syntax.

Good data scientists:

  1. Ask good questions — what story is the data telling?
  2. Stay curious — always learning new tools and techniques
  3. Communicate clearly — make complex findings accessible
  4. Work reproducibly — others can verify and build on your work
  5. Think critically — question assumptions, check for bias
  6. Persist through errors — debugging is part of the job

You’ve developed all of these skills this quarter.

The growth mindset

Ten weeks ago, many of you had never written a line of code.

Now you can:

  • Import and clean messy data
  • Create publication-quality visualizations
  • Wrangle complex datasets with joins and pivots
  • Handle missing data appropriately
  • Build reproducible reports

That’s incredible growth.

Errors are part of the process

Remember:

  • Everyone gets errors — even experienced programmers
  • Errors mean you’re learning
  • Each error you solve makes you better
  • The frustration is temporary; the skills are permanent

Important

If you take one thing from this course: You can learn hard things.

Thank you

Thank you for:

  • Showing up and participating
  • Helping each other
  • Asking questions
  • Persisting through challenges
  • Trusting the process

You’ve been a great class. I’m excited to see what you do with these skills. 🎉