R
The statistical programming language used by data scientists, researchers, and analysts for statistical modeling, data visualization, and academic research.
R is a programming language purpose-built for statistical computing and data visualization. It is the dominant language in academic research, clinical trials, social science, epidemiology, and quantitative finance — fields where statistical rigor matters as much as engineering. The R ecosystem is anchored by the tidyverse (dplyr, ggplot2, tidyr) for data manipulation and visualization, and R Markdown for reproducible research reports. While Python has overtaken R in industry data science roles, R remains the required language in many research, biostatistics, and econometrics positions, and the two languages frequently coexist in data science teams.
Typical time to job-readiness: ~3 months.
Learning R
Beginner
Learn R through R for Data Science (Hadley Wickham's free online book) — it teaches the tidyverse approach (dplyr for data manipulation, ggplot2 for visualization) which is the modern standard. Get comfortable with RStudio as your IDE and R Markdown for combining code and narrative.
Intermediate
Statistical modeling with lm(), glm(), and survival analysis; advanced ggplot2 customization; data cleaning with tidyr; and working with dates, strings, and list-columns in tibbles. Learn how to write reproducible analyses and share them via R Markdown or Quarto.
Advanced
Shiny for interactive web applications, Rcpp for performance-critical code, parallel computing with furrr, and packaging your code as an R package. Research roles assess R fluency through take-home analyses — prepare to explain your statistical choices, not just your code.
Key concepts
- Tidyverse: dplyr for data manipulation, ggplot2 for visualization, tidyr for reshaping — the modern R standard
- Vectors, data frames, and tibbles — R's primary data structures; everything is a vector at its base
- Pipe operator (|> or magrittr's %>%): chains operations without intermediate variables
- R Markdown and Quarto: mix code, output, and narrative in a single reproducible document
- CRAN ecosystem: 20,000+ packages for statistics, visualization, and domain-specific methods
- Statistical output: p-values, confidence intervals, and effect sizes must be interpreted, not just reported
Common interview topics
- Walk me through how you would clean and analyze a messy dataset in R
- How do you use dplyr to filter, group, and summarize a data frame
- What is ggplot2 and how do you use the grammar of graphics to build a chart
- Explain the difference between a linear model and a generalized linear model
- How do you make your R analysis reproducible for someone else to run