2027 Summer R Workshops @ UQ

Introduction to R and Statistical Modelling

Learn the R skills needed to analyse data confidently, build powerful statistical models, create publication-quality figures, and answer complex research questions. Whether you’re working with ecological, environmental, biomedical, social science, or other research datasets, R has become an essential tool for modern data analysis.

These hands-on workshops are designed for both beginners and researchers looking to strengthen their analytical skills. You’ll learn how to wrangle data efficiently, develop reproducible workflows, apply appropriate statistical methods, and avoid common analytical pitfalls. By the end of the workshop, you’ll have practical skills that can be immediately applied to your own research projects. Although most of our example datasets are ecological in nature, the approaches and skills are equally applicable in many other research contexts.

NoteWhy learn to code in the age of AI?

AI can generate code in seconds, but it cannot tell you whether the analysis is appropriate, the assumptions are valid, or the conclusions are scientifically defensible. Learning R gives you the skills to critically evaluate code, identify errors, and adapt analyses to your own research questions. Learning the fundamentals ensures you remain the expert in the AI loop.

This workshop series will run in-person only 8–12 February 2027 at the University of Queensland. You are free to attend as many or as few days as you would like.

The costs are $300 (regular) and $200 (students) per day. Proceeds from these workshops are used to support research at UQ and UniSC.

Day 1: Introduction to R and data wrangling in the tidyverse

In the “Introduction to R and data wrangling in the tidyverse” workshop, we will start slowly. We will help familiarise you with R (the programming language), RStudio (the interface we recommend to program R in) and the tidyverse (a set of programming packages we use to speed up analysis and plotting). By the end of the day, you will be confidently loading datasets and writing basic code in R to manipulate (wrangle) the data.

Difficulty rating: 🤓⚪⚪⚪⚪

Assumed knowledge: No prior programming or R experience required. This is a beginner-friendly introduction to R. You should be comfortable using your laptop, installing software, navigating file systems, and creating/organizing files and folders. Basic familiarity with spreadsheets (e.g., Excel, Google Sheets) is helpful but not essential.

NOTE: If you are starting with “Day 1: Introduction to R and data wrangling in the tidyverse”, we recommend you only register for Days 1–2 as Days 3–5 will cover material that requires a more robust knowledge of R programming. After completing Days 1–2 we would encourage you to consider our July workshops for further tips/trick in R programming, and then come back the following February to complete the advanced statistics covered on Days 3–5.

Day 2: Introduction to visualisation and linear modelling

Today we continue where we left off yesterday. We will start with an introduction to visualisation in R using ggplot2 the core plotting package of the tidyverse. We will explore the range of plot types available — including scatter plots, line plots, boxplot, histograms — before moving on to demonstrate how to change the aesthetics of a plot, add annotations and save your creation. After lunch we will shift gears and begin learning the foundations of statistical modelling so you can analyse your data. We’ll start with simple linear models and explore why we fit models, learn how to interpret model output and how to upgrade your code to design more advanced linear models. We then learn how to select the best model, examining model diagnostics and plotting the output using ggplot2.

Difficulty rating: 🤓⚪⚪⚪⚪

Assumed knowledge: Day-1 content (loading data, basic data wrangling), foundational statistics including hypothesis testing, p-values, confidence intervals, and basic concepts such as correlation and regression. Some familiarity with RStudio and running R scripts is essential.

Day 3: Generalised linear modelling

Not all response variables are continuous, unbounded and have normal errors — counts, binary outcomes, proportions, and strictly positive measurements each require their own error distributions and mean-variance relationships. Today we extend the linear modelling framework to Generalised Linear Models (GLMs), introducing the three building blocks of every GLM: the linear predictor, the link function, and the error family (which defines the mean-variance relationship). We work through real ecological datasets using Poisson and Negative Binomial errors for count data, Gamma errors for continuous positive data, and Binomial errors for binary and proportion data. Along the way, we cover how to interpret coefficients on the link scale, assess model fit using deviance residuals and simulation-based diagnostics, and select the best model using information criteria and likelihood-ratio tests. We finish with a brief introduction to more specialised distributions — Tweedie and Beta — for zero-inflated continuous data and proportions.

Difficulty rating: 🤓🤓⚪⚪⚪

Assumed knowledge: Day-2 content (fitting and interpreting linear models, model selection, reading diagnostic plots). Foundational statistics including hypothesis testing, p-values, confidence intervals, and residuals. Some familiarity with probability distributions is helpful but not required.

Day 4: Mixed modelling

In this workshop you will learn how to fit Linear Mixed Models (LMMs), Generalised Linear Mixed Models (GLMMs), Generalised Additive Models (GAMs), and Generalised Additive Mixed Models (GAMMs). Mixed-effects models build on generalised Linear Models (GLMs) and have observations that are grouped in some way. This explicit recognition of grouping of observations within the model structure resolves many of the frequently encountered challenges associated with non-independence of observations, nested (hierarchical) designs, and spatial and temporal structuring. GAMs are extensions of GLMs and use flexible smoothers (wiggly lines) rather than mathematical equations to describe the relationship between the response and a set of predictors.

Difficulty rating: 🤓🤓🤓⚪⚪

Assumed knowledge: Solid understanding of linear and Generalized Linear Models (Day-2 material). At least 6 months of regular R programming experience, including data wrangling with tidyverse, writing custom functions, and interpreting model diagnostics. Some familiarity with concepts like random effects, nested designs, and non-independence will be advantageous.

Day 5: Spatial modelling with temporal and spatial autocorrelation

Most observations in a dataset are not independent, and ignoring spatial and temporal correlation does not just violate assumptions, it produces incorrect answers. This workshop equips you to model complex, structured data with confidence using Generalised Additive Models (GAMs) and Generalised Mixed Linear Models (GLMMs) with correlated residual structures. You’ll learn to diagnose autocorrelation in time series, point-based spatial data, and gridded aerial datasets, then address it using structured and penalised random effects for spatial, temporal, and spatiotemporal dependencies. By covering topics including model checking, validation, and interpretation, you’ll learn to transform messy, correlated data into defensible insights.

Difficulty rating: 🤓🤓🤓🤓⚪

Assumed knowledge: Strong proficiency in linear, generalized, and mixed models, plus GAMs (material from Days 2–3). Minimum 6-12 months of active R programming with ability to confidently understand functions, debug code, and create visualisations. Understanding of spatial/temporal data structures, autocorrelation concepts, and matrix operations will be beneficial.