AI-R05 — AI for Quantitative Research and Data Analysis

Wishlist Share

About Course

Course Code: AI-R05  |  School: School of Artificial Intelligence  |  Cluster: Level 2 — Research & Academic Practice

Level: Intermediate  |  Duration: 6 weeks · 20–26 learning hours  |  Language: English  |  Certificate: Professional Certificate (non-degree)  |  Format: Self-paced with AI support under human supervision

Overview

A generative assistant will write analysis code that runs. Whether that code answers the question asked, and whether the answer means what the researcher thinks it means, are separate matters that the assistant cannot settle. This course is organised around that separation.

Learners work through the full quantitative pipeline: data cleaning with a decision log, exploratory description, model specification tied to the research question, estimation, diagnostics, and interpretation with honest uncertainty. Machine assistance is used for code drafting, for explaining unfamiliar output, and for suggesting diagnostics, always followed by verification against a reference text or a known-answer test.

Considerable attention is given to the practices that make quantitative results credible: pre-specifying the analysis where possible, documenting every deviation, reporting effect sizes and intervals rather than significance alone, and making the analysis reproducible by someone who has only the data and the script.

Learning outcomes

On completion, a successful learner will be able to:

  1. Clean and prepare a dataset with a complete decision log covering missingness, outliers and transformations.
  2. Specify a statistical model that answers the stated research question, and justify the specification.
  3. Draft, verify and correct analysis code produced with machine assistance, including through known-answer tests.
  4. Run and interpret appropriate diagnostics, and act on what they show.
  5. Report results with effect sizes and uncertainty intervals, avoiding inference that the design does not support.
  6. Produce a reproducible analysis package that a third party can re-run from raw data to reported figures.

Who this course is for

Researchers and analysts working with survey, administrative or measurement data; doctoral candidates in the social, health and policy sciences; institutional and government analysts.

Prerequisites

Prior study of introductory statistics, including regression. Working familiarity with one statistical environment such as R, Stata, Python or SPSS. AI-F03 recommended. Learners must bring a dataset they are permitted to analyse.

Syllabus

Module 1 — Cleaning as a documented set of decisions

Focus. Missingness mechanisms and their implications, outlier judgement, transformation, and the principle that every cleaning decision is an analytic decision requiring a recorded reason.

Lessons. 1.1 Missing completely at random, at random, and not at random. 1.2 Deletion, imputation and their consequences. 1.3 Outliers: error, extreme value, or finding. 1.4 The cleaning decision log.

Core reading. Andrew Gelman, Jennifer Hill & Aki Vehtari, Regression and Other Stories (Cambridge: Cambridge University Press, 2020), chapters 2–3.

Deliverable. Cleaned dataset with a complete decision log.

Module 2 — Description before modelling

Focus. Univariate and bivariate description, visualisation that reveals rather than decorates, and the practice of understanding the data thoroughly before fitting anything.

Lessons. 2.1 Distributions and what they warn you about. 2.2 Bivariate structure and confounding hints. 2.3 Graphics for the analyst, not the reader. 2.4 What description already answers.

Core reading. Gelman, Hill & Vehtari, Regression and Other Stories, chapters 1 and 4.

Deliverable. Descriptive report with annotated graphics and the questions they raise.

Module 3 — Specification tied to the question

Focus. Choosing a model because it corresponds to the claim being made. Functional form, covariate selection as a substantive rather than statistical decision, and the hazards of stepwise and data-driven selection.

Lessons. 3.1 From question to estimand. 3.2 Functional form. 3.3 Covariate choice as substantive judgement. 3.4 Why stepwise selection fails.

Core reading. Andrew Gelman & Jennifer Hill, Data Analysis Using Regression and Multilevel/Hierarchical Models (Cambridge: Cambridge University Press, 2007), chapters 3–4 and 9.

Deliverable. Specification memo with the estimand, model and justification of covariates.

Module 4 — Machine-assisted code, verified

Focus. Drafting analysis code with assistance and then proving it correct: known-answer tests, replication of a textbook example, and independent recomputation of key quantities.

Lessons. 4.1 Prompting for analysis code. 4.2 Known-answer tests. 4.3 Replicating a published example. 4.4 The error profile of generated statistical code.

Core reading. Roger D. Peng, “Reproducible Research in Computational Science”, Science 334, no. 6060 (2011): 1226–1227.

Deliverable. Verification notebook: generated code, the tests applied, and every error found and corrected.

Module 5 — Diagnostics and honest uncertainty

Focus. Residual and assumption checks, influence, sensitivity to specification, and the reporting of effect size and interval rather than significance as a verdict.

Lessons. 5.1 Residuals, influence and leverage. 5.2 Sensitivity across specifications. 5.3 Effect size and interval estimation. 5.4 What a p-value does not tell you.

Core reading. Ronald L. Wasserstein & Nicole A. Lazar, “The ASA’s Statement on p-Values: Context, Process, and Purpose”, The American Statistician 70, no. 2 (2016): 129–133. Joseph P. Simmons, Leif D. Nelson & Uri Simonsohn, “False-Positive Psychology”, Psychological Science 22, no. 11 (2011): 1359–1366.

Deliverable. Diagnostic and sensitivity report with revised estimates.

Module 6 — Reproducibility and reporting

Focus. Packaging the analysis so it can be re-run, documenting deviations from the analysis plan, and writing results that do not overstate what the design supports.

Lessons. 6.1 Project structure and dependency capture. 6.2 Documenting deviations. 6.3 FAIR data practice. 6.4 Writing results without overclaiming.

Core reading. Mark D. Wilkinson et al., “The FAIR Guiding Principles for Scientific Data Management and Stewardship”, Scientific Data 3 (2016): 160018. Brian A. Nosek et al., “The Preregistration Revolution”, PNAS 115, no. 11 (2018).

Deliverable. Final submission: reproducible analysis package, results section and AI-use disclosure.

Assessment

Component Weight
Cleaning decision log 15%
Descriptive report 12%
Specification memo 18%
Verification notebook 20%
Diagnostic and sensitivity report 20%
Reproducible package and results section 15%
Total 100%

Pass mark 70 per cent. All assessed components must be attempted. Every mark in this course is issued by a human assessor; no assessment outcome is generated automatically.

Rubric criteria

Each assessed artefact is marked against four criteria at four levels (distinction, pass with merit, pass, fail).

  1. Decision transparency: is every analytic choice recorded with a reason a reviewer could evaluate?
  2. Code correctness: is generated code demonstrably verified rather than merely executed without error?
  3. Inferential restraint: are conclusions limited to what the design and data support?
  4. Reproducibility: can a third party regenerate every reported figure from the supplied package?

Reading list

Core. Andrew Gelman, Jennifer Hill & Aki Vehtari, Regression and Other Stories (Cambridge: Cambridge University Press, 2020). Andrew Gelman & Jennifer Hill, Data Analysis Using Regression and Multilevel/Hierarchical Models (Cambridge: Cambridge University Press, 2007).

Peer-reviewed. Ronald L. Wasserstein & Nicole A. Lazar, “The ASA’s Statement on p-Values”, The American Statistician 70, no. 2 (2016). Joseph P. Simmons, Leif D. Nelson & Uri Simonsohn, “False-Positive Psychology”, Psychological Science 22, no. 11 (2011). Brian A. Nosek et al., “The Preregistration Revolution”, PNAS 115, no. 11 (2018). Mark D. Wilkinson et al., “The FAIR Guiding Principles”, Scientific Data 3 (2016). Roger D. Peng, “Reproducible Research in Computational Science”, Science 334 (2011).

All items are published works identifiable by author, title and publisher. Learners obtain them through an institutional library or the publisher. The Academy does not distribute copyrighted texts.

Academic integrity and use of AI

Generative tools may be used in producing assessed work under three conditions. Use must be disclosed in a short statement appended to each submission, naming the tool and the task it performed. Any factual or technical claim originating from a generative tool must be verified against a citable source before it enters assessed work, and the verification must be evidenced. The analytical judgement in each artefact must be the learner’s own and must be defensible in a short follow-up. Reported numbers must come from the learner’s own executed analysis. Figures produced by a language model rather than by computation are treated as fabricated data.

Instructor: pending owner confirmation. Pricing: pending owner approval. Reference list verified against publisher records; any later addition is marked for verification before publication.

Show More

Course Content

Module 0 — Start Here

  • Welcome and How This Course Works

Module 1 — Core Concepts

Module 2 — Frameworks and Standards

Module 3 — Evidence and Sources

Module 4 — Analysis

Module 5 — Cases and Application

Module 6 — Assessment Preparation

Module 7 — Final Project

Student Ratings & Reviews

No Review Yet
No Review Yet

Want to receive push notifications for all major on-site activities?