About Course
Course Code: AI-G03 | School: School of Artificial Intelligence | Cluster: Level 4 — Governance & Policy
Level: Advanced | Duration: 6 weeks · 24–30 learning hours | Language: English | Certificate: Executive Certificate (non-degree) | Format: Self-paced with AI support under human supervision
Overview
This course treats responsible AI as an engineering and control problem rather than a values statement. Its subject is the four exposures an institution actually carries when it deploys a machine learning or generative system: performance risk, allocative and representational bias, privacy leakage, and adversarial security.
Each is handled the same way. Learners study documented cases where the failure occurred, learn the measurement or testing technique that would have surfaced it, apply that technique to a system in their own environment, and specify the control that reduces the exposure. The output is a risk register with evidence, not a policy statement.
The course insists on a distinction that is often blurred: fairness is not a single quantity, and the common statistical definitions of it are mathematically incompatible with one another in most realistic settings. Learners are required to choose a definition, justify the choice against the decision being made, and state what that choice sacrifices.
Learning outcomes
On completion, a successful learner will be able to:
- Construct a risk register for a deployed or proposed AI system, with each entry traced to an observable failure mode.
- Measure group-level performance disparity in a system’s outputs and interpret the result without overclaiming.
- Select and justify a fairness criterion for a specific decision, and state explicitly what the choice trades away.
- Assess privacy exposure including training-data leakage, membership inference and inadvertent disclosure through prompts.
- Test for the principal security exposures of language-model applications, including prompt injection and insecure output handling.
- Specify proportionate controls and residual risk in a form an accountable executive can sign.
Who this course is for
Data protection officers, risk and compliance staff, IT security leads, institutional researchers, and technical staff responsible for systems already in production. Also suitable for consultants conducting AI assurance work.
Prerequisites
AI-F03 or equivalent understanding of how models are trained and evaluated. Comfort reading tables of results and simple statistics. No programming is required, but learners with access to system logs and outputs will produce stronger assessed work.
Syllabus
Module 1 — Risk as a register, not a sentiment
Focus. Moving from generic AI-risk language to specific, observable failure modes with owners, likelihood, consequence and controls. Establishing the register that the rest of the course populates.
Lessons. 1.1 From values to failure modes. 1.2 Likelihood and consequence for statistical systems. 1.3 Ownership and the accountable person. 1.4 What a defensible register looks like.
Core reading. NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), MAP and MEASURE functions.
Deliverable. Initial risk register for one system, with at least eight entries traced to observable failure modes.
Module 2 — Performance failure and the limits of accuracy
Focus. Why aggregate accuracy conceals the failures that matter, error analysis by subgroup and by consequence, and the behaviour of a system on inputs it was never trained to handle.
Lessons. 2.1 Aggregate metrics and what they hide. 2.2 Error analysis by subgroup and by cost. 2.3 Out-of-distribution behaviour. 2.4 Silent degradation after deployment.
Core reading. Solon Barocas, Moritz Hardt & Arvind Narayanan, Fairness and Machine Learning: Limitations and Opportunities (Cambridge, MA: MIT Press, 2023), chapters 1–2.
Deliverable. Disaggregated error analysis for one system output set.
Module 3 — Bias: allocative, representational, and incompatible definitions
Focus. Documented cases of allocative harm, the proxy problem, and the formal impossibility results that prevent satisfying multiple fairness criteria at once. Choosing and defending one criterion.
Lessons. 3.1 Allocative and representational harm. 3.2 Proxies and how they encode history. 3.3 Competing fairness definitions and their incompatibility. 3.4 Choosing a criterion and stating the sacrifice.
Core reading. Ziad Obermeyer et al., “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations”, Science 366, no. 6464 (2019): 447–453. Joy Buolamwini & Timnit Gebru, “Gender Shades”, Proceedings of FAT* (ACM, 2018). Ninareh Mehrabi et al., “A Survey on Bias and Fairness in Machine Learning”, ACM Computing Surveys 54, no. 6 (2021).
Deliverable. Fairness memorandum: chosen criterion, justification, measured disparity, and the stated trade-off.
Module 4 — Privacy: leakage, inference and disclosure
Focus. Training-data extraction and membership inference as demonstrated attacks, the specific institutional risk of confidential material entering prompts, and the controls that reduce both.
Lessons. 4.1 Extraction and memorisation. 4.2 Membership inference. 4.3 Prompt-side disclosure and vendor retention. 4.4 Minimisation, retention limits and access control.
Core reading. Nicholas Carlini et al., “Extracting Training Data from Large Language Models”, Proceedings of the 30th USENIX Security Symposium (2021). Reza Shokri et al., “Membership Inference Attacks Against Machine Learning Models”, IEEE Symposium on Security and Privacy (2017).
Deliverable. Privacy exposure assessment with a data-flow diagram for one AI-assisted process.
Module 5 — Security of language-model applications
Focus. Prompt injection as an architectural rather than a prompting problem, insecure output handling, excessive agency, supply-chain exposure, and the testing regime that surfaces each.
Lessons. 5.1 Direct and indirect prompt injection. 5.2 Insecure output handling and downstream execution. 5.3 Excessive agency and permission scope. 5.4 Supply chain and model provenance.
Core reading. OWASP, Top 10 for Large Language Model Applications (OWASP Foundation, current edition). Nicolas Papernot et al., “SoK: Security and Privacy in Machine Learning”, IEEE European Symposium on Security and Privacy (2018).
Deliverable. Security test report against the OWASP LLM categories for one application.
Module 6 — Controls, residual risk and sign-off
Focus. Selecting proportionate controls, distinguishing controls that reduce likelihood from those that reduce consequence, and writing a residual-risk statement an executive can sign without being misled.
Lessons. 6.1 Control selection and proportionality. 6.2 Likelihood controls versus consequence controls. 6.3 Assurance evidence and internal audit. 6.4 Writing residual risk honestly.
Core reading. Inioluwa Deborah Raji et al., “Closing the AI Accountability Gap”, Proceedings of FAT* (ACM, 2020). NIST AI RMF 1.0, MANAGE function.
Deliverable. Final submission: completed risk register with controls, evidence, and a signed-off residual-risk statement.
Assessment
| Component | Weight |
| Initial risk register | 12% |
| Disaggregated error analysis | 16% |
| Fairness memorandum | 22% |
| Privacy exposure assessment | 18% |
| Security test report | 17% |
| Final register and residual-risk statement | 15% |
| Total | 100% |
Pass mark 70 per cent. All assessed components must be attempted. Every mark in this course is issued by a human assessor; no assessment outcome is generated automatically.
Rubric criteria
Each assessed artefact is marked against four criteria at four levels (distinction, pass with merit, pass, fail).
- Evidential basis: is every register entry supported by measurement, testing or documentation rather than assertion?
- Technical correctness: are metrics, fairness definitions and attack categories used accurately?
- Proportionality of control: is each control matched to the exposure it addresses, without security theatre?
- Honesty about residual risk: does the final statement enable an accountable person to decide, including where the answer is uncomfortable?
Reading list
Core. Solon Barocas, Moritz Hardt & Arvind Narayanan, Fairness and Machine Learning: Limitations and Opportunities (Cambridge, MA: MIT Press, 2023).
Peer-reviewed. Ziad Obermeyer et al., “Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations”, Science 366, no. 6464 (2019). Joy Buolamwini & Timnit Gebru, “Gender Shades”, Proceedings of FAT* (ACM, 2018). Ninareh Mehrabi et al., “A Survey on Bias and Fairness in Machine Learning”, ACM Computing Surveys 54, no. 6 (2021). Nicholas Carlini et al., “Extracting Training Data from Large Language Models”, USENIX Security (2021). Reza Shokri et al., “Membership Inference Attacks Against Machine Learning Models”, IEEE S&P (2017). Laura Weidinger et al., “Taxonomy of Risks Posed by Language Models”, Proceedings of FAccT (ACM, 2022).
Standards. NIST, AI Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023). ISO/IEC 42001:2023. OWASP, Top 10 for Large Language Model Applications (current edition).
All items are published works identifiable by author, title and publisher. Learners obtain them through an institutional library or the publisher. The Academy does not distribute copyrighted texts.
Academic integrity and use of AI
Generative tools may be used in producing assessed work under three conditions. Use must be disclosed in a short statement appended to each submission, naming the tool and the task it performed. Any factual or technical claim originating from a generative tool must be verified against a citable source before it enters assessed work, and the verification must be evidenced. The analytical judgement in each artefact must be the learner’s own and must be defensible in a short follow-up. Assessed work in this course depends on real measurement. Fabricated or model-generated figures presented as measurement will be treated as an integrity breach, not a stylistic error.
Instructor: pending owner confirmation. Pricing: pending owner approval. Reference list verified against publisher records; any later addition is marked for verification before publication.
Course Content
Module 0 — Start Here
-
Welcome and How This Course Works