Week 8: AI Incident Brief
~2–3 pages | Due before Week 9's first class | Submit via Canvas
Assignment
Analyze a real AI incident with quantitative harm estimation and disparity analysis.
Components
Submit as one PDF, using the numbered items below as your section headings in order.
- Incident summary - what happened, what system, what context
- Failure chain - technical, human, organizational, and governance failures
- Harm analysis with quantitative estimation - who was harmed, how many, for how long, estimated dollar cost or measurable impact, subgroup disparities. Before computing your own numbers, prompt an AI tool for a harm/scale estimate from the same source material. Then compute your own using the course’s framework (disparate impact ratio, equalized odds, errors/month × cost). Note in your AI use log where the AI’s estimate was unsupported, missing subgroup breakdown, or invented.
- Missing evaluation with disparity analysis - what testing was absent, what would a disparate impact analysis have revealed, what threshold should have triggered review
- Missing governance - rules and accountability absent
- Recommended safeguard with quantitative threshold - concrete measure with a measurable trigger
- Connection to your work - could this happen in your domain, with estimated scale impact
Any code you wrote to compute κ or MDD needs to be submitted as .ipynb and rendered as .html. Also include your AI use log covering the harm-estimate check in Component 3 - see the AI use log guide.
Rubric
| Criterion | Excellent (5) | Adequate (3) | Needs revision (1) |
|---|---|---|---|
| Incident clarity | Clear, accurate with context | Present but missing context | Vague |
| Failure chain | All four layers with specifics | Some layers | Only technical |
| Quantitative harm | Specific numbers: people, cost, duration, with AI’s estimate compared and critiqued in the use log | Numbers present but no AI comparison, or comparison superficial | No harm analysis |
| Disparity analysis | Subgroup impact estimated or computed | Mentions subgroups but no metrics | No subgroup analysis |
| Safeguard with threshold | Concrete with quantitative trigger | Present but no threshold | Missing |
| Professional connection | Thoughtful with estimated scale | Thin | No connection |
Total: 30 points