Student chat
The COMPAS Recidivism Algorithm
Privacy Notice: Student names have been replaced with pseudonyms to protect privacy. The audio is AI-generated narration, not the students’ voices. Emoji reactions are illustrative. Plain transcript.
The conversation, then the understanding quiz and the post-chat survey that followed it. All three are written out below in order.
01Chat transcript
Guide
👋 Welcome, everyone!
Our topic is COMPAS. We'll be discussing this statement:
The COMPAS recidivism algorithm is/was biased against Blacks.
- Luna agrees with the statement.
- Justin disagrees with the statement.
My name is Guide, and I'm here to support your discussion. If you need anything, start your message with "Hey Guide" 🤓
Luna
Justin
Luna
Justin
Justin
Luna
Justin
Guide
Luna, Justin is making a key mathematical point — that you literally cannot have both calibration AND equal error rates across groups when base rates differ.
But you're arguing the algorithm is "built on" structural racism through proxies like neighborhood. How do you respond to his claim that these factors help predict recidivism accurately, even if they correlate with race?
Is there a way to balance predictive accuracy with avoiding the reinforcement of systemic disadvantages?
the next day
Luna
Justin
the next day
Luna
Guide
Justin, Luna seems to be wrapping up, but there's still a key disagreement here.
You've argued the algorithm itself isn't biased if it meets calibration standards. She's arguing that using racially correlated proxies makes it biased regardless of calibration.
Do you think there's a meaningful distinction between "technically unbiased" and "practically biased" algorithms? Or is calibration sufficient to call something fair?
Justin
You have completed the chat! You may continue chatting after the Understanding Quiz 🤠
02Understanding quiz
Sway asks each student five questions about what their partner argued. Scores are accurate. Questions and responses may be updated from time to time to reflect improvements to Sway.
Justindisagrees
Understanding Quiz
5 questions about what Luna argued.
- She argued that COMPAS lacked calibration within groups and assigned different scores to defendants with identical histories.
- She argued that COMPAS used proxies tied to structural inequality and produced racially unequal patterns of prediction errors.
- She argued that COMPAS relied mainly on human judgment and therefore reproduced the personal prejudices of judges.
- She argued that COMPAS openly used race and produced risk scores that were inaccurate for both racial groups.
correct Justification
Luna said COMPAS "uses proxies that correlate strongly with structural inequalities" and cited ProPublica's finding that it "overpredicted recidivism for Black people and underpredicted it for White people." She did not claim that COMPAS directly asked about race or lacked within-group calibration.
- She believed calibration was unreliable because defendants with the same score had substantially different reoffending rates by race.
- She believed calibration mattered only when an algorithm excluded information about neighborhood and prior criminal history.
- She believed unequal base rates disappeared once the effects of historical discrimination were included in the model.
- She believed unequal error distribution and racially correlated inputs could reinforce disadvantage despite statistical calibration.
correct Justification
Luna accepted that COMPAS "met calibration" but argued that "the distribution of errors matters" and that factors such as neighborhood "indirectly encode structural racism." Her objection was therefore that calibration did not address the inequities the model could reinforce.
- She treated the tradeoff as evidence that differences in recidivism base rates are created entirely by COMPAS itself.
- She treated the tradeoff as a reason to reassess whether calibration is the right fairness goal when its errors cause unequal harm.
- She treated the tradeoff as a reason to preserve calibration while moving all fairness concerns outside the algorithm.
- She treated the tradeoff as proof that equal error rates provide the only defensible definition of algorithmic fairness.
correct Justification
Luna replied, "That is exactly why we cannot stop at calibration" and asked "whether calibration is the best solution" if it produces more harmful errors for Black defendants. She used the mathematical constraint to motivate ethical reassessment, not to deny the constraint.
- She proposed retaining the same risk question while applying lower decision thresholds to every racial group.
- She proposed replacing COMPAS with human evaluators who could interpret neighborhood information more compassionately.
- She proposed reframing the model to identify who needs support to avoid reoffending rather than only who is most likely to reoffend.
- She proposed predicting reoffending with fewer variables while leaving sentencing and reintegration policies unchanged.
correct Justification
Luna proposed changing the question from "Who is most likely to reoffend?" to "Who needs support to avoid reoffending?" She explicitly said this would not be "throwing out accuracy" but would help ensure that the model "serves justice too."
- She accepted that public safety outweighs fairness concerns, while maintaining that human judges should verify each risk score.
- She concluded that calibration resolves concerns about proxies, while social support should be handled independently of COMPAS.
- She concluded that predictive proxies should be removed even if their removal makes the algorithm lose most of its practical value.
- She acknowledged limits to removing proxies and simplifying the model, while maintaining that its effects should be repeatedly reviewed for inequity.
correct Justification
Luna said you made her think about "how far we can simplify an algorithm before it loses its value" and acknowledged that "proxies cannot be fully removed." However, she still emphasized a duty to "constantly revisit whether this system is inadvertently reinforcing inequities."
Lunaagrees
Understanding Quiz
5 questions about what Justin argued.
- He thought the errors resulted mainly from judges applying otherwise neutral risk scores differently.
- He thought public safety made disparities in prediction errors irrelevant to evaluating algorithmic fairness.
- He thought differing base rates can produce unequal errors in a binary prediction system even when it is calibrated.
- He thought the proxy variables were too weakly related to race to affect the pattern of errors.
correct Justification
Justin argued that "variations in base probabilities between groups are bound to cause unequal false-negative and false-positive rates" and that this "doesn't prove unfairness." He instead emphasized that COMPAS was calibrated within groups.
- He said they measured personal choices rather than social conditions, and therefore could not operate as racial proxies.
- He said they could improve prediction, and calibration gave him no compelling evidence that COMPAS used them to discriminate.
- He said they corrected unequal base rates, and therefore helped COMPAS produce equal error rates across groups.
- He said they should be excluded from sentencing, although they could still be used to provide reintegration services.
correct Justification
Justin said it was sensible "to give the algorithm as much information as possible" to identify likely reoffending, while adding that he did not see "a compelling argument that the algorithm used this information to discriminate." He also noted that it remained calibrated across groups.
- They should be removed if defendants consider them unfair, but retained if judges consider them informative.
- They should be removed if they affect error rates, but retained if they preserve calibration within groups.
- They should be removed if accuracy changes little without them, but retained if they are necessary for accuracy.
- They should be removed if they reveal income, but retained if they reveal neighborhood-level support systems.
correct Justification
Justin explicitly said, "if they can be safely removed without much effect on the algorithm's accuracy then they should be. But if they are necessary they should be kept."
- Guide asked whether equal error rates should replace calibration, or whether both standards could be achieved together.
- Guide asked whether social support should replace risk prediction or merely supplement sentencing decisions.
- Guide asked whether an algorithm can be technically unbiased yet practically biased, or whether calibration alone establishes fairness.
- Guide asked whether human judges identify race more readily than algorithms, or whether proxy removal solves that problem.
correct Justification
Guide framed the issue as whether there is "a meaningful distinction between 'technically unbiased' and 'practically biased' algorithms" and asked, "Or is calibration sufficient to call something fair?"
- He withdrew his reliance on calibration and concluded that equal error rates were the only defensible fairness standard.
- He accepted that proxy factors could create practical bias even if the algorithm remained technically unbiased by calibration.
- He concluded that COMPAS was technically biased but practically fair because it improved on unaided human judgment.
- He decided that proxy factors should be removed because public safety mattered less than preserving defendants' anonymity.
correct Justification
Earlier, Justin said, "the bias in the use of the algorithm was in how it was used" and that "the algorithm itself fulfilled its duty." At the end, he conceded that proxy factors "can make the algorithm practically biased while being technically unbiased" and that neighborhood and income factors could introduce practical bias.
03Post-chat survey
Last, both students rate the statement again and then rate a randomly sampled subset of our post-chat survey items. Like their transcripts, individual student opinions are never revealed to instructors.
Justindisagrees
Now you’ve had a chance to discuss the topic, rate your level of agreement with the original statement again.
Remember, all responses to Sway surveys are private and never shown to your instructor.
The COMPAS recidivism algorithm is/was biased against Blacks.
| Strongly disagree | Moderately disagree | Slightly disagree | No idea | Slightly agree | Moderately agree | Strongly agree |
|---|---|---|---|---|---|---|
How much do you agree with each of these?
| Statement | Strongly Disagree | Disagree | Neutral | Agree | Strongly Agree |
|---|---|---|---|---|---|
| Guide's contributions improved the discussion | |||||
| Guide treated me and my partner with equal respect | |||||
| I felt comfortable sharing my honest opinions with my partner | |||||
| I was not offended by my partner's perspective | |||||
| My partner was respectful | |||||
| It was valuable to chat with a student who did NOT share my perspective |
Lunaagrees
Now you’ve had a chance to discuss the topic, rate your level of agreement with the original statement again.
Remember, all responses to Sway surveys are private and never shown to your instructor.
The COMPAS recidivism algorithm is/was biased against Blacks.
| Strongly disagree | Moderately disagree | Slightly disagree | No idea | Slightly agree | Moderately agree | Strongly agree |
|---|---|---|---|---|---|---|
How much do you agree with each of these?
| Statement | Strongly Disagree | Disagree | Neutral | Agree | Strongly Agree |
|---|---|---|---|---|---|
| Guide's contributions improved the discussion | |||||
| Guide treated me and my partner with equal respect | |||||
| I was not offended by my partner's perspective | |||||
| My partner had better reasons for their views than I expected | |||||
| It was valuable to chat with a student who did NOT share my perspective | |||||
| Sway helped me articulate my thoughts/feelings better |
How did this Sway chat affect your confidence discussing complex issues with people who hold different views from you?
Opinion change
Justindisagrees
+3
Lunaagrees
—
before the chat after the chat both, unchanged
Up next
All the chats
Sixteen transcripts of students paired with a classmate who disagreed with them.
Back to the examples →The student app
Every screen a student can reach in Sway, drawn as a wireframe you can click around in.
Explore the app →The research
What we measure after chats like this one: opinion change, and how students rate the people they disagree with.
Read the research →