Privacy Notice: Student names have been replaced with pseudonyms to protect privacy.
Students Discuss COMPAS
Prefer to listen to the chat? Click play below!
Audio narration is AI-generated from the text transcript and does not represent the students' actual voices.
0:00
0:00
GuideSeptember 26, 2025, 4:33 PM
👋 Welcome, everyone!
Our topic is COMPAS. We'll be discussing this statement:
The COMPAS recidivism algorithm is/was biased against Blacks.
Luna agrees with the statement.
Justin disagrees with the statement.
My name is Guide, and I'm here to support your discussion. If you need anything, start your message with "Hey Guide" 🤓
LunaSeptember 26, 2025, 4:47 PM
Hi Justin, nice to meet you! I'm Luna. It seems that we've been matched for the Sway discussion about the COMPAS topic. To be honest, this is my first time using Sway, so I'm trying to adapt to it. From the Guide's suggestion, it is a good idea if we talk about how we'd like to approach this discussion as the first step. Do you have any idea? I'm also okay if you want to discuss the topic directly, starting with your opinion/my opinion
JustinSeptember 26, 2025, 4:51 PM
Hey Luna! Nice to meet you as well :) yes i'm a bit confused about sway as well lol. Not really sure what we're being graded on. But yeah it seems like we are supposed to just start discussing the topic. If it works for you we could both lay out our opinions and then we can discuss each one
LunaSeptember 26, 2025, 5:19 PM
Sure! In my opinion, COMPAS recidivism algorithm is biased against Blacks. From what I got in yesterday's slide, this algorithm doesn't directly ask about race, but it uses proxies that correlate strongly with structural inequalities. Furthermore, ProPublica reported COMPAS systematically overpredicted recidivism for Black people and underpredicted it for White people. Eventhough Northpointe said it is because of different base rates, however a "fair" algorithm must consider not only prediction quality but also social context and justice in outcomes. Since higher recorded recidivism rates for Black people are partly the result of systemic inequality, discrimination, etc. What's your opinion? :)
JustinSeptember 26, 2025, 5:36 PM
Thanks for asking! I think you make some very strong points that i'm looking forward to discussing :). When I first read the prompt for this, my immediate reaction was to agree with the statement. To me, the most compelling supporting evidence for the COMPAS algorithm being biased was the drastically higher false positive (predicted to reoffend but they actually don't) rate for black defendants and the similarly disproportionately high false negative rate for white defendants (predicted not to reoffend but they actually do). However, I believe Brian Hedden, in his article "On Statistical Criteria of Algorithmic Fairness" makes a compelling case for why ensuring equal false-negative rates and equal false-positive rates across ethnic groups make for arbitrary criterion for a fair algorithm. Unfortunately, some groups are statistically more likely to reoffend compared to others. This is undeniably a direct cause of a long history of systematic oppression and discrimination against these groups. These injustices have caused lasting impacts on factors that can unavoidably help predict recidivism rates (less access to education, weaker support systems, etc.) As unfortunate as this is, it does help to explain the lopsided false-negative and false-positive rates across groups. Hedden's 'perfectly fair algorithm' thought experiment puts this in a way that is easy to digest. When the outcome is binary (either an individual reoffends or they don't) variations in base probabilities between groups are bound to cause unequal false-negative and false-positive rates. It doesn't prove unfairness. In fact, Hedden found that the only criteria he could not disprove for devising a fair algorithm was calibration within groups ("for each possible risk score, for the expected percentage of individuals assigned that risk score who are actually positive is the same for each relevant group and is equal in that risk score.") The COMPAS algorithm did uphold this criteria. That is, people who were assigned a risk score of 7 had the same reoffending rate whether they were white or black. To me, this indicates that the algorithm was reasonably fair and unbiased towards ethnic groups.
JustinSeptember 26, 2025, 5:46 PM
To reply to your stance, I had the same concern about COMPAS using proxies to discriminate by race. I feel uncomfortable about that and believe more should be done to obscure the race of defendants. I actually agree that how the algorithm was utilized was biased towards black people. You're completely right that there is vital background information about why recidivism rates are higher for black people than for other groups. This information should not be ignored as it currently is. And I think it makes sense to be more forgiving towards individuals who suffer from the lasting impacts of discrimination. However, I believe the bias in the use of the algorithm was in how it was used. I think the algorithm itself fulfilled its duty in fairly predicting recidivism regardless of ethnicity.
LunaSeptember 26, 2025, 7:09 PM
Thank you for such a thoughtful reply! :) From your opinion, I can see your point of view on this topic with Hedden's argument, especially his point that equal false-positive and false-negative rates might not be the right measure of fairness. Sometimes differences in base rates can naturally cause imbalances. I also agree with you that COMPAS was calibrated within groups, so it was "fair" in the way Hedden defines it. However, there is a part of me that still thinks COMPAS was biased against Black people. Even if Black & White defendants with the same scores are equally likely to reoffend, the distribution of errors matters. Moreover, as I said before, factors like prior neighborhood strongly correlate with systemic discrimination, so it still indirectly encodes structural racism even though the algorithm is technically fair. Therefore, the bias is not just how it is used, but in what it is built on. While I agree that COMPAS met calibration and did what it was supposed to do statistically, I do not think that makes it fair. A truly fair algorithm should not only reflect the base rate gap, but also needs to actively avoid reinforcing the disadvantages already faced by Black defendants🤔
JustinSeptember 26, 2025, 10:09 PM
I get where you're coming from about the higher prediction errors disproportionately affecting black people. I did some digging on that and it turns out that due to the differing base rates of recidivism it's mathematically impossible to have both calibration across groups and equal false-positive/false-negative rates. You quite literally have to choose one or the other. To your point about the algorithm considering factors like prior neighbourhood, part of me agrees with you that these questions might indirectly help identify race, but at the same time, these are convicted criminals. I think it's sensible to prioritize correctly identifying individuals who are likely to reoffend, and as such to give the algorithm as much information as possible. These factors at best can allow you to make an educated guess at one's race, but i don't see a compelling argument that the algorithm used this information to discriminate against certain ethnic groups. The algorithm was still calibrated across groups, despite having access to these potentially discriminatory factors. I do agree that the usage of the algorithm needs to be rethought in order to avoid reinforcing disadvantages, but i don't think the algorithm itself has anything to do with that.
GuideSeptember 26, 2025, 10:09 PM
Luna, Justin is making a key mathematical point — that you literally cannot have both calibration AND equal error rates across groups when base rates differ.
But you're arguing the algorithm is "built on" structural racism through proxies like neighborhood. How do you respond to his claim that these factors help predict recidivism accurately, even if they correlate with race?
Is there a way to balance predictive accuracy with avoiding the reinforcement of systemic disadvantages?
LunaSeptember 27, 2025, 3:32 PM
Hi Justin, sorry for late reply! Your arguments really make sense, especially the algorithm cannot have calibration and equal rates at the same time. That is exactly why we cannot stop at calibration. We still need to rethink our decision frequently (based on the Ethical Thinking definition). If sticking with calibration leads to more harmful errors for Black defendants, then the question arises whether calibration is the best solution for the fairness problem. I think the key issue is what kind of accuracy we are aiming for. If proxies like neighborhood data are strongly correlated with race, then it may make the algorithm better at reproducing biased outcomes than at measuring someone's true likelihood of reoffending. That is why I think we have to be careful about saying "more predictive features" with "more fairness". And as you said, the algorithm needs to be rethought, since the real challenge is balancing accuracy with fairness. In my opinion, one possible way is to reframe the question instead of just asking "Who is most likely to reoffend?" We could ask, "Who needs support to avoid reoffending?" By addressing this kind of approach, we are not throwing out accuracy, but we are making sure the model serves justice too:)
JustinSeptember 27, 2025, 5:46 PM
No worries at all! Thanks for such a thoughtful reply :) your compassion really comes through in these messages. I am curious how predictive these proxy factors actually are, and how much worse the algorithm would be without them. I think that if they can be safely removed without much effect on the algorithm's accuracy then they should be. But if they are necessary they should be kept. Like I was saying, at the end of the day the safety of society should be prioritized over the anonymity of a convicted criminal. I think it's worth keeping in mind that the algorithm already obscures a defendant's ethnicity much more than having humans judge a defendant's likelihood of reoffending (obviously lol). And i do wonder that if you remove these factors, if there will be a new bunch of factors to take issue with. I imagine it's probably easier than we both expect to take a pretty good guess at one's race no matter how much you strip down the questions asked. It only really makes sense to me to be looking at removing these questions if there is reason to believe the algorithm is biased, which i don't see any evidence for. I think your point about supporting these defendants is a great one and i agree with that fully. The current system definitely does not do enough to help reintegrate people into society.
LunaSeptember 28, 2025, 4:34 PM
Thank you so much, Justin! I've really appreciated this discussion:) You've made me think more critically about calibration and how far we can simplify an algorithm before it loses its value. I agree with you that public safety is important, but we must also consider fairness in this context. Even if proxies cannot be fully removed, but we have a responsibility to constantly revisit whether this system is inadvertently reinforcing inequities. At the end of the day, I think we're both on the same page that the algorithm alone is not enough. We need better social support to make reintegration possible😊
GuideSeptember 28, 2025, 4:34 PM
Justin, Luna seems to be wrapping up, but there's still a key disagreement here.
You've argued the algorithm itself isn't biased if it meets calibration standards. She's arguing that using racially correlated proxies makes it biased regardless of calibration.
Do you think there's a meaningful distinction between "technically unbiased" and "practically biased" algorithms? Or is calibration sufficient to call something fair?
JustinSeptember 28, 2025, 5:11 PM
That's so nice of you to say, i've really appreciated this discussion as well! I agree, it seems like we're on the same page that the algorithm can be unbiased but still unfair if used in the wrong manner. To answer the AI question, yes i think there's a meaningful distinction there. The proxy factors are a great example of something that can make the algorithm practically biased while being technically unbiased. There's certainly a case to be made that factors like previous neighborhoods introduce bias. Since you're being judged based on your neighbours actions, your level of income, etc. i can see the point to be made for that introducing practical bias to the algorithm.
Understanding Quiz
Justin
At the beginning of the discussion, what reasoning did Luna use to support her concern about COMPAS?
Justification
She argued that COMPAS lacked calibration within groups and assigned different scores to defendants with identical histories.
She argued that COMPAS used proxies tied to structural inequality and produced racially unequal patterns of prediction errors.
She argued that COMPAS relied mainly on human judgment and therefore reproduced the personal prejudices of judges.
She argued that COMPAS openly used race and produced risk scores that were inaccurate for both racial groups.
After acknowledging your point about calibration, why did Luna continue to regard COMPAS as unfair?
Justification
She believed calibration was unreliable because defendants with the same score had substantially different reoffending rates by race.
She believed calibration mattered only when an algorithm excluded information about neighborhood and prior criminal history.
She believed unequal base rates disappeared once the effects of historical discrimination were included in the model.
She believed unequal error distribution and racially correlated inputs could reinforce disadvantage despite statistical calibration.
When you explained that calibration and equal error rates cannot both be achieved when base rates differ, how did Luna use that point in her response?
Justification
She treated the tradeoff as evidence that differences in recidivism base rates are created entirely by COMPAS itself.
She treated the tradeoff as a reason to reassess whether calibration is the right fairness goal when its errors cause unequal harm.
She treated the tradeoff as a reason to preserve calibration while moving all fairness concerns outside the algorithm.
She treated the tradeoff as proof that equal error rates provide the only defensible definition of algorithmic fairness.
How did Luna propose balancing predictive goals with justice rather than simply discarding accuracy?
Justification
She proposed retaining the same risk question while applying lower decision thresholds to every racial group.
She proposed replacing COMPAS with human evaluators who could interpret neighborhood information more compassionately.
She proposed reframing the model to identify who needs support to avoid reoffending rather than only who is most likely to reoffend.
She proposed predicting reoffending with fewer variables while leaving sentencing and reintegration policies unchanged.
Toward the end of the discussion, how had your arguments affected Luna's position on predictive proxies and algorithm design?
Justification
She accepted that public safety outweighs fairness concerns, while maintaining that human judges should verify each risk score.
She concluded that calibration resolves concerns about proxies, while social support should be handled independently of COMPAS.
She concluded that predictive proxies should be removed even if their removal makes the algorithm lose most of its practical value.
She acknowledged limits to removing proxies and simplifying the model, while maintaining that its effects should be repeatedly reviewed for inequity.
Luna
At the beginning of the discussion, why did Justin think the unequal error rates cited by ProPublica did not establish that COMPAS itself was unfair?
Justification
He thought the errors resulted mainly from judges applying otherwise neutral risk scores differently.
He thought public safety made disparities in prediction errors irrelevant to evaluating algorithmic fairness.
He thought differing base rates can produce unequal errors in a binary prediction system even when it is calibrated.
He thought the proxy variables were too weakly related to race to affect the pattern of errors.
When you argued that neighborhood and similar factors encoded structural racism, how did Justin initially defend their inclusion?
Justification
He said they measured personal choices rather than social conditions, and therefore could not operate as racial proxies.
He said they could improve prediction, and calibration gave him no compelling evidence that COMPAS used them to discriminate.
He said they corrected unequal base rates, and therefore helped COMPAS produce equal error rates across groups.
He said they should be excluded from sentencing, although they could still be used to provide reintegration services.
Later, what test did Justin propose for deciding whether racially correlated proxy factors should be removed?
Justification
They should be removed if defendants consider them unfair, but retained if judges consider them informative.
They should be removed if they affect error rates, but retained if they preserve calibration within groups.
They should be removed if accuracy changes little without them, but retained if they are necessary for accuracy.
They should be removed if they reveal income, but retained if they reveal neighborhood-level support systems.
Near the end, how did Guide sharpen the unresolved disagreement between you and Justin?
Justification
Guide asked whether equal error rates should replace calibration, or whether both standards could be achieved together.
Guide asked whether social support should replace risk prediction or merely supplement sentencing decisions.
Guide asked whether an algorithm can be technically unbiased yet practically biased, or whether calibration alone establishes fairness.
Guide asked whether human judges identify race more readily than algorithms, or whether proxy removal solves that problem.
How did Justin's final position differ from his earlier separation between the algorithm and its use?
Justification
He withdrew his reliance on calibration and concluded that equal error rates were the only defensible fairness standard.
He accepted that proxy factors could create practical bias even if the algorithm remained technically unbiased by calibration.
He concluded that COMPAS was technically biased but practically fair because it improved on unaided human judgment.
He decided that proxy factors should be removed because public safety mattered less than preserving defendants' anonymity.
Survey results
Opinion Changes
Students rated the following statement: The COMPAS recidivism algorithm is/was biased against Blacks.
Strongly disagree
Moderately disagree
Slightly disagree
No idea
Slightly agree
Moderately agree
Strongly agree
Justin
+3
Luna
—
Pre-chat opinion
Post-chat opinion
Moved toward agreement
Moved toward disagreement
Partner Ratings
Statement
Strongly Disagree
Disagree
Neutral
Agree
Strongly Agree
Guide's contributions improved the discussion
—
—
—
Justin
Luna
Guide treated me and my partner with equal respect
—
—
—
—
JustinLuna
I felt comfortable sharing my honest opinions with my partner
—
—
—
—
Justin
I was not offended by my partner's perspective
—
—
—
—
JustinLuna
My partner was respectful
—
—
—
—
Justin
My partner had better reasons for their views than I expected
—
—
—
—
Luna
It was valuable to chat with a student who did NOT share my perspective
—
—
—
—
JustinLuna
Sway helped me articulate my thoughts/feelings better
—
—
—
—
Luna
Optional open feedback
"How did this Sway chat affect your confidence discussing complex issues with people who hold different views from you?"
Luna: "Sway chat is really helpful, especially to express my opinion with other people! I like some of the features, like the Guide who assists us during the discussion and the Sway notification that always reminds me about the discussion. Thank you!"