In this study, we assessed whether ML predictions of inpatient aggression in acute psychiatric care are unfair. To our knowledge, this is the most comprehensive fairness assessment of ML as related to this outcome, and builds on previous work by Dobbins et al. 19 by examining a wider range of social determinants and applying an intersectional approach. A random forest model was trained on a range of demographic, clinical, admission, and risk assessment data, yielding an F1 score of 0.22, a PR-AUC of 0.13, and a ROC-AUC of 0.81. Although maximizing predictive performance was not an emphasis of this study, the model achieved comparable performance to ML algorithms reported in prior research trained on tabular data in clinically heterogenous psychiatric populations (ROC-AUC obtained in Suchting et al. = 0.7823, Menger et al. = 0.7626, Wang et al. = 0.6327, Danielsen et al. = 0.8728). The fairness assessment revealed the algorithm violates both disparate mistreatment and equalized odds: there were significant disparities in FPR, TPR, and ROC-AUC curves across race/ethnicity, gender, admission mode, citizenship, and housing status. Relative to other groups, FPR was elevated in individuals who are Middle Eastern and Black, those who identify as male, are admitted into emergency care by the police, Canadian citizens, and with unstable or supportive forms of housing. Intersectional analyses revealed that Middle Eastern men had the highest FPR among all groups. There were significant differences in TPR and ROC-AUC curves in relation to the FPR of each group, suggesting the nature of algorithmic unfairness differs between groups. For example, in the case of patients who are Middle Eastern, in unstable or no housing, or admitted by police, FPR and TPR were both elevated relative to other groups, suggesting the model was calibrated to increase overall predictive accuracy at the expense of higher FPR. Conversely, for other groups like Black patients, models had high FPR and low TPR, suggesting poor overall performance.
Importantly, observational measures of unfairness, such as TPR and FPR are merely outcome measures that do not explain how unfair predictions arise. Rather, these results must be understood in the context of underlying social and structural inequities that can give rise to unfair predictions in the first place, such as racial profiling in the criminal justice system, racial residential segregation, or barriers to accessing mental healthcare29. We discuss some of these parallels in the section below.
Black individuals are less likely to receive adequate outpatient psychiatric treatment, they are more likely to be involuntarily admitted into inpatient treatment, and they may also present with more severe psychotic symptoms, compared to White individuals13,14. Black men in particular face significant barriers in accessing mental health care, and they are more likely to be misdiagnosed with psychotic disorders, as compared to White men16,17,18. Interpersonal bias is also possibility, where structurally reinforced stereotypes may lead to higher risk perceptions for racially marginalized individuals on clinical risk instruments like the DASA, though research is largely inconclusive on whether these instruments are themselves biased. Both male gender and Black race have been found to be significantly associated with violence in psychiatric settings. Findings from our study suggest that these associations can become embedded in clinical datasets, which may lead to unfair treatment by ML algorithms, both via increased false positive predictions and poorer performance in identifying at-risk individuals2,30.
Police apprehension for admission into the ED is also communicated among clinicians to be a relevant factor in risk assessment due to an increased likelihood of aggression in patients admitted involuntarily, and/or referred by the police2,31. It is therefore perhaps not surprising that this mode of admission was associated with the highest FPR of any other predictor in the fairness assessment. Patients apprehended by police for admission into emergency psychiatric care are indeed more likely to become violent or aggressive, which is likely to account for relatively high FPRs and TPRs for this group32. At the same time, racially marginalized and Indigenous groups have increased rates of involuntary admissions into psychiatric care by police, likely due to various factors, such as barriers to accessing mental health care or racial profiling31,33,34,35. This tendency may in part explain the finding of higher FPRs among Black men, and potentially Middle Eastern and Indigenous individuals as well.
The fairness assessment also highlights housing as a potential source of algorithmic unfairness, specifically for those with unstable or supportive forms of housing. On a social level, unstable housing has been associated with psychiatric conditions, such as trauma and substance use, as well as a lower educational attainment and disrupted support networks36,37. Conditions of unstable housing may contribute to food or water insecurity, sleep deprivation, and hyper vigilance, which can lead to the expression of behaviors that are rated as precursors of aggression on clinical instruments, such as the DASA (e.g., irritability, sensitivity to provocation, and unwillingness to follow instructions). Structurally, current psychiatric care systems are not well-equipped to meet the constellation of needs of unhoused individuals, which may contribute to their increased ED use and higher false positive predictions for the risk of violence in inpatient care37,38,39,40,41. Supportive housing services for people with severe mental illness offer more stability, but they are in high demand and extremely under-resourced, often unable to meet complex, individual needs42.
We also identified performance disparities that are not linked to well-researched inequities. For example, while qualitative analyses have shown a general distrust of biomedical mental health services among Middle Eastern individuals, there is a considerable research gap in characterizing how they interact with these systems43. Although our analyses suggest that high FPR for Middle Eastern patients may be in part related to improved model TPR/sensitivity, social and structural determinants likely play a role in the way their risk of violence or aggression is perceived; these may be related to cultural communication barriers, or expressions of distrust manifesting as increased irritability or an unwillingness to follow instructions. However, the gender discrepancy in FPR (but not in TPR) for this group suggests this effect may only extend to men. Similarly, the algorithm displayed modest FPR differences based on citizenship, which is also not a well-documented demographic feature in the psychiatric literature. Nevertheless, citizenship may be an important factor to consider in future fairness assessments of ML models in healthcare, given its impact on access to community, social, and health services.
These findings highlight the importance of thoughtful documentation and processing of demographic data, which is a strength of our study. Specifically, access to high-quality and diverse sociodemographic information is necessary for evaluating ML models for fairness, making it critical that these data are measured or not lost during processing44. Middle Eastern ethnicity, for example, does not appear to be commonly encoded as a unique racial or ethnic category in research datasets, which inevitably precludes the discovery of important trends in this population as identified in our study45. Demographics in our dataset were drawn from CAMH’s health equity form, which was designed to capture a range of rich features which are not frequently characterized, such as specific ethnic and gender minorities46.
Overall, our results suggest that if fairness is not properly considered, the deployment of ML algorithms to support the prediction of aggression in acute psychiatric care and other clinical settings has the potential to cause significant harms with respect to both disparate mistreatment and equalized odds in socially and structurally disadvantaged groups45. Bias in ML algorithms has already been shown to reduce clinician accuracy47; in psychiatric risk assessment, the unwarranted use of interventions based on a false positive prediction can lead to unnecessary distress, disruption of trust in a therapeutic relationship or the health system, and may even precipitate violent or aggressive incidents when they otherwise would not have occurred48. Furthermore, there is extensive literature highlighting the cyclical nature of algorithmic unfairness: algorithms can reproduce and amplify existing inequalities, which can then become embedded in new datasets used to develop ML algorithms or inform care45,49. Even if an unfair recommendation is not followed, disagreement between providers and ML algorithms may lead providers to fear legal implications against them, which may negatively impact care50. Given these concerns, algorithmic unfairness is recognized by both patients and providers as a major barrier in the clinical implementation of predictive risk models50,51.
There exists a range of algorithmic methods to improve a model’s fairness, such as integrating fairness benchmarks into optimization criteria during model training, resampling the input data itself to improve fairness, or enforcing specific fairness criteria using group-specific prediction thresholds52. Several studies have now applied “debiasing” methods to clinical ML algorithms, demonstrating promising results53,54,55. Our findings highlight the necessity to properly assess fairness so that these measures can be applied as appropriate to predictive risk models before they are deployed. An important consideration, however, is that most debiasing methods use the ground truth outcome label as a benchmark to determine whether a model is fair56. In other words, most methods seek to faithfully replicate “the world as it is” – no more, but no less unfair than the input data. However, we have discussed how data relating to inpatient aggression, particularly the administration of coercive interventions, is deeply intertwined with societal inequities. As such, debiasing metrics and methods in this context must use some “true” notion of fairness that represents “the world as it should be”. Algorithmic interventions, therefore, do not constitute a complete solution. To enable algorithmic debiasing approaches, practitioners first must define how a fair and equitable ML algorithm should behave – this is a social question, not a technical one.
Ultimately, ML systems do not operate in a vacuum, but rather as part of highly complex sociotechnical systems where algorithms and societal inequities interact in complex ways. We highlight that ML fairness assessments can identify inequities across large, complex datasets to help target further investigation. However, fairness analysis alone cannot deeply characterize these social and structural drivers of unfairness, nor the exact processes by which they ultimately result in unfair predictions. When seeking to understand algorithmic fairness, therefore, it is important to characterize and understand these biases and inequities on a social level, such as through qualitative approaches that reveal patient and provider experience57,58.
It is also important to note that there is no single optimal way to assess the fairness of ML algorithms. There are over 70 definitions of fairness, many of which are conflicting, making it impossible to simultaneously satisfy all possible definitions59. We restricted our analysis to a single a priori perspective of what constitutes a fair ML model with a focus on disparate mistreatment and equalized odds, making it possible that our analysis missed other relevant fairness considerations or perspectives. For example, in contrast to the group notion of fairness used in this study, individual fairness postulates that similar individuals should receive similar ML predictions, drawing from philosophies of consistency and individual justice rather than anti-discrimination frameworks49,60. Individual fairness often relies on counterfactual or explanation-based ways to define fairness, neither of which were assessed in this study60. Additionally, our investigation examined the fairness of only one model architecture, since our aim was to evaluate models that could be advanced for further testing and implementation (i.e., those performing best on training or validation data). Although underlying societal inequities are likely to impact different types of ML models in similar ways, there is evidence that fairness performance can vary based on model architecture55,61. As such, future research could consider how fairness characteristics differ between model types, and investigate impacts of integrating fairness considerations into model selection itself62,63.
Additionally, there are limitations within the dataset used for this study. Our algorithm was trained using an urban Canadian population—although underlying inequities appear pervasive across populations, our findings may not generalize to other populations64. Moreover, the analysis relied on EHR data which is known to vary in quality. For instance, it is possible that some aggressive incidents were not documented, or modes of admission were mislabelled. Following prior work23, we included restraints in the outcome under the assumption that they were only applied when aggressive incidents were imminent, which may not always hold. Additionally, segmentation of the dataset into subgroups reduced the sample size for the fairness assessment, especially with respect to minority and intersectional groups. For example, limited sample sizes necessitated us to collapse granular descriptions of ethnic heritage into “Black” as a big-bucket category, which may mask additional disparities in ML fairness within this heterogenous group65. Similar limitations were present with gender, as we grouped all genders that were not male or female into a single category, which still lacked the sufficient size to perform intersectional analysis. We were also limited in our exploration of other indicators of socioeconomic status beyond housing, such as income or area-level deprivation and marginalization, which may have offered valuable insights into how these features impact model fairness. As such, we encourage future ML studies in this context to perform fairness assessments, particularly by leveraging rich dataset features, such as granular ethnic breakdowns, larger sample sizes for intersectional groups, and including multiple indicators of socioeconomic status. This will enable a more nuanced and thorough understanding of algorithmic fairness and how it may differ across populations.
In conclusion, ML predictions of aggression in acute psychiatric care and other clinical settings have the potential to be unfairly biased. However, this is not meant to be an argument against the use of ML in such contexts. Rather, we suggest that it is critical to be aware of fairness-related considerations prior to their implementation, and illustrate how performing such analyses can shed light on underlying inequities. To this end, we encourage future ML work in psychiatry to consider fairness as a critical element of evaluation and to conduct further research to interrogate these identified inequities.
