Published May 4, 2019 | Version v1
Conference paper Open

Reddit C-SSRS Suicide Dataset

  • 1. Kno.e.sis Center
  • 2. University of Kentucky
  • 3. Department of Psychiatry, Wright State University
  • 4. Cornell University

Description

Knowledge-aware Assessment of Severity of Suicide Risk for Early Intervention

Mental health illness such as depression is a significant risk factor for suicide ideation, behaviors, and attempts. A report by Substance Abuse and Mental Health Services Administration (SAMHSA) shows that 80% of the patients suffering from Borderline Personality Disorder (BPD) have suicidal behavior, 5-10% of whom commit suicide. While multiple initiatives have been developed and implemented for suicide prevention, a key challenge has been the social stigma associated with mental disorders, which deters patients from seeking help or sharing their experiences directly with others including clinicians. This is particularly true for teenagers and younger adults where suicide is the second highest cause of death in the US Prior research involving surveys and questionnaires (e.g. PHQ-9) for suicide risk prediction failed to provide a quantitative assessment of risk that informed timely clinical decision-making for intervention. Our interdisciplinary study concerns the use of Reddit as an unobtrusive data source for gleaning information about suicidal tendencies and other related mental health conditions afflicting depressed users. We provide details of our learning framework that incorporates domain-specific knowledge to predict the severity of suicide risk for an individual. Our approach involves developing a suicide risk severity lexicon using medical knowledge bases and suicide ontology to detect cues relevant to suicidal thoughts and actions. We also use language modeling, medical entity recognition, and normalization and negation detection to create a dataset of 2181 redditors that have discussed or implied suicidal ideation, behavior, or attempt. Given the importance of clinical knowledge, our gold standard dataset of 500 redditors (out of 2181) was developed by four practicing psychiatrists following the guidelines outlined in Columbia.Suicide Severity Rating Scale (C-SSRS), with the pairwise annotator agreement of 0.79 and group-wise agreement of 0.73. Compared to the existing four-label classification scheme (no risk, low risk, moderate risk, and high risk), our proposed C-SSRS-based 5-label classification scheme distinguishes people who are supportive, from those who show different severity of suicidal tendency. Our 5-label classification scheme outperforms the state-of-the-art schemes by improving the graded recall by 4.2% and reducing the perceived risk measure by 12.5%. Convolutional neural network (CNN) provided the best performance in our scheme due to the discriminative features and use of domain-specific knowledge resources, in comparison to SVM-L that has been used in the state-of-the-art tools over similar dataset

Files

500_Reddit_users_posts_labels.csv

Files (3.7 MB)

Name Size Download all
md5:be0addddd2bb64129d866ac41b625d04
3.6 MB Preview Download
md5:6f6b6d2b33ead618a4cb7bffbbd528a0
4.2 kB Preview Download
md5:88b5a7dfca9f6574204de8a5934f2127
3.6 kB Preview Download
md5:17f6cacf1deae66201f9469706f36de0
14.9 kB Preview Download
md5:1c96b872c7e305aba3f898ee84770c25
53.4 kB Preview Download

Additional details

Related works

References
10.1145/3308558.3313698 (DOI)