kth.sePublications KTH
Change search
Link to record
Permanent link

Direct link
Velupillai, SumithraORCID iD iconorcid.org/0000-0002-4178-2980
Publications (10 of 26) Show all publications
Dutta, R., Gkotsis, G., Velupillai, S., Bakolis, I. & Stewart, R. (2021). Temporal and diurnal variation in social media posts to a suicide support forum. BMC Psychiatry, 21(1), Article ID 259.
Open this publication in new window or tab >>Temporal and diurnal variation in social media posts to a suicide support forum
Show others...
2021 (English)In: BMC Psychiatry, E-ISSN 1471-244X, Vol. 21, no 1, article id 259Article in journal (Refereed) Published
Abstract [en]

BackgroundRates of suicide attempts and deaths are highest on Mondays and these occur more frequently in the morning or early afternoon, suggesting weekly temporal and diurnal variation in suicidal behaviour. It is unknown whether there are similar time trends on social media, of posts relevant to suicide. We aimed to determine temporal and diurnal variation in posting patterns on the Reddit forum SuicideWatch, an online community for individuals who might be at risk of, or who know someone at risk of suicide.MethodsWe used time series analysis to compare date and time stamps of 90,518 SuicideWatch posts from 1st December 2008 to 31st August 2015 to (i) 6,616,431 posts on the most commonly subscribed general subreddit, AskReddit and (ii) 66,934 of these AskReddit posts, which were posted by the SuicideWatch authors.ResultsMondays showed the highest proportion of posts on SuicideWatch. Clear diurnal variation was observed, with a peak in the early morning (2:00-5:00h), and a subsequent decrease to a trough in late morning/early afternoon (11:00-14:00h). Conversely, the highest volume of posts in the control data was between 20:00-23:00h.ConclusionsPosts on SuicideWatch occurred most frequently on Mondays: the day most associated with suicide risk. The early morning peak in SuicideWatch posts precedes the time of day during which suicide attempts and deaths most commonly occur. Further research of these weekly and diurnal rhythms should help target populations with support and suicide prevention interventions when needed most.

Place, publisher, year, edition, pages
Springer Nature, 2021
Keywords
Suicide, Suicide timing, Social media, Diurnal rhythms, Temporal pattern
National Category
Cardiology and Cardiovascular Disease
Identifiers
urn:nbn:se:kth:diva-298111 (URN)10.1186/s12888-021-03268-1 (DOI)000658331200001 ()34011346 (PubMedID)2-s2.0-85106291746 (Scopus ID)
Note

QC 20210629

Available from: 2021-06-29 Created: 2021-06-29 Last updated: 2025-02-10Bibliographically approved
Viani, N., Tissot, H., Bernardino, A. & Velupillai, S. (2019). Annotating Temporal Information in Clinical Notes for Timeline Reconstruction: Towards the Definition of Calendar Expressions. In: SIGBIOMED WORKSHOP ON BIOMEDICAL NATURAL LANGUAGE PROCESSING (BIONLP 2019): . Paper presented at 18th SIGBioMed Workshop on Biomedical Natural Language Processing (BioNLP); Florence, ITALY; AUG 01, 2019 (pp. 201-210). ASSOC COMPUTATIONAL LINGUISTICS-ACL
Open this publication in new window or tab >>Annotating Temporal Information in Clinical Notes for Timeline Reconstruction: Towards the Definition of Calendar Expressions
2019 (English)In: SIGBIOMED WORKSHOP ON BIOMEDICAL NATURAL LANGUAGE PROCESSING (BIONLP 2019), ASSOC COMPUTATIONAL LINGUISTICS-ACL , 2019, p. 201-210Conference paper, Published paper (Refereed)
Abstract [en]

To automatically analyse complex trajectory information enclosed in clinical text (e.g. timing of symptoms, duration of treatment), it is important to understand the related temporal aspects, anchoring each event on an absolute point in time. In the clinical domain, few temporally annotated corpora are currently available. Moreover, underlying annotation schemas - which mainly rely on the TimeML standard - are not necessarily easily applicable for applications such as patient timeline reconstruction. In this work, we investigated how temporal information is documented in clinical text by annotating a corpus of medical reports with time expressions (TIMEXes), based on TimeML. The developed corpus is available to the NLP community. Starting from our annotations, we analysed the suitability of the TimeML TIMEX schema for capturing timeline information, identifying challenges and possible solutions. As a result, we propose a novel annotation schema that could be useful for timeline reconstruction: CALendar EXpression (CALEX).

Place, publisher, year, edition, pages
ASSOC COMPUTATIONAL LINGUISTICS-ACL, 2019
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-272684 (URN)000521946800021 ()2-s2.0-85094861470 (Scopus ID)
Conference
18th SIGBioMed Workshop on Biomedical Natural Language Processing (BioNLP); Florence, ITALY; AUG 01, 2019
Note

QC 20200512

Part of proceedings: ISBN 978-1-950737-28-4

Available from: 2020-05-12 Created: 2020-05-12 Last updated: 2025-02-07Bibliographically approved
Viani, N., Kam, J., Yin, L., Verma, S., Stewart, R., Patel, R. & Velupillai, S. (2019). Annotating temporal relations to determine the onset of psychosis symptoms. In: 17th World Congress on Medical and Health Informatics, MEDINFO 2019: . Paper presented at 17th World Congress on Medical and Health Informatics, MEDINFO 2019; Lyon; France; 25 August 2019 through 30 August 2019 (pp. 418-422). IOS Press
Open this publication in new window or tab >>Annotating temporal relations to determine the onset of psychosis symptoms
Show others...
2019 (English)In: 17th World Congress on Medical and Health Informatics, MEDINFO 2019, IOS Press, 2019, p. 418-422Conference paper, Published paper (Refereed)
Abstract [en]

For patients with a diagnosis of schizophrenia, determining symptom onset is crucial for timely and successful intervention. In mental health records, information about early symptoms is often documented only in free text, and thus needs to be extracted to support clinical research. To achieve this, natural language processing (NLP) methods can be used. Development and evaluation of NLP systems requires manually annotated corpora. We present a corpus of mental health records annotated with temporal relations for psychosis symptoms. We propose a methodology for document selection and manual annotation to detect symptom onset information, and develop an annotated corpus. To assess the utility of the created corpus, we propose a pilot NLP system. To the best of our knowledge, this is the first temporally-annotated corpus tailored to a specific clinical use-case.

Place, publisher, year, edition, pages
IOS Press, 2019
Series
Studies in Health Technology and Informatics, ISSN 0926-9630 ; 264
Keywords
Electronic Health Records, Natural Language Processing, Schizophrenia
National Category
Other Health Sciences
Identifiers
urn:nbn:se:kth:diva-262523 (URN)10.3233/SHTI190255 (DOI)000569653400084 ()31437957 (PubMedID)2-s2.0-85071455534 (Scopus ID)
Conference
17th World Congress on Medical and Health Informatics, MEDINFO 2019; Lyon; France; 25 August 2019 through 30 August 2019
Note

QC 20191028

Part of ISBN 9781643680026

Available from: 2019-10-28 Created: 2019-10-28 Last updated: 2024-10-15Bibliographically approved
Velupillai, S., Epstein, S., Bittar, A., Stephenson, T., Dutta, R. & Downs, J. (2019). Identifying suicidal adolescents from mental health records using natural language processing. In: 17th World Congress on Medical and Health Informatics, MEDINFO 2019: . Paper presented at 17th World Congress on Medical and Health Informatics, MEDINFO 2019; Lyon; France; 25 August 2019 through 30 August 2019 (pp. 413-417). IOS Press, 264
Open this publication in new window or tab >>Identifying suicidal adolescents from mental health records using natural language processing
Show others...
2019 (English)In: 17th World Congress on Medical and Health Informatics, MEDINFO 2019, IOS Press, 2019, Vol. 264, p. 413-417Conference paper, Published paper (Refereed)
Abstract [en]

Suicidal ideation is a risk factor for self-harm, completed suicide and can be indicative of mental health issues. Adolescents are a particularly vulnerable group, but few studies have examined suicidal behaviour prevalence in large cohorts. Electronic Health Records (EHRs) are a rich source of secondary health care data that could be used to estimate prevalence. Most EHR documentation related to suicide risk is written in free text, thus requiring Natural Language Processing (NLP) approaches. We adapted and evaluated a simple lexicon- and rule-based NLP approach to identify suicidal adolescents from a large EHR database. We developed a comprehensive manually annotated EHR reference standard and assessed NLP performance at both document and patient level on data from 200 patients (~5000 documents). We achieved promising results (>80% f1 score at both document and patient level). Simple NLP approaches can be successfully used to identify patients who exhibit suicidal risk behaviour, and our proposed approach could be useful for other populations and settings.

Place, publisher, year, edition, pages
IOS Press, 2019
Series
Studies in Health Technology and Informatics, ISSN 0926-9630 ; 264
Keywords
Electronic Health Records, Natural Language Processing, Suicide
National Category
Other Health Sciences
Identifiers
urn:nbn:se:kth:diva-262522 (URN)10.3233/SHTI190254 (DOI)000569653400083 ()31437956 (PubMedID)2-s2.0-85071484866 (Scopus ID)
Conference
17th World Congress on Medical and Health Informatics, MEDINFO 2019; Lyon; France; 25 August 2019 through 30 August 2019
Note

QC 20191028

Part of ISBN 9781643680026

Available from: 2019-10-28 Created: 2019-10-28 Last updated: 2024-10-25Bibliographically approved
Wang, Z., Ive, J., Velupillai, S. & Specia, L. (2019). Is artificial data useful for biomedical Natural Language Processing algorithms?. In: SIGBIOMED WORKSHOP ON BIOMEDICAL NATURAL LANGUAGE PROCESSING (BIONLP 2019): . Paper presented at 18th SIGBioMed Workshop on Biomedical Natural Language Processing (BioNLP), Florence, ITALY, AUG 1, 2019 (pp. 240-249). Association for Computational Linguistics (ACL)
Open this publication in new window or tab >>Is artificial data useful for biomedical Natural Language Processing algorithms?
2019 (English)In: SIGBIOMED WORKSHOP ON BIOMEDICAL NATURAL LANGUAGE PROCESSING (BIONLP 2019), Association for Computational Linguistics (ACL) , 2019, p. 240-249Conference paper, Published paper (Refereed)
Abstract [en]

A major obstacle to the development of Natural Language Processing (NLP) methods in the biomedical domain is data accessibility. This problem can be addressed by generating medical data artificially. Most previous studies have focused on the generation of short clinical text, and evaluation of the data utility has been limited. We propose a generic methodology to guide the generation of clinical text with key phrases. We use the artificial data as additional training data in two key biomedical NLP tasks: text classification and temporal relation extraction. We show that artificially generated training data used in conjunction with real training data can lead to performance boosts for data-greedy neural network algorithms. We also demonstrate the usefulness of the generated data for NLP setups where it fully replaces real training data.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL), 2019
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-272687 (URN)10.18653/v1/W19-5026 (DOI)000521946800026 ()2-s2.0-85096602481 (Scopus ID)
Conference
18th SIGBioMed Workshop on Biomedical Natural Language Processing (BioNLP), Florence, ITALY, AUG 1, 2019
Note

Part of ISBN 9781950737284 

QC 20260714

Available from: 2020-05-12 Created: 2020-05-12 Last updated: 2026-07-14Bibliographically approved
Velupillai, S., Hadlaczky, G., Baca-Garcia, E., Gorrell, G. M., Werbeloff, N., Nguyen, D., . . . Dutta, R. (2019). Risk Assessment Tools and Data-Driven Approaches for Predicting and Preventing Suicidal Behavior. Frontiers in Psychiatry, 10, Article ID 36.
Open this publication in new window or tab >>Risk Assessment Tools and Data-Driven Approaches for Predicting and Preventing Suicidal Behavior
Show others...
2019 (English)In: Frontiers in Psychiatry, E-ISSN 1664-0640, Vol. 10, article id 36Article in journal (Refereed) Published
Abstract [en]

Risk assessment of suicidal behavior is a time-consuming but notoriously inaccurate activity for mental health services globally. In the last 50 years a large number of tools have been designed for suicide risk assessment, and tested in a wide variety of populations, but studies show that these tools suffer from low positive predictive values. More recently, advances in research fields such as machine learning and natural language processing applied on large datasets have shown promising results for health care, and may enable an important shift in advancing precision medicine. In this conceptual review, we discuss established risk assessment tools and examples of novel data-driven approaches that have been used for identification of suicidal behavior and risk. We provide a perspective on the strengths and weaknesses of these applications to mental health-related data, and suggest research directions to enable improvement in clinical practice.

Place, publisher, year, edition, pages
FRONTIERS MEDIA SA, 2019
Keywords
suicide risk prediction, suicidality, suicide risk assessment, clinical informatics, machine learning, natural language processing
National Category
Health Care Service and Management, Health Policy and Services and Health Economy
Identifiers
urn:nbn:se:kth:diva-245137 (URN)10.3389/fpsyt.2019.00036 (DOI)000458715600001 ()30814958 (PubMedID)2-s2.0-85062732148 (Scopus ID)
Note

QC 20190313

Available from: 2019-03-13 Created: 2019-03-13 Last updated: 2024-01-17Bibliographically approved
Bittar, A., Velupillai, S., Roberts, A. & Dutta, R. (2019). Text classification to inform suicide risk assessment in electronic health records. In: 17th World Congress on Medical and Health Informatics, MEDINFO 2019: . Paper presented at 17th World Congress on Medical and Health Informatics, MEDINFO 2019; Lyon; France; 25 August 2019 through 30 August 2019 (pp. 40-44). IOS Press, 264
Open this publication in new window or tab >>Text classification to inform suicide risk assessment in electronic health records
2019 (English)In: 17th World Congress on Medical and Health Informatics, MEDINFO 2019, IOS Press, 2019, Vol. 264, p. 40-44Conference paper, Published paper (Refereed)
Abstract [en]

Assessing a patient's risk of an impending suicide attempt has been hampered by limited information about dynamic factors that change rapidly in the days leading up to an attempt. The storage of patient data in electronic health records (EHRs) has facilitated population-level risk assessment studies using machine learning techniques. Until recently, most such work has used only structured EHR data and excluded the unstructured text of clinical notes. In this article, we describe our experiments on suicide risk assessment, modelling the problem as a classification task. Given the wealth of text data in mental health EHRs, we aimed to assess the impact of using this data in distinguishing periods prior to a suicide attempt from those not preceding such an attempt. We compare three different feature sets, one structured and two text-based, and show that inclusion of text features significantly improves classification accuracy in suicide risk assessment.

Place, publisher, year, edition, pages
IOS Press, 2019
Series
Studies in Health Technology and Informatics, ISSN 0926-9630 ; 264
Keywords
Natural Language Processing, Risk Assessment, Suicide
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:kth:diva-262521 (URN)10.3233/SHTI190179 (DOI)000569653400008 ()31437881 (PubMedID)2-s2.0-85071499215 (Scopus ID)
Conference
17th World Congress on Medical and Health Informatics, MEDINFO 2019; Lyon; France; 25 August 2019 through 30 August 2019
Note

QC 20191018

Part of ISBN 9781643680026

Available from: 2019-10-18 Created: 2019-10-18 Last updated: 2024-10-23Bibliographically approved
Neveol, A., Dalianis, H., Velupillai, S., Savova, G. & Zweigenbaum, P. (2018). Clinical Natural Language Processing in languages other than English: opportunities and challenges. Journal of Biomedical Semantics, 9, Article ID 12.
Open this publication in new window or tab >>Clinical Natural Language Processing in languages other than English: opportunities and challenges
Show others...
2018 (English)In: Journal of Biomedical Semantics, E-ISSN 2041-1480, Vol. 9, article id 12Article, review/survey (Refereed) Published
Abstract [en]

Background: Natural language processing applied to clinical text or aimed at a clinical outcome has been thriving in recent years. This paper offers the first broad overview of clinical Natural Language Processing (NLP) for languages other than English. Recent studies are summarized to offer insights and outline opportunities in this area. Main Body: We envision three groups of intended readers: (1) NLP researchers leveraging experience gained in other languages, (2) NLP researchers faced with establishing clinical text processing in a language other than English, and (3) clinical informatics researchers and practitioners looking for resources in their languages in order to apply NLP techniques and tools to clinical practice and/or investigation. We review work in clinical NLP in languages other than English. We classify these studies into three groups: (i) studies describing the development of new NLP systems or components de novo, (ii) studies describing the adaptation of NLP architectures developed for English to another language, and (iii) studies focusing on a particular clinical application. Conclusion: We show the advantages and drawbacks of each method, and highlight the appropriate application context. Finally, we identify major challenges and opportunities that will affect the impact of NLP on clinical practice and public health studies in a context that encompasses English as well as other languages.

Place, publisher, year, edition, pages
BioMed Central, 2018
Keywords
Natural Language Processing, Clinical Decision-Making, Languages other than English
National Category
Bioinformatics (Computational Biology)
Identifiers
urn:nbn:se:kth:diva-226788 (URN)10.1186/s13326-018-0179-8 (DOI)000428928300001 ()29602312 (PubMedID)2-s2.0-85044782945 (Scopus ID)
Funder
Swedish Research Council, 2015-00359
Note

QC 20180426

Available from: 2018-04-26 Created: 2018-04-26 Last updated: 2024-03-18Bibliographically approved
Ive, J., Gkotsis, G., Dutta, R., Stewart, R. & Velupillai, S. (2018). Hierarchical neural model with attention mechanisms for the classification of social media text related to mental health. In: Proceedings of the 5th Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic, CLPsych 2018 at the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HTL 2018: . Paper presented at 5th Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic, CLPsych 2018 at the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HTL 2018, New Orleans, United States (pp. 69-77). Association for Computational Linguistics (ACL)
Open this publication in new window or tab >>Hierarchical neural model with attention mechanisms for the classification of social media text related to mental health
Show others...
2018 (English)In: Proceedings of the 5th Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic, CLPsych 2018 at the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HTL 2018, Association for Computational Linguistics (ACL) , 2018, p. 69-77Conference paper, Published paper (Refereed)
Abstract [en]

Mental health problems represent a major public health challenge. Automated analysis of text related to mental health is aimed to help medical decision-making, public health policies and to improve health care. Such analysis may involve text classification. Traditionally, automated classification has been performed mainly using machine learning methods involving costly feature engineering. Recently, the performance of those methods has been dramatically improved by neural methods. However, mainly Convolutional neural networks (CNNs) have been explored. In this paper, we apply a hierarchical Recurrent neural network (RNN) architecture with an attention mechanism on social media data related to mental health. We show that this architecture improves overall classification results as compared to previously reported results on the same data. Benefitting from the attention mechanism, it can also efficiently select text elements crucial for classification decisions, which can also be used for in-depth analysis.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL), 2018
National Category
Information Systems
Identifiers
urn:nbn:se:kth:diva-385470 (URN)2-s2.0-85061048924 (Scopus ID)
Conference
5th Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic, CLPsych 2018 at the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HTL 2018, New Orleans, United States
Note

Part of ISBN 9781948087124

QC 20260714

Available from: 2026-07-14 Created: 2026-07-14 Last updated: 2026-07-14Bibliographically approved
Ive, J., Viani, N., Chandran, D., Bittar, A. & Velupillai, S. (2018). KCL-Health-NLP@CLEF eHealth 2018 Task 1: ICD-10 coding of French and Italian death certificates with character-level convolutional neural networks. In: CEUR Workshop Proceedings: . Paper presented at 19th Working Notes of CLEF Conference and Labs of the Evaluation Forum, CLEF 2018, Avignon, France, 10 September 2018 through 14 September 2018. CEUR-WS, 2125
Open this publication in new window or tab >>KCL-Health-NLP@CLEF eHealth 2018 Task 1: ICD-10 coding of French and Italian death certificates with character-level convolutional neural networks
Show others...
2018 (English)In: CEUR Workshop Proceedings, CEUR-WS , 2018, Vol. 2125Conference paper, Published paper (Refereed)
Abstract [en]

In this paper we describe the participation of the KCL-Health-NLP team in the CLEF eHealth 2018 lab, specifically Task 1: Multilingual Information Extraction-ICD10 coding. The task involves the automatic coding of causes of death in death certificates in French, Italian and Hungarian according to the ICD-10 taxonomy. Choosing to work on the two Romance languages, we treated the task as a sequence-to-sequence prediction problem. Our system has an encoder-decoder architecture, with convolutional neural networks based on character em-beddings as encoders and recurrent neural network decoders. Our hypothesis was that a character-level representation would allow our model to generalise across two genealogically related languages. Results obtained by pre-training our Italian model on the French data set confirmed this intuition. We also explored the impact of character-level features extracted from dictionary-matched ICD codes. We obtained F-measures of 0.72/0.64 and 0.78 on the French aligned/raw and Italian raw internal test data, respectively. On the blind test set released by the task organisers, our top results were 0.65/0.52 and 0.69 F-measure, respectively.

Place, publisher, year, edition, pages
CEUR-WS, 2018
Series
CEUR Workshop Proceedings, ISSN 1613-0073
Keywords
Convolutional neural networks, Encoder-decoder architecture, Recurrent neural networks
National Category
Other Computer and Information Science
Identifiers
urn:nbn:se:kth:diva-234052 (URN)2-s2.0-85051060301 (Scopus ID)
Conference
19th Working Notes of CLEF Conference and Labs of the Evaluation Forum, CLEF 2018, Avignon, France, 10 September 2018 through 14 September 2018
Note

QC 20180906

Available from: 2018-09-06 Created: 2018-09-06 Last updated: 2022-06-26Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-4178-2980

Search in DiVA

Show all publications