Open this publication in new window or tab >>Show others...
2025 (English)In: 2025 IEEE 12th International Conference on Data Science and Advanced Analytics, DSAA 2025, Institute of Electrical and Electronics Engineers (IEEE) , 2025Conference paper, Published paper (Refereed)
Abstract [en]
A careful analysis of the Large Language Model (LLM) results, generated through anonymized representations of the original dataset, is crucial to precisely evaluate the data-sharing procedure's limitations and facilitate valuable collaborations among Internet-based cognitive behavioral therapy (ICBT) companies and third parties. This paper presents an experimental study of fine-Tuning 27 LMs for a multiclass classification task to identify depression severity using 40,191 tweets labeled by human annotators. We fine-Tune 14 Bidirectional Encoder Representations from Transformers (BERT), 6 Robustly Optimized BERT Pretraining Approaches (RoBerta), 3 Generative Pretraining (GPT), and 4 Text-To-Text Transfer Transformer (T5) based LMs to classify confidential and anonymized tweets. We report that T5, through conditional generation, outperforms widely adopted BERT, RoBerta, and GPT types for classifying confidential and anonymized tweets. Anonymizing personal information safeguards user privacy and often increases LM performance. Case sensitivity can potentially improve or harm the performance of domain-specific LMs for original and anonymized text.
Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Keywords
Depression Severity, Google T5, Internet-based Cognitive Behavioral Therapy, Robust LMs
National Category
Natural Language Processing
Identifiers
urn:nbn:se:kth:diva-377819 (URN)10.1109/DSAA65442.2025.11247966 (DOI)2-s2.0-105029896486 (Scopus ID)
Conference
12th IEEE International Conference on Data Science and Advanced Analytics, DSAA 2025, Birmingham, United Kingdom of Great Britain, Oct 9 2025 - Oct 12 2025
Note
Part of ISBN 979-8-3315-1179-1
QC 20260310
2026-03-102026-03-102026-03-10Bibliographically approved