• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Articles
  • Using large language models for extracting and pre-annotating texts on mental health from noisy data in a low-resource language
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 18, 2026
When Pictures Hinder Understanding: Illustrations May Impede Learning of Abstract Ideas
Illustrations can help remember specific actions but do not always make abstract ideas easier to learn. Researchers from HSE University and Humboldt University compared how people learn from texts with different levels of abstractness. They found that participants remembered illustrations better and performed better on related tasks after reading a multimedia text about yoga asanas than after reading an abstract text about the Nash equilibrium. The findings could help improve the selection of illustrations for educational and informational materials. The study has been published in Learning and Instruction.
September 17, 2026
'I Wish That People Would Place Greater Trust in Science'
When Tatiana Eremicheva chose Fundamental and Computational Linguistics as her field of study, she thought it would be about learning languages. Instead, she discovered it was about helping people. In this interview for the HSE Young Scientists project, she discusses science as a way of understanding the world, billiards as a team-building activity, and why learning to read is not always as easy as it seems.
September 15, 2026
Immunity to Chaos: How Personal Resources Help Us Cope with the Challenges of a Turbulent World
International conflicts, crises and digital overload—the modern world puts our minds to the test every day. Traditional psychology often focuses on the consequences: anxiety, depression, and psychosomatic disorders. But what if we looked at the problem differently—through the lens of the resources that prevent us from breaking down? Psychological immunity is precisely this set of resources. Alena Zolotareva and her group, Psychological Immunity as a Resource for Positive Functioning, are developing an integrative model of this phenomenon, adapting diagnostic tools and preparing for large-scale empirical research. Why do psychologists need to collaborate with medical professionals, and how could their research transform preventive care in clinics and corporations?

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Using large language models for extracting and pre-annotating texts on mental health from noisy data in a low-resource language

PeerJ Computer Science, США. 2024. Vol. 10. Article e2395 .
Sergei Koltcov, Surkov A., Koltsova O., Ignatenko V.

 

Recent advancements in large language models (LLMs) have opened new possibilities for developing conversational agents (CAs) in various subfields of mental healthcare. However, this progress is hindered by limited access to high-quality training data, often due to privacy concerns and high annotation costs for low-resource languages. A potential solution is to create human-AI annotation systems that utilize extensive public domain user-to-user and user-to-professional discussions on social media. These discussions, however, are extremely noisy, necessitating the adaptation of LLMs for fully automatic cleaning and pre-classification to reduce human annotation effort. To date, research on LLM-based annotation in the mental health domain is extremely scarce. In this article, we explore the potential of zero-shot classification using four LLMs to select and pre-classify texts into topics representing psychiatric disorders, in order to facilitate the future development of CAs for disorder-specific counseling. We use 64,404 Russian-language texts from online discussion threads labeled with seven most commonly discussed disorders: depression, neurosis, paranoia, anxiety disorder, bipolar disorder, obsessive-compulsive disorder, and borderline personality disorder. Our research shows that while preliminary data filtering using zero-shot technology slightly improves classification, LLM fine-tuning makes a far larger contribution to its quality. Both standard and natural language inference (NLI) modes of fine-tuning increase classification accuracy by more than three times compared to non-fine-tuned training with preliminarily filtered data. Although NLI fine-tuning achieves slightly higher accuracy (0.64) than the standard approach, it is six times slower, indicating a need for further experimentation with NLI hypothesis engineering. Additionally, we demonstrate that lemmatization does not affect classification quality and that multilingual models using texts in their original language perform slightly better than English-only models using automatically translated texts. Finally, we introduce our dataset and model as the first openly available Russian-language resource for developing conversational agents in the domain of mental health counseling.

Research target: Computer Science Psychology
Language: English
Full text
DOI
Text on another site
Keywords: natural language inferencelarge language model (LLM)Большие языковые модели (LLMs)Zero shot classificationPsychological text dataлогический вывод на естественном языкетекстовые психологические данные
Publication based on the results of:
Modelling information and communication behaviour in computer-mediated environments and improving algorithms for behavioural data analysis (2024)
Similar publications
Improving the Accuracy of Automatic Wildlife Detection in Nature Reserves Using Infrared Imaging
Aleksei Samarin, Alexander Savelev, Aleksei Toropov et al., Pattern Recognition and Image Analysis 2026 Vol. 36 No. 2 P. 323–334
In this paper, an improved approach for automatic wildlife detection in natural environments based on the integration of a neural network architecture with a two-stream attention mechanism and a novel preclassification step based on infrared data has been presented. The proposed method addresses one of the key challenges in environmental monitoring: the need for scalable ...
Added: September 19, 2026
IDAP++: Advancing Divergence-Aware Pruning with Joint Filter and Layer Optimization
Aleksei Samarin, Nazarenko A., Kotenko E. et al., Proceedings of the ACM on Management of Data, USA 2026 Vol. 4 No. 1 P. 1–28
Modern knowledge and large volumes of data are increasingly encoded within neural networks, making the task of simplifying their structures and reducing the number of parameters especially relevant, both to improve efficiency and to facilitate deployment in resource-constrained environments. This paper presents a novel approach to neural network compression that addresses redundancy at both the ...
Added: September 19, 2026
Automated Feature Engineering-Based Approach for Micrococci Microscopic Image Classification and Taxonomic Characteristics Determination
Aleksei Samarin, Alexander Savelev, Aleksei Toropov et al., Pattern Recognition and Image Analysis 2025 Vol. 35 No. 2 P. 148–158
This paper describes our research on creating classifiers for microbial images (micrococci microscopy images) obtained from pictures of unfixed microscopic scenes. In our work, we propose an AutoML approach based on the automatic generation and analysis of the feature space for constructing the most optimal descriptors of microorganism images for subsequent classification. This makes it ...
Added: September 19, 2026
Improvement in Microbial Classification Quality Using Synthetic Microscopic Images Generated by Large Visual-Language Models
Aleksei Samarin, Alexander Savelev, Aleksei Toropov et al., Pattern Recognition and Image Analysis 2026 Vol. 36 No. 2 P. 302–312
The lack of annotated microscopic datasets remains a major obstacle to training robust deep learning models for microbial classification. In this paper, a novel data augmentation pipeline that uses visual–linguistic large-scale models to generate synthetic microscopic images of six different bacterial and nonbacterial classes has been proposed. Synthetic samples have gradually been added to the ...
Added: September 19, 2026
Advances in Neural Computation, Machine Learning, and Cognitive Research IX
Springer, Cham, 2026.
computer vision ...
Added: September 19, 2026
Proceedings of 18th International Conference on Machine Learning and Computing
Springer, Cham, 2026.
Added: September 19, 2026
Proceedings of the 35th Conference of Open Innovations Association FRUCT
FRUCT Oy, 2024.
Added: September 19, 2026
Proceedings of the 36th Conference of Open Innovations Association FRUCT
FRUCT Oy, 2024.
Added: September 19, 2026
Proceedings of the 37th Conference of Open Innovations Association FRUCT
FRUCT Oy, 2025.
Added: September 19, 2026
Proceedings of the 39th Conference of Open Innovations Association FRUCT
FRUCT Oy, 2026.
Added: September 19, 2026
Flow-Guided Neural Pruning: Signal-Flow Framework for Multi-Architecture Model Compression
Aleksei Samarin, Nazarenko A., Kotenko E. et al., Machine Learning and Knowledge Extraction 2026 Vol. 8 No. 8 P. 1–26
This paper presents a novel method for pruning deep neural networks based on the concept of flow, derived from the continuous modeling of signal propagation across layers. We derive flow functions for fully connected, convolutional, and self-attention architectures, and we propose a new iterative pruning algorithm, Iterative Flow-Aware Pruning (IFAP), that leverages these measures to ...
Added: September 19, 2026
Толерантность к неопределенности в спортивном туризме: стаж, опыт и навыки автономности на маршруте
Салихова А. А., Polikanova I., Вестник Московского университета. Серия 14: Психология 2026 Т. 49 № 3 С. 9–34
Background. An adaptive attitude toward uncertainty may facilitate goals achievement, however, there is a lack of scientific research on this issue among extreme sport athletes. Sports tourism (ST) is a promising area for investigation of the psychological mechanisms underlying adaptation to extreme stressful environments. Objective. The goal is to establish the connection between tolerance for uncertainty, route experience, autonomy skills, and ...
Added: September 18, 2026
Электрофизиологические корреляты влияния трансформационных психологических методов на совладающее поведение
Сизикова Т. Э., Леонов С. В., Polikanova I., Сибирский психологический журнал 2026 № 101 С. 127–145
This work aims to investigate the impact of a transformational psychological game on changes in psychological parameters of coping behavior and to identify electrophysiological correlates, taking into account the age characteristics of the subjects. The intervention was carried out using the transformational psychological game "Shambhala – 5" (by T.E. Sizikova). The article analyzes psychologist views on transformation, ...
Added: September 18, 2026
Dynamic Pattern Analysis: Method Overview and Trajectory Assessment of Object Development
Myachin A. L., Procedia Computer Science 2026 Vol. 287 P. 193–200
We extend the static pattern analysis method to the temporal dimension by introducing a six-type trajectory taxonomy that classifies objects according to the frequency and structure of pattern switches over an observation window of T > 8 periods. For each object, a reference pattern is designated as the most frequently occupied group over the observation ...
Added: September 18, 2026
Neurocognitive mechanisms of source credibility: how work experience and patient ratings modulate the processing of medical veracity cues
Monahhova E., Morozova A., Gorodnicheva Y. et al., Frontiers in Human Neuroscience 2026 No. 20 Article 1867901
Inroduction:  Source credibility is fundamental to how medical information is processed,yet the underlying neurocognitive mechanisms remain poorly understood. We investigated how two credibility cues — a doctor’s expertise (years of work experience) and aggregate patient ratings (star ratings) — modulate brain responses to veracity cues (‘True’ vs. ‘False’) presented after health- related headlines. Methods:  We recorded EEG from ...
Added: September 18, 2026
FROM LECTURER TO CHATBOT: EVALUATING A HYBRID TEACHING MODEL
Ravedovskaya U., Didenko A., Journal of Teaching English for Specific and Academic Purposes 2025 Vol. 13 No. 3 P. 529–538
Emergence of generative artificial intelligence (GenAI) is revolutionizing teaching and evaluation in higher education. Early adopters have already demonstrated how large language model (LLM) conversational agents can serve as on-demand tutors, yet empirical evidence regarding their effectiveness in facilitating conceptual learning in non-computational domains is scant. Based on constructivist learning theory, this paper presents a ...
Added: September 18, 2026
Light in the dark: Cross-sectional and longitudinal investigation of the network of Dark Triad and Big Five personality traits, resilience and anxiety
Papageorgiou K., Лиханов М. В., Li J. et al., Personality and Individual Differences 2026 Vol. Volume 263 Article 114019
Multivariate approaches are needed to gain a better understanding of the positive and negative impacts of Dark Triad traits within personality networks under different socio-cultural and personal conditions. In three independent studies in two countries (total N of participants = 4700), we applied network analyses and structural equation modeling to investigate the network of personality (Dark Triad ...
Added: September 18, 2026
On the Efficiency of Bounded Multi-Source Shortest Path Algorithm
Громов Р. С., Нестеров Р.А., Proceedings of the Institute for System Programming of the RAS 2026 Vol. 38 No. 4 P. 23–44
This paper explores the performance criteria of the newest algorithm for solving the problem of finding shortest paths on a graph from a given vertex – Bounded Multi-Source Shortest Path Algorithm (BM-SSP). The algorithm was published in 2025 and, as its creators claim, it is asymptotically superior to Dijkstra’s deterministic algorithm. However, in the publication devoted ...
Added: September 18, 2026
Метамотивационные основания существования социальных групп: концептуальный анализ
Tatarko A., Организационная психология 2026 Т. 16 № 2 С. 9–44
Цель. Цель статьи состояла в разработке новой концептуальной модели метамотивационных оснований существования социальных групп, которая позволила бы объяснить их устойчивость, целенаправленное функционирование и формирование коллективной идентичности. Автор стремился выявить и систематизировать ключевые элементы, придающие смысл и стратегическое направление групповому существованию. Подход. Работа основана на синтезе концепций мотивации А. Маслоу, культурно-исторического подхода (Л. С. Выготский, А. Н. Леонтьев), теории ценностей ...
Added: September 17, 2026
Implicit Nature of Visual Statistical Learning: The Importance of Participants’ Goals and Experimental Paradigm
Anton Rogachev, Logvinenko T., Rebreikina A. et al., Cognitive Science 2026 Vol. 50 No. 9 Article 70263
In a recent Letter to the Editor in Cognitive Science, Németh et al. (2026) published a critical commentary in response to our experimental article on the study of visual statistical learning (visual SL) in children aged 3–9 years. The authors present three critical points regarding (1) the inaccurate interpretation of the results of their article, (2) ...
Added: September 17, 2026
A functional systems view on neural tracking of natural speech
Anton Rogachev, Sysoeva O., Frontiers in Systems Neuroscience 2025 Vol. 19 Article 1658243
Added: September 17, 2026
Visual Statistical Learning in Children Aged 3−9 Years
Anton Rogachev, Logvinenko T., Rebreikina A. et al., Cognitive Science 2025 Vol. 49 No. 10 Article e70130
Visual statistical learning (visual SL) is the ability to implicitly extract statistical patterns from visual stimuli. Visual SL could be assessed using online measures, evaluating reaction times (RTs) to stimuli during task performance, and offline measures, which assess recognition of the presented patterns. We examined 96 children aged 3−9 years using a visual SL task ...
Added: September 17, 2026
Proceedings of the 4th Workshop on NLP for Music and Audio (NLP4MusA 2026)
Buzaev F., Mullakhmetov R., Bogachev R. et al., Association for Computational Linguistics, 2026.
Playlist generation based on textual queries using large language models (LLMs) is becoming an important interaction paradigm for music streaming platforms. User queries span a wide spectrum from highly personalized intent to essentially catalog-style requests. Existing systems typically rely on non-personalized retrieval/ranking or apply a fixed level of preference conditioning to every query, which can ...
Added: June 22, 2026
Benchmarking DNA large language models on quadruplexes
Cherednichenko O., Herbert A., Poptsova M., Computational and Structural Biotechnology Journal 2025 Vol. 27 P. 992–1000
Large language models (LLMs) in genomics have successfully predicted various functional genomic elements. While their performance is typically evaluated using genomic benchmark datasets, it remains unclear which LLM is best suited for specific downstream tasks, particularly for generating whole-genome annotations. Current LLMs in genomics fall into three main categories: transformer-based models, long convolution-based models, and state-space models ...
Added: June 19, 2026
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit