• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Bridging Gaps in Russian Language Processing: AI and Everyday Conversations
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 22, 2026
Personal Interest in Doctoral Thesis Topic Most Important for Confidence in Successful Defence
A researcher at HSE University analysed data on 1,539 doctoral students from 161 Russian universities to identify which features of a thesis topic are associated with academic success and engagement. The most important factor was found to be personal interest in the research topic, which was associated with almost all key aspects of doctoral programme experience—from engaging with the academic supervisor to research activity and confidence about successfully defending the thesis. The findings have been published in Higher Education.
September 21, 2026
Researchers Develop Methodology to Assess the Quality of Legal Representation in Criminal Proceedings
Having a good defence attorney in criminal proceedings can largely determine whether a defendant retains their freedom, health and good name. Researchers at HSE University propose a method for predicting an attorney’s performance based on the outcomes of their previous cases. The methodology takes into account the severity of the charges, the complexity of the cases, and the most likely outcome, drawing on judicial statistics.
September 21, 2026
Algebra, Geometry, and AI: Russian and Vietnamese Mathematicians Discuss Current Research
A delegation of scientists from Hanoi visited the HSE Faculty of Computer Science and then took part in a Russian-Vietnamese conference in St Petersburg. The events were part of the three-year project ‘Flexibility and Computational Methods.’ Over the course of the project, the researchers have prepared joint publications and obtained new mathematical results.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Bridging Gaps in Russian Language Processing: AI and Everyday Conversations

P. 253–258.
Tatiana Sherstinova, Nikolay Mikhaylovskiy, Evgenia Kolpashchikova, Violetta Kruglikova

Contemporary advancements in NLP and neural network techniques are paving the way to enhance and harness traditional linguistic resources and corpora, as well as expand the methods of applying neural networks for complex language material. Thus, a weak point for both theoretical and applied linguistic tasks is the processing of spontaneous everyday speech. Two experiments described in this article are dedicated to the analysis of how successfully modern neural models cope with the recognition and generation of everyday Russian speech. The material for the experiments is the well-known ORD speech corpus, the largest collection of professional and mundane dialogues in Russian. The first experiment targets the pressing issue of increasing the volume of transcribed speech data through state-of-the-art automatic speech recognition techniques. Experimental recognition was conducted using two diverse methods – the NTR Acoustic Model and OpenAI's Whisper system. The second experiment zeroes in on refining generative language models tailored for Russian using a conversational dataset. A prototype dialogue system, derived from the enhanced ruGPT-3 Small model, exemplifies the transformative potential of fine-tuning in dialogue generation tasks. The acquired results are utilized to enrich datasets for recognizing everyday Russian speech and for constructing chatbots that emulate spontaneous Russian conversations.

Language: English
Full text
DOI
Text on another site
Keywords: neural networksNLPeveryday speechautomatic speech recognition
Publication based on the results of:
Text as Big Data: methods and models for big text data analysis (2024)

In book

Proceedings of the 35th Conference of Open Innovations Association FRUCT, 24-26 April 2024, Tampere, Finland
Issue 1. , FRUCT Oy, 2024.
Similar publications
A Model Based on Universal Filters for Image Color Correction
Aleksei Samarin, Nazarenko A., Alexander Savelev et al., Pattern Recognition and Image Analysis 2024 Vol. 34 No. 3 P. 844–854
Improving image quality is becoming an increasingly popular task, especially when working with mobile devices. One common approach to image enhancement is the use of convolutional neural networks. However, to achieve good results, such networks must be large enough, otherwise there is a risk of unwanted artifacts. In addition, large convolutional neural networks require significant ...
Added: September 21, 2026
Filter-Based Preprocessing Neural Network Model for Microorganism Detection Improvement
Aleksei Samarin, Aleksei Toropov, Alexander Savelev et al., , in: Pattern Recognition. ICPR 2024 International Workshops and Challenges.: Springer, Cham, 2025. P. 308–320.
This research explores an innovative approach to enhancing the accuracy of detecting small microorganisms in complex microscopic environments. Our study introduces a streamlined, hybrid image pre-processing model specifically designed to address the challenges of identifying diplococci in live microscopy of dynamic samples. By integrating pre-defined filtering techniques with predictive adjustments for optimal applicability, our method ...
Added: September 21, 2026
Модель глубокого обучения для автоматизированной интерпретации медицинских электрофизиологических данных
Lebedev O. B., Шмелева А. Г., Гежа Н. С., Информатика и автоматизация (Труды СПИИРАН) 2026 Т. 25 № 3 С. 720–750
This paper describes the development of a neural network model for automated analysis of medical data in electrophysiology based on deep learning methods. The relevance of this work stems from the growing need to improve the objectivity, speed, and accuracy of processing complex spatiotemporal signals, such as ECG or EEG. Convolutional neural networks (CNNs), which ...
Added: September 10, 2026
Scalable machine learning approach to disordered s-wave superconductors
Неверов В. Д., Красавин А. В., Vagov A. et al., Physical Review B: Condensed Matter and Materials Physics 2026 Vol. 113 P. 1–6
We develop a neural network approach to solve the self-consistent Bogoliubov-de Gennes equations in strongly disordered s-wave superconductors. The method accurately reproduces inhomogeneous gap distributions and generalizes to system sizes far larger than those used in training. It reduces computational scaling from O(N6 ) to O(N2), enabling quantitative analysis of percolation phenomena and the superconductor-insulator ...
Added: September 5, 2026
Proceedings of the 43rd International Conference on Machine Learning (ICML 2026)
Seul: PMLR, 2026.
Added: June 4, 2026
Измерение ИИ-грамотности взрослых россиян: методика и результаты
Davydov S. G., Федоров В. В., Социологические исследования 2026 № 5 С. 141–147
Представлены результаты измерения ИИ-грамотности взрослого населения России. Исследование решает проблему отсутствия эмпирических данных о фактическом уровне владения компетенциями в сфере искусственного интеллекта среди граждан. Методика основана на самооценке владения пятью типами ИИ-инструментов по 5‑балльной шкале и последующем индексировании. Сбор информации осуществлен методом телефонного опроса (CATI) на общероссийской выборке проекта «ВЦИОМ–Спутник» (N = 1600). Выявлен уровень ...
Added: May 13, 2026
Granular computing-based deep learning for text classification
Behzadidoost R., Mahan F., Izadkhah H., Information Sciences 2024 Vol. 652 Article 119746
Granular computing involves a comprehensive process that encompasses theories, methodologies, and techniques to solve complex problems, rather than being just an algorithm. As the volume of generated data continues to grow rapidly, data-driven problems have become increasingly complex. Although deep learning models have outperformed traditional machine learning models in solving complex problems, there is still room for enhancing their performance. ...
Added: March 12, 2026
30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, Kanazawa, Japan, July 4–6, 2025, Proceedings, Part I. Natural Language Processing and Information Systems. (LNCS, volume 15836)
Springer, 2025.
The two-volume set LNCS 15836 and 15837 constitutes the proceedings of the 30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, held in Kanazawa, Japan, during July 4–6, 2025. The 33 full papers, 19 short papers and 2 demo papers presented in this volume were carefully reviewed and selected from 120 submissions. ...
Added: February 3, 2026
Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021)
INCOMA Ltd, 2021.
Added: January 28, 2026
Screen-Cam Imitation Module for Improving Data Hiding Robustness
Dzhanashia K., Aleksandr Fedosov, Oleg Evsutin, Sensors 2025 Vol. 25 No. 23 Article 7726
Using an attack-simulation module is a well-recognized approach to improving the robustness of end-to-end neural-network-based data-hiding schemes. However, most proposed attack simulators are limited in the types of attacks they cover, usually handling only a basic set of digital transformations. Real, in-demand use cases for data-hiding methods may involve modifications that cannot be modeled by ...
Added: November 28, 2025
Understanding the training dynamics of CoLaNET by its simplified model
O.A. Goryunov, Maslennikov O. V., Kiselev M. V. et al., Chaos, Solitons and Fractals 2026 Vol. 203 Article 117663
Training complex, biologically plausible Spiking Neural Networks (SNNs) with local learning rules is a significant challenge for theoretical analysis. Here we address this problem by developing a comprehensive analytical theory for the learning dynamics of CoLaNET, a recently proposed columnar SNN. In particular, we consider a simplified model that captures the core algorithmic logic of ...
Added: November 28, 2025
Смежные права на результаты интеллектуальной деятельности, созданные искусственным интеллектом: философско-правовой анализ замены критерия творчества на критерий инвестиций
Pakshin P., Актуальные проблемы российского права 2025 Т. 20 № 11 С. 11–18
The paper substantiates the necessity of providing legal protection for the results of intellectual works created by artificial intelligence through the mechanism of related rights. It examines ways to reduce legal risks associated with the creation of intellectual property using artificial intelligence technologies and offers a philosophical and legal analysis of the proposed hypothesis, namely, ...
Added: November 27, 2025
Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025)
Association for Computational Linguistics, 2025.
Added: November 17, 2025
2025 International Joint Conference on Neural Networks (IJCNN)
IEEE, 2025.
Added: November 15, 2025
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
Anton R., Mikhalchuk M., Rahmatullaev T. et al., , in: Findings of the Association for Computational Linguistics: NAACL 2025.: Association for Computational Linguistics, 2025. P. 7757–7764.
We introduce methods to quantify how Large Language Models (LLMs) encode and store contextual information, revealing that tokens often seen as minor (e.g., determiners, punctuation) carry surprisingly high context. Notably, removing these tokens — especially stopwords, articles, and commas — consistently degrades performance on MMLU and BABILong-4k, even if removing only irrelevant tokens. Our analysis ...
Added: November 6, 2025
Segmentation of Vertebral Arteries on the MR Images
Prikhodko R., Moshkin A., Romanov A., , in: 2025 International Russian Automation Conference (RusAutoCon).: IEEE, 2025. P. 273–278.
The vertebral arteries are one of the most important sources of blood supply to the brain, therefore any pathological changes in them can be the reason behind serious diseases. Magnetic Resonance Imaging (MRI) allows diagnosticians to examine main arteries, which is exceptionally important for effective diagnosis. However, because of the small size of arteries relative ...
Added: November 6, 2025
Машинное обучение и представление информации: новые возможности цифровых архивов (рецензия на книгу: Artificial Intelligence, Archives and Manuscripts. New Relationships between the Virtual Archive and Its Referent. Edinburgh: University of Edinburgh, 2025
Penskaja E., Имагология и компаративистика 2025 № 23 С. 380–389
The book Artificial Intelligence, Archives and Manuscripts. New Relationships between the Virtual Archive and Its Referent (2025) is presented. This collective monograph discusses both technological and legal, intellectual issues that researchers and archivists face in automated work with manuscript heritage, artificial intelligence and neural networks. ...
Added: October 30, 2025
Free energy of neural network can predict accuracy after pruning
Surkov A., Sergei Koltcov, Ignatenko V. et al., Physica A: Statistical Mechanics and its Applications 2025 Vol. 681 Article 131085
Neural networks are powerful tools capable of achieving state-of-the-art performance across a wide range of tasks; however, their effectiveness often comes at the cost of extremely large numbers of parameters, which can hinder their deployment in resource-constrained environments. To address this issue, various pruning techniques have been proposed to reduce model size and complexity while ...
Added: October 30, 2025
Исследования благополучия с помощью передовых методов обработки естественного языка (NLP): перспективы и ограничения
Voevodina E., Современная зарубежная психология 2025 Т. 14 № 3 С. 172–181
Context and relevance. Well-being research faces methodological limitations of conventional psychometric measures, criticized for poor ecological validity, limited information yield, and inadequate capture of multidimensional construct of well-being. Advanced natural language processing (NLP) technologies offer solutions to these constraints. Objective. To evaluate opportunities and challenges of transformer-based NLP for well-being research. Methods and materials. We conducted an analytical review of ...
Added: October 9, 2025
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit