• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • Neural Networks Compression for Language Modeling
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 17, 2026
'I Wish That People Would Place Greater Trust in Science'
When Tatiana Eremicheva chose Fundamental and Computational Linguistics as her field of study, she thought it would be about learning languages. Instead, she discovered it was about helping people. In this interview for the HSE Young Scientists project, she discusses science as a way of understanding the world, billiards as a team-building activity, and why learning to read is not always as easy as it seems.
September 15, 2026
Immunity to Chaos: How Personal Resources Help Us Cope with the Challenges of a Turbulent World
International conflicts, crises and digital overload—the modern world puts our minds to the test every day. Traditional psychology often focuses on the consequences: anxiety, depression, and psychosomatic disorders. But what if we looked at the problem differently—through the lens of the resources that prevent us from breaking down? Psychological immunity is precisely this set of resources. Alena Zolotareva and her group, Psychological Immunity as a Resource for Positive Functioning, are developing an integrative model of this phenomenon, adapting diagnostic tools and preparing for large-scale empirical research. Why do psychologists need to collaborate with medical professionals, and how could their research transform preventive care in clinics and corporations?
September 11, 2026
How to Assess Students Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

Neural Networks Compression for Language Modeling

P. 351–357.
Grachev A., Ignatov D. I., Savchenko A.

In this paper, we consider several compression techniques for the language modeling problem based on recurrent neural networks (RNNs). It is known that conventional RNNs, e.g., LSTM-based networks in language modeling, are characterized with either high space complexity or substantial inference time. This problem is especially crucial
for mobile applications, in which the constant interaction with the remote server is inappropriate. By using the Penn Treebank (PTB) dataset we compare pruning, quantization, low-rank factorization, tensor train decomposition for LSTM networks in terms of model size and suitability for fast inference.

Language: English
Full text
DOI
Text on another site
Keywords: quantizationlanguage modelingPruningLSTMRNNLow-rank factorization

In book

Pattern Recognition and Machine Intelligence. 7th International Conference, PReMI 2017, Kolkata, India, December 5-8, 2017, Proceedings. Lecture Notes in Computer Science book series (LNCS, volume 10597)
Springer, 2017.
Similar publications
MinMAE calibration method for convolutional neural network quantization
Vasilev A., Kapitanov A., Roman Solovyev et al., PeerJ Computer Science 2026 Vol. 12 Article 3724
This article introduces MinMAE, a novel activation calibration method for Post-Training Quantization (PTQ) that significantly reduces accuracy loss in Convolutional Neural Networks (CNN). Motivated by the need for high-fidelity quantization without costly retraining, MinMAE directly minimizes the Mean Absolute Error (MAE) between original and dequantized activations, making it robust to outliers that degrade standard methods. ...
Added: May 3, 2026
Free energy of neural network can predict accuracy after pruning
Surkov A., Sergei Koltcov, Ignatenko V. et al., Physica A: Statistical Mechanics and its Applications 2025 Vol. 681 Article 131085
Neural networks are powerful tools capable of achieving state-of-the-art performance across a wide range of tasks; however, their effectiveness often comes at the cost of extremely large numbers of parameters, which can hinder their deployment in resource-constrained environments. To address this issue, various pruning techniques have been proposed to reduce model size and complexity while ...
Added: October 30, 2025
Применение методов машинного обучения для прогнозирования нефтяных котировок
Nazarova V., Lodiagin B., Круглов Ф. А. et al., AlterEconomics (ранее - Журнал экономической теории) 2025 № 22(3) С. 482–502
This paper examines methods for forecasting oil prices, comparing traditional autoregressive mo dels (ARIMA, SARIMAX) with machine learning approaches (LSTM). The target variable is the price of WTI crude oil. The dataset covers 2015–2019 and includes both WTI price data and a set of exogenous varia bles: the Wilshire 5000, Dow Jones, and DXY indices; ...
Added: October 5, 2025
Application of Large Language Models to Solving Differential Equations: Constructing Baseline Models with LSTM and GRU
Surkov A., Zakharov V., Sergei Koltcov et al., , in: Smart Technologies, Systems and Applications: 4th International Conference, SmartTech-IC 2024, Quito, Ecuador, December 2–4, 2024, Revised Selected Papers, Part IIVol. 2: Revised Selected Papers, Part II.: Springer, 2025. P. 239–252.
Currently, large language models are actively developing and beginning to be used to solve some mathematical problems. With the emergence of xLSTM model, which demonstrates the results comparable with transformer-based models, there has been a surge of interest in recurrent neural networks. This paper considers the application of baseline recurrent models such as LSTM and ...
Added: September 11, 2025
Deep learning deciphers the related role of master regulators and G-quadruplexes in tissue specification
Artem B., Andreasyan A., Konovalov D. et al., Scientific Reports 2025 Vol. 15 Article 23119
G-quadruplexes (GQs) are non-canonical DNA structures encoded by G-flipons with potential roles in gene regulation and chromatin structure. Here, we explore the role of G-flipons in tissue specification. We present a deep learning-based framework for the genome-wide G-flipon predictions across 14 human tissue types. The model was trained using high-confidence experimental maps of GQ-forming sequences ...
Added: August 8, 2025
Functional models of elementary discursive units in Russian eSports commentary
Микулинский А. Д., , in: Синергия языков и культур 2022: междисциплинарные исследования.: St. Petersburg: -, 2023. P. 335–351.
The paper is devoted to the issue of the local structure modeling of the eSports commentary spoken genre on an example of the Dota 2 computer discipline. ESports commentary is a spontaneous and creative speech aimed at describing of what is happening on the computer-gaming field. The main factors that force us to study it ...
Added: May 12, 2024
DeepZ: A Deep Learning Approach for Z-DNA Prediction
Beknazarov N., , in: Z-DNA: Methods and Protocols.: United States of America: Springer, 2023. P. 217–226.
Here we describe an approach that uses deep learning neural networks such as CNN and RNN to aggregate information from DNA sequence; physical, chemical, and structural properties of nucleotides; and omics data on histone modifications, methylation, chromatin accessibility, and transcription factor binding sites and data from other available NGS experiments. We explain how with the ...
Added: December 26, 2023
Application of the Method of Multivariate Multi-stage Forecasting Based on the LSTM Deep Learning Model for Bitcoin Price Time Series
Natalia Sizykh, Said Dandamaev, Dmitry Sizykh, , in: 16th International Conference Management of large-scale system development (MLSD).: IEEE, 2023. P. 1–5.
Forecasting data and research on cryptocurrency price forecasting methods are increasing in importance. So far, methods based on LSTM deep learning architecture have shown the best results in forecasting cryptocurrency prices. In order to improve the accuracy of forecasting data, this paper investigates the application of a multivariate multistep forecasting method based on the LSTM ...
Added: December 22, 2023
Classification of Short Scientific Texts
I. K. Kusakin, Fedorets O. V., A. Y. Romanov, Scientific and Technical Information Processing 2023 Vol. 50 No. 3 P. 176–183
This paper discusses modern approaches to natural language processing and the application of machine learning models to the task of classifying short scientific texts in Russian. This study is devoted to the analysis of methods for vectorization of textual information, selection of a model for scientific paper clas- sification, and training of linguistic model BERT ...
Added: November 4, 2023
TAPE: Assessing Few-shot Russian Language Understanding
Taktasheva E., Shavrina T., Fenogenova A. et al., , in: Findings of the Association for Computational Linguistics: EMNLP 2022.: Association for Computational Linguistics, 2022. P. 2472–2497.
Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of research, we propose TAPE (Text Attack and Perturbation Evaluation), a novel benchmark that includes six ...
Added: September 22, 2023
Исследование методов машинного обучения для классификации научных текстов на русском языке
Кусакин И. К., Федорец О. В., Romanov A., Научно-техническая информация. Серия 2: Информационные процессы и системы 2022 Т. 12 С. 6–9
This paper discusses modern approaches to natural language processing and appliance of artificial intelligence technologies in the task of classifying scientific texts in Russian. The report contains an analysis of implementations of text vectorization methods, a description of experiments with training various classifier models: from classical machine learning algorithms to neural network transformer architectures. ...
Added: January 31, 2023
Comparison of different coding schemes for 1-bit ADC
Osipov D., / Series arXiv "math". 2022. No. 1.
This paper devotes to comparison of different cod- ing schemes (various constructions of Polar and LDPC codes, Product codes and BCH codes) for the case when information is transmitted over AWGN channel with quantization with lowest possible complexity and resolution: 1-bit. We examine performance (in terms of Frame-error-rate — FER) for schemes mentioned above and ...
Added: December 27, 2022
Ad Astra or Astray: Exploring Linguistic Knowledge of Multilingual BERT through NLI Task
Tikhonova M., Mikhailov V., Dina Pisarevskaya et al., Natural Language Engineering 2022 P. 1–30
Recent research has reported that standard fine-tuning approaches can be unstable due to being prone to various sources of randomness, including but not limited to weight initialization, training data order, and hardware. Such brittleness can lead to different evaluation results, prediction confidences, and generalization inconsistency of the same models independently fine-tuned under the same experimental setup. ...
Added: May 21, 2022
Are CDS spreads predictable during the Covid-19 pandemic? Forecasting based on SVM, GMDH, LSTM and Markov switching autoregression
Vukovic D., Romanyuk K., Ivashchenko S. et al., Expert Systems with Applications 2022 Vol. 194 No. May 2022 Article 116553
This paper investigates the forecasting performance for credit default swap (CDS) spreads by Support Vector Machines (SVM), Group Method of Data Handling (GMDH), Long Short-Term Memory (LSTM) and Markov switching autoregression (MSA) for daily CDS spreads of the 513 leading US companies, in the period 2009–2020. The goal of this study is to test the forecasting performance of ...
Added: February 4, 2022
Geometric Methods in Physics XXXVIII. Workshop, Białowieża, Poland, 2019
Cham: Birkhäuser, 2020.
The book consists of articles based on the XXXVIII Białowieża Workshop on Geometric Methods in Physics, 2019. The series of Białowieża workshops, attended by a community of experts at the crossroads of mathematics and physics, is a major annual event in the field. The works in this book, based on presentations given at the workshop, ...
Added: November 3, 2021
Predictive models for metrological data of engineering systems
Lukankin Alexander, Slastnikov Sergey, Journal of Physics: Conference Series 2021 Vol. 1740 P. 1–6
Paper is devoted to the predictive models for metrological indicators on the real estate engineering infrastructure. The solution is in demand among many enterprises both in terms of security and economic considerations. The key task is to build a mathematical model performing predictions on the real data samples. We study both classical predictive models (ARIMA, ...
Added: February 2, 2021
Deep learning approach for predicting functional Z-DNA regions using omics data
Beknazarov N., Jin S., Poptsova M., Scientific Reports 2020 Vol. 10 P. 19134
Computational methods to predict Z-DNA regions are in high demand to understand the functional role of Z-DNA. The previous state-of-the-art method Z-Hunt is based on statistical mechanical and energy considerations about B- to Z-DNA transition using sequence information. Z-DNA CHiP-seq experiment results showed little overlap with Z-Hunt predictions implying that sequence information only is not ...
Added: December 11, 2020
Linearly Converging Error Compensated SGD
Eduard Gorbunov, Kovalev D., Makarenko D. et al., , in: Advances in Neural Information Processing Systems 33 (NeurIPS 2020).: Curran Associates, Inc., 2020. P. 20889–20900.
Added: December 7, 2020
Extensions of vertex algebras. Constructions and applications
Feigin B. L., Russian Mathematical Surveys 2017 Vol. 72 No. 4 P. 707–763
This paper discusses the main known constructions of vertex operator algebras. The starting point is the lattice algebra. Screenings distinguish subalgebras of lattice algebras. Moreover, one can construct extensions of vertex algebras. Combining these constructions gives most of the known examples. A large class of algebras with big centres is constructed. Such algebras have applications ...
Added: November 5, 2020
Leveraging Emotional Signals for Credibility Detection
Giachanou A., Россо П., Crestani F., , in: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’19).: NY: Association for Computing Machinery (ACM), 2019. P. 877–880.
The spread of false information on the Web is one of the main problems of our society. Automatic detection of fake news posts is a hard task since they are intentionally written to mislead the readers and to trigger intense emotions to them in an attempt to be disseminated in the social networks. Even though ...
Added: October 29, 2020
Towards trigonometric deformation of 𝔰𝔩ˆ2 coset VOA
Feigin B. L., Jimbo M., Mukhin E., Journal of Mathematical Physics 2019 Vol. 60 No. 7 P. 073507-1–073507-16
We discuss the quantization of the ̂ sl 2 coset vertex operator algebra W D(2,1;α) using the bosonization technique. We show that after quantization, there exist three  families of commuting integrals of motion coming from three copies of the quantum toroidal algebra associated with gl 2 . ...
Added: December 10, 2019
Efficient Language Modeling with Automatic Relevance Determination in Recurrent Neural Networks
Kodryan M., Grachev A., Ignatov D. I. et al., , in: Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)Issue W19-43.: Association for Computational Linguistics, 2019. P. 40–48.
Reduction of the number of parameters is one of the most important goals in Deep Learning. In this article we propose an adaptation of Doubly Stochastic Variational Inference for Automatic Relevance Determination (DSVI-ARD) for neural networks compression. We find this method to be especially useful in language modeling tasks, where large number of parameters in ...
Added: November 1, 2019
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit