?
Neural Networks Compression for Language Modeling
P. 351–357.
In this paper, we consider several compression techniques for the language modeling problem based on recurrent neural networks (RNNs). It is known that conventional RNNs, e.g., LSTM-based networks in language modeling, are characterized with either high space complexity or substantial inference time. This problem is especially crucial
for mobile applications, in which the constant interaction with the remote server is inappropriate. By using the Penn Treebank (PTB) dataset we compare pruning, quantization, low-rank factorization, tensor train decomposition for LSTM networks in terms of model size and suitability for fast inference.
Vasilev A., Kapitanov A., Roman Solovyev et al., PeerJ Computer Science 2026 Vol. 12 Article 3724
This article introduces MinMAE, a novel activation calibration method for Post-Training Quantization (PTQ) that significantly reduces accuracy loss in Convolutional Neural Networks (CNN). Motivated by the need for high-fidelity quantization without costly retraining, MinMAE directly minimizes the Mean Absolute Error (MAE) between original and dequantized activations, making it robust to outliers that degrade standard methods. ...
Added: May 3, 2026
Surkov A., Sergei Koltcov, Ignatenko V. et al., Physica A: Statistical Mechanics and its Applications 2025 Vol. 681 Article 131085
Neural networks are powerful tools capable of achieving state-of-the-art performance across a wide range of tasks; however, their effectiveness often comes at the cost of extremely large numbers of parameters, which can hinder their deployment in resource-constrained environments. To address this issue, various pruning techniques have been proposed to reduce model size and complexity while ...
Added: October 30, 2025
Nazarova V., Lodiagin B., Круглов Ф. А. et al., AlterEconomics (ранее - Журнал экономической теории) 2025 № 22(3) С. 482–502
This paper examines methods for forecasting oil prices, comparing traditional autoregressive mo dels (ARIMA, SARIMAX) with machine learning approaches (LSTM). The target variable is the price of WTI crude oil. The dataset covers 2015–2019 and includes both WTI price data and a set of exogenous varia bles: the Wilshire 5000, Dow Jones, and DXY indices; ...
Added: October 5, 2025
Surkov A., Zakharov V., Sergei Koltcov et al., , in: Smart Technologies, Systems and Applications: 4th International Conference, SmartTech-IC 2024, Quito, Ecuador, December 2–4, 2024, Revised Selected Papers, Part IIVol. 2: Revised Selected Papers, Part II.: Springer, 2025. P. 239–252.
Currently, large language models are actively developing and beginning to be used to solve some mathematical problems. With the emergence of xLSTM model, which demonstrates the results comparable with transformer-based models, there has been a surge of interest in recurrent neural networks. This paper considers the application of baseline recurrent models such as LSTM and ...
Added: September 11, 2025
Artem B., Andreasyan A., Konovalov D. et al., Scientific Reports 2025 Vol. 15 Article 23119
G-quadruplexes (GQs) are non-canonical DNA structures encoded by G-flipons with potential roles in gene regulation and chromatin structure. Here, we explore the role of G-flipons in tissue specification. We present a deep learning-based framework for the genome-wide G-flipon predictions across 14 human tissue types. The model was trained using high-confidence experimental maps of GQ-forming sequences ...
Added: August 8, 2025
Микулинский А. Д., , in: Синергия языков и культур 2022: междисциплинарные исследования.: St. Petersburg: -, 2023. P. 335–351.
The paper is devoted to the issue of the local structure modeling of the eSports commentary spoken genre on an example of the Dota 2 computer discipline. ESports commentary is a spontaneous and creative speech aimed at describing of what is happening on the computer-gaming field. The main factors that force us to study it ...
Added: May 12, 2024
Beknazarov N., , in: Z-DNA: Methods and Protocols.: United States of America: Springer, 2023. P. 217–226.
Here we describe an approach that uses deep learning neural networks such as CNN and RNN to aggregate information from DNA sequence; physical, chemical, and structural properties of nucleotides; and omics data on histone modifications, methylation, chromatin accessibility, and transcription factor binding sites and data from other available NGS experiments. We explain how with the ...
Added: December 26, 2023
Natalia Sizykh, Said Dandamaev, Dmitry Sizykh, , in: 16th International Conference Management of large-scale system development (MLSD).: IEEE, 2023. P. 1–5.
Forecasting data and research on cryptocurrency price forecasting methods are increasing in importance. So far, methods based on LSTM deep learning architecture have shown the best results in forecasting cryptocurrency prices. In order to improve the accuracy of forecasting data, this paper investigates the application of a multivariate multistep forecasting method based on the LSTM ...
Added: December 22, 2023
I. K. Kusakin, Fedorets O. V., A. Y. Romanov, Scientific and Technical Information Processing 2023 Vol. 50 No. 3 P. 176–183
This paper discusses modern approaches to natural language processing and the application of machine learning models to the task of classifying short scientific texts in Russian. This study is devoted to the analysis of methods for vectorization of textual information, selection of a model for scientific paper clas- sification, and training of linguistic model BERT ...
Added: November 4, 2023
Taktasheva E., Shavrina T., Fenogenova A. et al., , in: Findings of the Association for Computational Linguistics: EMNLP 2022.: Association for Computational Linguistics, 2022. P. 2472–2497.
Recent advances in zero-shot and few-shot learning have shown promise for a scope of research and practical purposes. However, this fast-growing area lacks standardized evaluation suites for non-English languages, hindering progress outside the Anglo-centric paradigm. To address this line of research, we propose TAPE (Text Attack and Perturbation Evaluation), a novel benchmark that includes six ...
Added: September 22, 2023
Кусакин И. К., Федорец О. В., Romanov A., Научно-техническая информация. Серия 2: Информационные процессы и системы 2022 Т. 12 С. 6–9
This paper discusses modern approaches to natural language processing and appliance of artificial intelligence technologies in the task of classifying scientific texts in Russian. The report contains an analysis of implementations of text vectorization methods, a description of experiments with training various classifier models: from classical machine learning algorithms to neural network transformer architectures. ...
Added: January 31, 2023
Osipov D., / Series arXiv "math". 2022. No. 1.
This paper devotes to comparison of different cod- ing schemes (various constructions of Polar and LDPC codes, Product codes and BCH codes) for the case when information is transmitted over AWGN channel with quantization with lowest possible complexity and resolution: 1-bit. We examine performance (in terms of Frame-error-rate — FER) for schemes mentioned above and ...
Added: December 27, 2022
Tikhonova M., Mikhailov V., Dina Pisarevskaya et al., Natural Language Engineering 2022 P. 1–30
Recent research has reported that standard fine-tuning approaches can be unstable due to being prone to various sources of randomness, including but not limited to weight initialization, training data order, and hardware. Such brittleness can lead to different evaluation results, prediction confidences, and generalization inconsistency of the same models independently fine-tuned under the same experimental setup. ...
Added: May 21, 2022
Vukovic D., Romanyuk K., Ivashchenko S. et al., Expert Systems with Applications 2022 Vol. 194 No. May 2022 Article 116553
This paper investigates the forecasting performance for credit default swap (CDS) spreads by Support Vector
Machines (SVM), Group Method of Data Handling (GMDH), Long Short-Term Memory (LSTM) and Markov
switching autoregression (MSA) for daily CDS spreads of the 513 leading US companies, in the period
2009–2020. The goal of this study is to test the forecasting performance of ...
Added: February 4, 2022
Cham: Birkhäuser, 2020.
The book consists of articles based on the XXXVIII Białowieża Workshop on Geometric Methods in Physics, 2019. The series of Białowieża workshops, attended by a community of experts at the crossroads of mathematics and physics, is a major annual event in the field. The works in this book, based on presentations given at the workshop, ...
Added: November 3, 2021
Lukankin Alexander, Slastnikov Sergey, Journal of Physics: Conference Series 2021 Vol. 1740 P. 1–6
Paper is devoted to the predictive models for metrological indicators on the real estate engineering infrastructure. The solution is in demand among many enterprises both in terms of security and economic considerations. The key task is to build a mathematical model performing predictions on the real data samples. We study both classical predictive models (ARIMA, ...
Added: February 2, 2021
Beknazarov N., Jin S., Poptsova M., Scientific Reports 2020 Vol. 10 P. 19134
Computational methods to predict Z-DNA regions are in high demand to understand the functional role of Z-DNA. The previous state-of-the-art method Z-Hunt is based on statistical mechanical and energy considerations about B- to Z-DNA transition using sequence information. Z-DNA CHiP-seq experiment results showed little overlap with Z-Hunt predictions implying that sequence information only is not ...
Added: December 11, 2020
Eduard Gorbunov, Kovalev D., Makarenko D. et al., , in: Advances in Neural Information Processing Systems 33 (NeurIPS 2020).: Curran Associates, Inc., 2020. P. 20889–20900.
Added: December 7, 2020
Feigin B. L., Russian Mathematical Surveys 2017 Vol. 72 No. 4 P. 707–763
This paper discusses the main known constructions of vertex operator algebras. The starting point is the lattice algebra. Screenings distinguish subalgebras of lattice algebras. Moreover, one can construct extensions of vertex algebras. Combining these constructions gives most of the known examples. A large class of algebras with big centres is constructed. Such algebras have applications ...
Added: November 5, 2020
Giachanou A., Россо П., Crestani F., , in: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’19).: NY: Association for Computing Machinery (ACM), 2019. P. 877–880.
The spread of false information on the Web is one of the main problems of our society. Automatic detection of fake news posts is a hard task since they are intentionally written to mislead the readers and to trigger intense emotions to them in an attempt to be disseminated in the social networks. Even though ...
Added: October 29, 2020
Feigin B. L., Jimbo M., Mukhin E., Journal of Mathematical Physics 2019 Vol. 60 No. 7 P. 073507-1–073507-16
We discuss the quantization of the ̂ sl 2 coset vertex operator algebra W D(2,1;α) using the bosonization technique. We show that after quantization, there exist three families of commuting integrals of motion coming from three copies of the quantum toroidal algebra associated with gl 2 . ...
Added: December 10, 2019
Kodryan M., Grachev A., Ignatov D. I. et al., , in: Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)Issue W19-43.: Association for Computational Linguistics, 2019. P. 40–48.
Reduction of the number of parameters is one of the most important goals in Deep Learning. In this article we propose an adaptation of Doubly Stochastic Variational Inference for Automatic Relevance Determination (DSVI-ARD) for neural networks compression. We find this method to be especially useful in language modeling tasks, where large number of parameters in ...
Added: November 1, 2019