• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • FastWave: Optimized Diffusion Model for Audio Super-Resolution
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
October 8, 2026
HSE Experts Take Part in 23rd Annual Meeting of Valdai Discussion Club
The 23rd Annual Meeting of the Valdai Discussion Club was held from September 28 to October 1, 2026 under the theme ‘Responsibility for the Future: Limits of the Possible, or Limitless Possibilities?’ The forum brought together 120 experts from 40 countries, including representatives of China, the United States, India, Brazil, the United Kingdom, Germany, Egypt, Iran, and Japan.
October 7, 2026
‘Our Team Consists of True Leaders in Their Respective Academic Disciplines
The HSE International Centre of Decision Choice and Analysis studies a wide range of methods for analysing decision-making and possible scenarios for the development of natural, socio-economic, and political phenomena using various mathematical models. The application of advanced mathematical methods to forecasting helps to prevent negative outcomes and avoid erroneous decisions. The HSE News Service spoke to the centre’s director, Prof. Fuad Aleskerov, about its work.
October 6, 2026
International N5 Symposium ‘Neural Networks and Nonlinearity in Nizhny Novgorod Brings Together Scientists from Russia and Serbia
The International N5 Symposium ‘Neural Networks and Nonlinearity in Nizhny Novgorod’ was held at the Nizhny Novgorod House of Scientists from September 23 to 26. The event was organised by HSE University–Nizhny Novgorod and the Nizhny Novgorod House of Scientists, with the participation of Sberbank and the Institute of Physics Belgrade. The symposium was held for the second time: the first conference took place in 2025 and attracted considerable interest from the academic community.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

FastWave: Optimized Diffusion Model for Audio Super-Resolution

P. 4506–4510.
Kuznetsov N., Kaledin M.

Audio Super-Resolution is a set of techniques aimed at high-quality estimation of the given signal as if it would be sampled with higher sample rate. Among suggested methods there are diffusion and flow models (which are considered slower), generative adversarial networks (which are considered faster), however both approaches are currently presented by high-parametric networks, requiring high computational costs both for training and inference. We propose a solution to both these problems by re-considering the recent advances in the training of diffusion models and applying them to super-resolution from any to 48 kHz sample rate. Our model called FastWave has around 50 GFLOPs of computational complexity and 1.3 M parameters and can be trained with less resources and significantly faster than the majority of recently proposed diffusion- and flow-based solutions. FastWave also has comparable performance to state-of-the-art models. We provide our implementation on GitHub.

Language: English
Full text
Text on another site
Keywords: speech processingbandwidth extensionDiffusion ModelsAudio super-resolution

In book

Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2026
International Speech Communication Association, 2026.
Similar publications
OrthoFuse: Training-free Riemannian Fusion of Orthogonal Style-Concept Adapters for Diffusion Models
Ali Aliev, Garifullin K., Yudin N. et al., , in: 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).: IEEE, 2026. P. 36009–36018.
In a rapidly growing field of model training there is a constant practical interest in parameter-efficient fine-tuning and various techniques that use a small amount of training data to adapt the model to a narrow task. However, there is an open question: how to combine several adapters tuned for different tasks into one which is ...
Added: July 13, 2026
GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
Oganov A., Bykov I., Neudachina E. et al., , in: The Fourteenth International Conference on Learning Representations (ICLR 2026).: ICLR, 2026. Ch. 18372.
While diffusion models achieve state-of-the-art generation quality, they still suffer from computationally expensive sampling. Recent works address this issue with gradient-based optimization methods that distill a few-step ODE diffusion solver from the full sampling process, reducing the number of function evaluations from dozens to just a few. However, these approaches often rely on intricate training ...
Added: July 13, 2026
LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters
Vladimir Bogachev, Aletov V., Alexander Molozhavenko et al., , in: The Fourteenth International Conference on Learning Representations (ICLR 2026).: ICLR, 2026. Ch. 20503 P. 1–26.
This work presents a novel, fully Riemannian framework for Low-Rank Adaptation (LoRA) that geometrically treats low-rank adapters by optimizing them directly on the fixed-rank manifold. This formulation eliminates the parametrization ambiguity present in standard Euclidean optimizers. Our framework integrates three key components to achieve this: (1) we derive Riemannion, a new Riemannian optimizer on the fixed-rank ...
Added: April 29, 2026
Neuro-oscillatory models of cortical speech processing
Dogonasheva O., Giraud A., Zakharov D. et al., Neural Networks 2025 Vol. 195 Article 108194
In this review, we examine computational models that explore the role of neural oscillations in speech perception, spanning from early auditory processing to higher cognitive stages. We focus on models that use rhythmic brain activities, such as gamma, theta, and delta oscillations, to encode phonemes, segment speech into syllables and words, and integrate linguistic elements ...
Added: October 16, 2025
CA-SER: Cross-Attention Feature Fusion for Speech Emotion Recognition
Deeb B., Savchenko A., Makarov I., , in: ECAI 2024. 27th European Conference on Artificial Intelligence, October 19 – 24 October 2024, Santiago de Compostela, Spain – Including 13th Conference on Prestigious Applications of Intelligent Systems (PAIS 2024).: IOS Press, 2024. P. 4479–4482.
In this paper, we introduce a novel tool for speech emotion recognition, CA-SER, that borrows self-supervised learning to extract semantic speech representations from a pre-trained wav2vec 2.0 model and combine them with spectral audio features to improve speech emotion recognition. Our approach involves a self-attention encoder on MFCC features to capture meaningful patterns in audio ...
Added: February 15, 2025
Нейрофизиологические корреляты автоматической обработки нулевой морфемы: данные вызванных потенциалов.
Alexeeva M., Myachykov A., Shtyrov Y., Журнал высшей нервной деятельности им. И.П. Павлова 2022 Т. 72 № 5 С. 666–677
Language functioning as a communicative system is described by a multitude of linguistic theories, which are not always consistent with each other and do not have strong cognitive and/or neurobiological bases. One of the most striking examples is the “zero morpheme” proposed by the Universal Grammar theory, which has only an abstract meaning and no ...
Added: September 20, 2022
Gain-optimized spectral distortions for pronunciation training
Savchenko A., Savchenko V., Savchenko L., Optimization Letters 2022 Vol. 16 No. 7 P. 2095–2113
This paper considers an assessment and evaluation of speech sound pronunciation quality in computer-aided language learning systems. We examine the gain optimization of spectral distortion measures between the speech signals of a native speaker and a learner. During training, a learner has to achieve stable pronunciation of all sounds. This is measured by computing the ...
Added: August 18, 2022
Критерий гарантированного уровня значимости в задаче автоматической сегментации речевого сигнала
Савченко В. В., Savchenko A., Радиотехника и электроника 2020 Т. 65 № 11 С. 1101–1108
The article considers the problem of automatic segmentation of a speech signal into phonetic units in conditions of their a priori uncertain spectral composition and correlation properties. A guaranteed significance level criterion is developed based on the information–theoretic approach. An example of practical application of this criterion is considered; a full-scale experiment is set up ...
Added: November 28, 2020
A Method for the Real-Time Updating of Voice Samples in the Unified Biometric System
Savchenko V., Savchenko A., Measurement Techniques 2020 Vol. 63 No. 5 P. 391–400
The article continues our previous work [V. V. Savchenko and A. V. Savchenko, Izmeritel’naya Tekhnika, No. 12, 46–52 (2019)]. We study the problem of automatic control of the quality of voice samples recorded and stored in the Unified Biometric System. We propose a solution of the problem of timely updates of the collected samples required ...
Added: October 1, 2020
Method for Measuring the Indicator of Acoustic Quality of Audio Recordings Prepared for Registration and Processing in the Unified Biometric System
Savchenko V., Savchenko A., Measurement Techniques 2020 Vol. 62 No. 12 P. 1071–1078
The problem of automated quality control of audio recordings containing voice samples of individuals is considered. It is shown that in solving this problem the most acute impediment is the problem of small samples of observations. To overcome the problem, a new, high-speed method of acoustic measurements is proposed, based on the principle of relative ...
Added: April 22, 2020
Speech and Computer. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) Volume 11658 LNAI
Springer, 2019.
Speech and Computer ...
Added: October 28, 2019
Criterion of Significance Level for Selection of Order of Spectral Estimation of Entropy Maximum
Savchenko A., Savchenko V., Radioelectronics and Communications Systems 2019 Vol. 62 No. 5 P. 223–231
It is researched a wide class of parametric estimations of power spectral density based on principle of entropy maximum and autoregression observation model. At that there is distinguished the key parameter which is used model order. It is considered a problem of a priori uncertainty when true value of order is a priori unknown. It ...
Added: August 16, 2019
A Method for Measuring the Pitch Frequency of Speech Signals for the Systems of Acoustic Speech Analysis
Savchenko A., Savchenko V., Measurement Techniques 2019 Vol. 62 No. 3 P. 282–288
We developed a new method for measuring the pitch frequency of speech signals with elevated noise immunity. The problem of protection against intense background noise is solved in this method by the frequency selection of vocalized segments of speech signals according to a scheme with comb filter of interperiodic accumulation. The efficiency of the method ...
Added: August 16, 2019
Анализ качества речи в информационной теории восприятия речи
Karpov N., Системы управления и информационные технологии 2012 Т. 48 № 2.1 С. 145–149
In this article analyzed some methods of speech quality estimation based on State Standard and Informational Theory of Speech Perception. Experimentally examine effectiveness and boundaries of free methods for speech parameterization and using it with deferent metrics. ...
Added: September 11, 2012
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit