• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • A
  • A
  • A
  • A
  • A
Обычная версия сайта
  • RU
  • EN
  • HSE University
  • Publications
  • Book chapter
  • DEDPUL: Difference-of-Estimated-Densities-based Positive-Unlabeled Learning
  • RU
  • EN
Расширенный поиск
Высшая школа экономики
Национальный исследовательский университет
Priority areas
  • business informatics
  • economics
  • engineering science
  • humanitarian
  • IT and mathematics
  • law
  • management
  • mathematics
  • sociology
  • state and public administration
by year
  • 2028
  • 2027
  • 2026
  • 2025
  • 2024
  • 2023
  • 2022
  • 2021
  • 2020
  • 2019
  • 2018
  • 2017
  • 2016
  • 2015
  • 2014
  • 2013
  • 2012
  • 2011
  • 2010
  • 2009
  • 2008
  • 2007
  • 2006
  • 2005
  • 2004
  • 2003
  • 2002
  • 2001
  • 2000
  • 1999
  • 1998
  • 1997
  • 1996
  • 1995
  • 1994
  • 1993
  • 1992
  • 1991
  • 1990
  • 1989
  • 1988
  • 1987
  • 1986
  • 1985
  • 1984
  • 1983
  • 1982
  • 1981
  • 1980
  • 1979
  • 1978
  • 1977
  • 1976
  • 1975
  • 1974
  • 1973
  • 1972
  • 1971
  • 1970
  • 1969
  • 1968
  • 1967
  • 1966
  • 1965
  • 1964
  • 1963
  • 1958
  • More
Subject
News
September 11, 2026
How to Assess Students Knowledge in the Age of AI
A researcher at HSE University has proposed a flowchart to help lecturers decide how to assess students who use artificial intelligence. It shows where the use of AI should be restricted and where it can be incorporated into the learning process. The article has been published in IT Professional.
September 9, 2026
‘Balkan Hospitality Opens Doors: Studying Dialects on the Verge of Extinction
You cannot study spoken dialects from books. Instead, you need to go to a village, seek out its elders, and earn the trust of local residents before you can record hours of spontaneous stories. This is how Natalia Muravleva, Associate Professor at the Faculty of Humanities, conducts her research. Her internship in Serbia continued her long-standing study of dialects spoken by Macedonian settlers. In this interview, she discusses how diaspora cultural centres help researchers reach informants, why native speakers need to be interviewed only in their own language (otherwise, as she puts it, they may 'break'), and how a single field season helped her finalise her monograph. She also shares warm memories of autumn in Belgrade and of colleagues with whom grammar can be discussed in three languages at once.
September 9, 2026
Scientists Train Neural Network to Generate Process Plans from 3D Models
Researchers at the HSE FCS AI and Digital Science Institute have developed CAD2TechSpec, a framework that converts 3D models of mechanical parts into machining process plans—step-by-step instructions for machine tools. The solution aims to reduce the time required for the design and preparation of technical process documentation in mechanical engineering, aircraft manufacturing, and other high-tech industries. The study findings have been published in PeerJ Computer Science.

 

Have you spotted a typo?
Highlight it, click Ctrl+Enter and send us a message. Thank you for your help!

Publications
  • Books
  • Articles
  • Chapters of books
  • Working papers
  • Report a publication
  • Research at HSE

?

DEDPUL: Difference-of-Estimated-Densities-based Positive-Unlabeled Learning

P. 782–790.
Dmitry Ivanov

Positive-Unlabeled (PU) learning is an analog to supervised binary classification for the case when only the positive
sample is clean, while the negative sample is contaminated with latent instances of positive class and hence can be considered as an unlabeled mixture. The objectives are to classify the unlabeled sample and train an unbiased positive-negative classifier, which generally requires to identify the mixing proportions of positives and negatives first. Recently, unbiased risk estimation framework has achieved state-of-the-art performance in PU learning. This approach, however, exhibits two major bottlenecks. First, the mixing proportions are assumed to be identified, i.e. known in the domain or estimated with additional methods. Second, the approach relies on the classifier being a neural network. In this paper, we propose DEDPUL, a method that solves PU Learning without the aforementioned issues. The mechanism behind DEDPUL is to apply a computationally cheap postprocessing procedure to the predictions of any classifier trained to distinguish positive and unlabeled data. Instead of assuming the proportions to be identified, DEDPUL estimates them alongside with classifying unlabeled sample. Experiments show that DEDPUL
outperforms the current state-of-the-art in both proportion estimation and PU Classification and is flexible in the choice of the classifier.

Language: English
Full text
DOI
Text on another site
Keywords: density estimationSemi-supervised learning Positive-Unlabeled ClassificationMixture Proportions Estimation
Publication based on the results of:
Allocation and social choice mechanisms: axioms, incentives, algorithms (2020)

In book

2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA 2020)
Miami: IEEE, 2020.
Similar publications
Refining the ONCE Benchmark With Hyperparameter Tuning
Maksim Golyadkin, Alexander Gambashidze, Nurgaliev I. et al., IEEE Access 2024 Vol. 12 P. 3805–3814
In response to the growing demand for 3D object detection in applications such as autonomous driving, robotics, and augmented reality, this work focuses on the evaluation of semi-supervised learning approaches for point cloud data. The point cloud representation provides reliable and consistent observations regardless of lighting conditions, thanks to advances in LiDAR sensors. Data annotation ...
Added: March 13, 2024
Social media mining for ideation: Identification of sustainable solutions and opinions
Ozcan S., Suloglu M., Sakar C. O. et al., Technovation 2021 Vol. 107 No. September 2021 P. 1–12
The availability of social media-based data creates opportunities to obtain information about consumers, trends, companies and technologies using text mining techniques. However, the quality of the data is a significant concern for social media-based analyses. The aim of this study was to mine tweets (microblogs) to explore trends and retrieve ideas for various purposes such ...
Added: December 12, 2021
2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA 2020)
Miami: IEEE, 2020.
Positive-Unlabeled (PU) learning is an analog to supervised binary classification for the case when only the positive sample is clean, while the negative sample is contaminated with latent instances of positive class and hence can be considered as an unlabeled mixture. The objectives are to classify the unlabeled sample and train an unbiased positive-negative classifier, ...
Added: October 15, 2020
Density deconvolution under general assumptions on the distribution of measurement errors
Belomestny D., Goldenshluger A., Annals of Statistics 2021 Vol. 49 No. 2 P. 615–649
In  this paper we study   the problem of density deconvolution under general assumptions on the measurement error distribution. Typically deconvolution estimators are constructed using Fourier transform techniques, and it is assumed that the  characteristic function of the measurement errors does not have zeros on the real line. This assumption is rather strong and is not fulfilled in many cases of interest. ...
Added: April 7, 2020
Об оценке плотности распределения с помощью ряда Фурье
Belomestny D., Iosipoi L., Управление большими системами: сборник трудов 2019 № 82 С. 28–43
In this paper, we consider the classical statistical problem of probability density estimation based on a sample from this distribution. This problem naturally arises in many applications when one aims at investigation of a probability structure in a random process. For instance, it is possible to identify some structure in a complex system using density ...
Added: October 21, 2019
Semi-Conditional Normalizing Flows for Semi-Supervised Learning
Atanov A., Volokhova A., Ashukha A. et al., Workshop on Invertible Neural Nets and Normalizing Flows, International Conference on Machine Learning 2019 P. 1–9
This paper proposes a semi-conditional normalizing flow model for semi-supervised learning. The model uses both labeled and unlabeled data to learn an explicit model of joint distribution over objects and labels. Semi-conditional architecture of the model allows us to efficiently compute a value and gradients of the marginal likelihood for unlabeled objects. The conditional part ...
Added: July 11, 2019
On the prediction loss of the lasso in the partially labeled setting
Bellec P., Dalalyan A., Grappin E. et al., Electronic journal of statistics 2018 Vol. 12 No. 2 P. 3443–3472
In this paper we revisit the risk bounds of the lasso estimator in the context of transductive and semi-supervised learning. In other terms, the setting under consideration is that of regression with random design under partial labeling. The main goal is to obtain user-friendly bounds on the off-sample prediction risk. To this end, the simple ...
Added: November 9, 2018
Sobolev-Hermite versus Sobolev nonparametric density estimation on R
Belomestny D., Comte F., Genon-Catalot V., Annals of the Institute of Statistical Mathematics 2019 Vol. 71 No. 1 P. 29–62
In this spaper, our aim is to revisit the nonparametric estimation of a square integrable density f on R, by using projection estimators on a Hermite basis. These estimators are studied from the point of view of their mean integrated squared error on R. A model selection method is described and proved to perform an ...
Added: May 5, 2018
Generalized Post–Widder inversion formula with application to statistics
Belomestny D., Mai H., Schoenmakers J., Journal of Mathematical Analysis and Applications 2017 No. 455 P. 89–104
In this work we derive an inversion formula for the Laplace transform of a density observed on a curve in the complex domain, which generalizes the well known Post– Widder formula. We establish convergence of our inversion method and derive the corresponding convergence rates for the case of a Laplace transform of a smooth density. ...
Added: September 22, 2017
Asymptotics for in-sample density forecasting
Lee Y. K., Mammen E., Nielsen J. et al., Annals of Statistics 2015 No. 43 P. 620–645
This paper generalizes recent proposals of density forecasting models and it develops theory for this class of models. In density forecasting, the density of observations is estimated in regions where the density is not observed. Identification of the density in such regions is guaranteed by structural assumptions on the density that allows exact extrapolation. In ...
Added: December 12, 2014
  • About
  • About
  • Key Figures & Facts
  • Sustainability at HSE University
  • Faculties & Departments
  • International Partnerships
  • Faculty & Staff
  • HSE Buildings
  • HSE University for Persons with Disabilities
  • Public Enquiries
  • Studies
  • Admissions
  • Programme Catalogue
  • Undergraduate
  • Graduate
  • Exchange Programmes
  • Summer University
  • Summer Schools
  • Semester in Moscow
  • Business Internship
  • Research
  • International Laboratories
  • Research Centres
  • Research Projects
  • Monitoring Studies
  • Conferences & Seminars
  • Academic Jobs
  • Yasin (April) International Academic Conference on Economic and Social Development
  • Media & Resources
  • Publications by staff
  • HSE Journals
  • Publishing House
  • iq.hse.ru: commentary by HSE experts
  • Library
  • Economic & Social Data Archive
  • Video
  • HSE Repository of Socio-Economic Information
  • HSE1993–2026
  • Contacts
  • Copyright
  • Privacy Policy
  • Site Map
Edit