Volume 23, 2019
|Page(s)||310 - 337|
|Published online||17 June 2019|
Optimal survey schemes for stochastic gradient descent with applications to M-estimation
Telecom ParisTech LTCI, Université Paris Saclay,
46 Rue Barrault,
2 Université Paris Ouest, MODAL’X, 200 Avenue de la République, Nanterre 92000, France.
3 Mines ParisTech, PSL University, Centre de géosciences, 35 Rue Saint Honoré, Fontainebleau 77305, France.
* Corresponding author: email@example.com
Accepted: 22 October 2018
Iterative stochastic approximation methods are widely used to solve M-estimation problems, in the context of predictive learning in particular. In certain situations that shall be undoubtedly more and more common in the Big Data era, the datasets available are so massive that computing statistics over the full sample is hardly feasible, if not unfeasible. A natural and popular approach to gradient descent in this context consists in substituting the “full data” statistics with their counterparts based on subsamples picked at random of manageable size. It is the main purpose of this paper to investigate the impact of survey sampling with unequal inclusion probabilities on stochastic gradient descent-based M-estimation methods. Precisely, we prove that, in presence of some a priori information, one may significantly increase statistical accuracy in terms of limit variance, when choosing appropriate first order inclusion probabilities. These results are described by asymptotic theorems and are also supported by illustrative numerical experiments.
Mathematics Subject Classification: 62D05
Key words: Asymptotic analysis / central limit theorem / Horvitz–Thompson estimator / M-estimation / Poisson sampling / stochastic gradient descent / survey scheme
© EDP Sciences, SMAI 2019
Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.
Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.
Initial download of the metrics may take a while.