Subsemble: an ensemble method for combining subset-specific algorithm fits
Ensemble methods using the same underlying algorithm trained on different subsets of observations have recently received increased attention as practical prediction tools for massive data sets. We propose Subsemble: a general subset ensemble prediction method, which can be used for small, moderate, or large data sets. Subsemble partitions the full data set into subsets of observations, fits a specified underlying algorithm on each subset, and uses a clever form of <italic>V</italic>-fold cross-validation to output a prediction function that combines the subset-specific fits. We give an oracle result that provides a theoretical performance guarantee for Subsemble. Through simulations, we demonstrate that Subsemble can be a beneficial tool for small- to moderate-sized data sets, and often has better prediction performance than the underlying algorithm fit just once on the full data set. We also describe how to include Subsemble as a candidate in a SuperLearner library, providing a practical way to evaluate the performance of Subsemble relative to the underlying algorithm fit just once on the full data set.
Year of publication: |
2014
|
---|---|
Authors: | Sapp, Stephanie ; Laan, Mark J. van der ; Canny, John |
Published in: |
Journal of Applied Statistics. - Taylor & Francis Journals, ISSN 0266-4763. - Vol. 41.2014, 6, p. 1247-1259
|
Publisher: |
Taylor & Francis Journals |
Saved in:
Online Resource
Saved in favorites
Similar items by person
-
Locally Efficient Estimation With Bivariate Right-Censored Data
Quale, Christopher M., (2006)
-
Rosenblum, Michael, (2009)
-
Schnitzer, Mireille E., (2014)
- More ...