| Journal of Data Science | |
| Understanding Variable Effects from Black Box Prediction: Quantifying Effects in Tree Ensembles Using Partial Dependence | |
| article | |
| Guy Cafri1  Barbara A. Bailey1  | |
| [1] University of California, San Diego San Diego State University | |
| 关键词: Bagging; Bootstrap; Data Mining; | |
| DOI : 10.6339/JDS.201601_14(1).0005 | |
| 学科分类:土木及结构工程学 | |
| 来源: JDS | |
PDF
|
|
【 摘 要 】
Scientific interest often centers on characterizing the effect of one or more variables on an outcome. While data mining approaches such as random forests are flexible alternatives to conventional parametric models, they suffer from a lack of interpretability because variable effects are not quantified in a substantively meaningful way. In this paper we describe a method for quantifying variable effects using partial dependence, which produces an estimate that can be interpreted as the effect on the response for a one unit change in the predictor, while averaging over the effects of all other variables. Most importantly, the approach avoids problems related to model misspecification and challenges to implementation in high dimensional settings encountered with other approaches (e.g., multiple linear regression). We propose and evaluate through simulation a method for constructing a point estimate of this effect size. We also propose and evaluate interval estimates based on a non-parametric bootstrap. The method is illustrated on data used for the prediction of the age of abalone.
【 授权许可】
CC BY
【 预 览 】
| Files | Size | Format | View |
|---|---|---|---|
| RO202307150000233ZK.pdf | 602KB |
PDF