plot.subsample.rfsrc.RdDisplay confidence intervals for variable importance (VIMP) from a
subsample object. Select predictors, an outcome, and an
interval method using the saved replicate estimates. Prediction-error
and joint-VIMP rows can also be displayed when requested during
subsampling.
# S3 method for class 'rfsrc'
plot.subsample(x, alpha = .01, xvar.names,
standardize = TRUE, normal = TRUE, jknife = FALSE, target, m.target = NULL,
pmax = 75, main = "", sorted = TRUE, show.plots = TRUE, ...)An object returned by subsample.rfsrc.
Significance level for the displayed intervals, whose
nominal confidence level is \(1-\alpha\). The plotting default
.01 gives 99 percent intervals. Use alpha = .05 to
match the default of extract.subsample and the print method.
Names of predictor rows to display. If omitted,
all available rows are considered, subject to pmax.
This selects stored results and does not regrow the forest.
For regression outcomes, divide VIMP by the
variance of the selected response in the full training data.
The same variance is used for all replicate estimates and any
error row. Other families are unchanged. Set FALSE to
display the original error scale.
Use normal-approximation intervals when TRUE
and nonparametric intervals when FALSE. For double-bootstrap
objects, these are bootstrap normal and percentile intervals.
Select the delete-\(d\) jackknife standard error for
normal intervals. Use normal = TRUE, jknife = TRUE for this
display. The default normal display uses the subsampling standard
error. Ignored when normal = FALSE and for double-bootstrap
objects.
For classification, an integer or class label selecting
class-specific VIMP; the default is overall VIMP. For competing
risks, an integer from 1 to J selecting an event type;
the default is the first event. This selects a statistic within
the outcome identified by m.target.
Name of one response in a multivariate or mixed-outcome
forest. If omitted, a default response is selected. This selects
an outcome in the saved results; it is separate from choosing
a class or event with target.
Maximum number of rows retained for display, selected by VIMP.
Main plot title.
Order the displayed rows by importance.
Draw the plot. Set FALSE to obtain its
plot-data return value without opening or updating a plot.
Graphical arguments for the interval display, including
xlim, ylim, xlab, ylab, boxfill,
border, whisklty, whisklwd, and axis controls.
A named pars list is also accepted; direct arguments take
precedence. See Details.
The default uses normal intervals and the subsampling standard
error. Use normal = TRUE, jknife = TRUE for normal intervals
with the jackknife standard error, or normal = FALSE for
nonparametric intervals. All use the estimates already stored
in x; changing alpha or the display method does not
perform another round of subsampling.
The subsampling standard error measures replicate dispersion
around the replicate mean. The jackknife calculation measures
dispersion around the full-data estimate, incorporating the
displacement between those two centers. The jackknife standard
error is not necessarily larger for a finite replicate set.
The nonparametric interval reverses empirical quantiles of the
centered subsampling roots. See subsample.rfsrc
for the calculations.
The display summarizes uncertainty in VIMP, not the distribution of the observed predictor values. Its intervals are calculated separately for each displayed statistic. A positive lower VIMP endpoint identifies a positive-importance result under that interval procedure; the display does not apply a multiple-testing correction. Prediction-error rows summarize the error itself.
xvar.names selects rows and m.target selects a response.
For a classification response or competing-risk outcome,
target additionally selects its class-specific or
event-specific statistic. These selections are made after
resampling, so one saved object can be used for several displays.
Set alpha explicitly when comparing results with
print or extract.subsample; their default is .05,
whereas the plotting default is .01.
The default display is horizontal. xlim controls its
horizontal VIMP/error axis and ylim controls its vertical
row positions. These arguments always refer to the displayed axes;
the wrapper handles the internal limit reversal used by bxp
for horizontal boxes. horizontal = FALSE gives a vertical
display, with VIMP/error on the vertical axis.
Use boxfill (or col) for the box fill,
border for its border, and whisklty and
whisklwd for interval whiskers. By default, boxes are red
when their lower endpoint is positive and blue otherwise. User
styles are retained when the intervals are redrawn above the guides.
outline = FALSE remains the default. The whiskers are
confidence limits and the boxes show the inner interval summaries,
rather than quartiles of the observed predictor values.
cex.axis, col.axis, and las control axis
labels. Set xaxt = "n" or yaxt = "n" to suppress
one axis, or axes = FALSE to suppress both. A scalar
ylab is a vertical-axis title. names supplies one
row label per selected statistic before sorting and trimming; labels
follow the same ordering as the statistics. at supplies
positions for the final displayed intervals.
show.plots = FALSE returns the plot data invisibly without
opening a graphics device. extract.subsample(x, raw = TRUE)
returns the interval matrices and replicate estimates directly.
Invisibly returns a boxplot-summary list. Its stats component
is the selected five-row confidence-interval matrix, in displayed
order, and names contains the row labels. The remaining
components are the boxplot scaffold, not additional subsampling
confidence intervals. The same list is returned with
show.plots = FALSE.
Ishwaran H. and Lu M. (2019). Standard errors and confidence intervals for variable importance in random forest regression, classification, and survival. Statistics in Medicine, 38, 558-582.
Politis, D.N. and Romano, J.P. (1994). Large sample confidence regions based on subsamples under minimal assumptions. The Annals of Statistics, 22(4):2031-2050.
Shao, J. and Wu, C.J. (1989). A general theory for jackknife variance estimation. The Annals of Statistics, 17(3):1176-1197.
# \donttest{
## Small settings are for illustration; increase B for final inference.
set.seed(19)
dta <- na.omit(airquality)
o <- rfsrc(Ozone ~ ., data = dta, ntree = 100,
importance = "permute", block.size = 1)
smp <- subsample(o, B = 25, verbose = FALSE)
## Three interval displays, all at the same confidence level.
plot.subsample(smp, alpha = .05, main = "Subsampling normal intervals")
plot.subsample(smp, alpha = .05, jknife = TRUE,
main = "Jackknife normal intervals")
plot.subsample(smp, alpha = .05, normal = FALSE,
main = "Nonparametric intervals")
## Restrict the display and customize the main axes and whiskers.
plot.subsample(smp, alpha = .05,
xvar.names = c("Solar.R", "Wind", "Temp"),
xlim = c(-.1, .8), las = 2, cex.axis = .75,
whisklty = 1, whisklwd = 1.5)
plot.data <- plot.subsample(smp, alpha = .05, show.plots = FALSE)
print(plot.data)
## Multivariate regression: one subsample bank, two response displays.
mv <- rfsrc(cbind(Ozone, Temp) ~ ., data = dta, ntree = 100,
importance = "permute", block.size = 1)
mv.smp <- subsample(mv, B = 25, verbose = FALSE)
plot.subsample(mv.smp, m.target = "Ozone", alpha = .05,
main = "Ozone")
plot.subsample(mv.smp, m.target = "Temp", alpha = .05,
main = "Temperature")
print(extract.subsample(mv.smp, m.target = "Temp", alpha = .05)$var.sel.Z)
# }