Display confidence intervals for variable importance (VIMP) from a subsample object. Select predictors, an outcome, and an interval method using the saved replicate estimates. Prediction-error and joint-VIMP rows can also be displayed when requested during subsampling.

# S3 method for class 'rfsrc'
plot.subsample(x, alpha = .01, xvar.names,
 standardize = TRUE, normal = TRUE, jknife = FALSE, target, m.target = NULL,
 pmax = 75, main = "", sorted = TRUE, show.plots = TRUE, ...)

Arguments

x

An object returned by subsample.rfsrc.

alpha

Significance level for the displayed intervals, whose nominal confidence level is \(1-\alpha\). The plotting default .01 gives 99 percent intervals. Use alpha = .05 to match the default of extract.subsample and the print method.

xvar.names

Names of predictor rows to display. If omitted, all available rows are considered, subject to pmax. This selects stored results and does not regrow the forest.

standardize

For regression outcomes, divide VIMP by the variance of the selected response in the full training data. The same variance is used for all replicate estimates and any error row. Other families are unchanged. Set FALSE to display the original error scale.

normal

Use normal-approximation intervals when TRUE and nonparametric intervals when FALSE. For double-bootstrap objects, these are bootstrap normal and percentile intervals.

jknife

Select the delete-\(d\) jackknife standard error for normal intervals. Use normal = TRUE, jknife = TRUE for this display. The default normal display uses the subsampling standard error. Ignored when normal = FALSE and for double-bootstrap objects.

target

For classification, an integer or class label selecting class-specific VIMP; the default is overall VIMP. For competing risks, an integer from 1 to J selecting an event type; the default is the first event. This selects a statistic within the outcome identified by m.target.

m.target

Name of one response in a multivariate or mixed-outcome forest. If omitted, a default response is selected. This selects an outcome in the saved results; it is separate from choosing a class or event with target.

pmax

Maximum number of rows retained for display, selected by VIMP.

main

Main plot title.

sorted

Order the displayed rows by importance.

show.plots

Draw the plot. Set FALSE to obtain its plot-data return value without opening or updating a plot.

...

Graphical arguments for the interval display, including xlim, ylim, xlab, ylab, boxfill, border, whisklty, whisklwd, and axis controls. A named pars list is also accepted; direct arguments take precedence. See Details.

Details

Choosing an interval display

The default uses normal intervals and the subsampling standard error. Use normal = TRUE, jknife = TRUE for normal intervals with the jackknife standard error, or normal = FALSE for nonparametric intervals. All use the estimates already stored in x; changing alpha or the display method does not perform another round of subsampling.

The subsampling standard error measures replicate dispersion around the replicate mean. The jackknife calculation measures dispersion around the full-data estimate, incorporating the displacement between those two centers. The jackknife standard error is not necessarily larger for a finite replicate set. The nonparametric interval reverses empirical quantiles of the centered subsampling roots. See subsample.rfsrc for the calculations.

The display summarizes uncertainty in VIMP, not the distribution of the observed predictor values. Its intervals are calculated separately for each displayed statistic. A positive lower VIMP endpoint identifies a positive-importance result under that interval procedure; the display does not apply a multiple-testing correction. Prediction-error rows summarize the error itself.

Selecting predictors and outcomes

xvar.names selects rows and m.target selects a response. For a classification response or competing-risk outcome, target additionally selects its class-specific or event-specific statistic. These selections are made after resampling, so one saved object can be used for several displays.

Set alpha explicitly when comparing results with print or extract.subsample; their default is .05, whereas the plotting default is .01.

Graphical customization

The default display is horizontal. xlim controls its horizontal VIMP/error axis and ylim controls its vertical row positions. These arguments always refer to the displayed axes; the wrapper handles the internal limit reversal used by bxp for horizontal boxes. horizontal = FALSE gives a vertical display, with VIMP/error on the vertical axis.

Use boxfill (or col) for the box fill, border for its border, and whisklty and whisklwd for interval whiskers. By default, boxes are red when their lower endpoint is positive and blue otherwise. User styles are retained when the intervals are redrawn above the guides. outline = FALSE remains the default. The whiskers are confidence limits and the boxes show the inner interval summaries, rather than quartiles of the observed predictor values.

cex.axis, col.axis, and las control axis labels. Set xaxt = "n" or yaxt = "n" to suppress one axis, or axes = FALSE to suppress both. A scalar ylab is a vertical-axis title. names supplies one row label per selected statistic before sorting and trimming; labels follow the same ordering as the statistics. at supplies positions for the final displayed intervals.

show.plots = FALSE returns the plot data invisibly without opening a graphics device. extract.subsample(x, raw = TRUE) returns the interval matrices and replicate estimates directly.

Value

Invisibly returns a boxplot-summary list. Its stats component is the selected five-row confidence-interval matrix, in displayed order, and names contains the row labels. The remaining components are the boxplot scaffold, not additional subsampling confidence intervals. The same list is returned with show.plots = FALSE.

Author

Hemant Ishwaran and Udaya B. Kogalur

References

Ishwaran H. and Lu M. (2019). Standard errors and confidence intervals for variable importance in random forest regression, classification, and survival. Statistics in Medicine, 38, 558-582.

Politis, D.N. and Romano, J.P. (1994). Large sample confidence regions based on subsamples under minimal assumptions. The Annals of Statistics, 22(4):2031-2050.

Shao, J. and Wu, C.J. (1989). A general theory for jackknife variance estimation. The Annals of Statistics, 17(3):1176-1197.

Examples

# \donttest{
## Small settings are for illustration; increase B for final inference.
set.seed(19)
dta <- na.omit(airquality)
o <- rfsrc(Ozone ~ ., data = dta, ntree = 100,
            importance = "permute", block.size = 1)
smp <- subsample(o, B = 25, verbose = FALSE)

## Three interval displays, all at the same confidence level.
plot.subsample(smp, alpha = .05, main = "Subsampling normal intervals")
plot.subsample(smp, alpha = .05, jknife = TRUE,
                 main = "Jackknife normal intervals")
plot.subsample(smp, alpha = .05, normal = FALSE,
                 main = "Nonparametric intervals")

## Restrict the display and customize the main axes and whiskers.
plot.subsample(smp, alpha = .05,
                 xvar.names = c("Solar.R", "Wind", "Temp"),
                 xlim = c(-.1, .8), las = 2, cex.axis = .75,
                 whisklty = 1, whisklwd = 1.5)
plot.data <- plot.subsample(smp, alpha = .05, show.plots = FALSE)
print(plot.data)

## Multivariate regression: one subsample bank, two response displays.
mv <- rfsrc(cbind(Ozone, Temp) ~ ., data = dta, ntree = 100,
             importance = "permute", block.size = 1)
mv.smp <- subsample(mv, B = 25, verbose = FALSE)
plot.subsample(mv.smp, m.target = "Ozone", alpha = .05,
                 main = "Ozone")
plot.subsample(mv.smp, m.target = "Temp", alpha = .05,
                 main = "Temperature")
print(extract.subsample(mv.smp, m.target = "Temp", alpha = .05)$var.sel.Z)
# }