impute.learn.rfsrc.RdLearns a predictive imputer from training data for later use on new data.
If the training data contain missing values, the function first
imputes them using impute. It then fits one saved full-sweep
learner per selected target on the completed training data and reuses
those learners later to update missing values in new data without
refitting on the test set.
The same saved learner bank can also be used to score new data for out-of-distribution (OOD) behavior. Note that OOD scores are available even when new data have missing values. Each selected target is reconstructed from its saved conditional learner and compared with the observed value. Target-wise discrepancies are calibrated against a training reference calculated from out-of-bag predictions computed during training.
If the training data are complete and target.mode = "all",
the initial training-data imputation step is skipped and the
full-sweep learners are fit directly from the complete training data.
If supervised.formula is supplied, the function also fits an
internal supervised forest from the training data. The supervised
forest is fit after training imputation and provides auxiliary
predictors for test-time imputation and OOD scoring. The auxiliary
predictors use supervised information learned from the training
outcomes. Supervised outcomes supplied with new data are dropped and
are not used at deployment time. Leave supervised.formula
unspecified to use the learned imputer without these auxiliary
predictors.
impute.learn.rfsrc(formula, data,
ntree = 100, nodesize = 1, nsplit = 10,
nimpute = 2, fast = FALSE, blocks,
mf.q, max.iter = 10, eps = 0.01,
ytry = NULL, always.use = NULL, verbose = TRUE,
...,
supervised.formula = NULL,
supervised.args = list(),
full.sweep.options = list(ntree = 100, nsplit = 10),
target.mode = c("missing.only", "all"),
deployment.xvars = NULL,
anonymous = TRUE,
learner.prefix = "impute.learner.",
learner.root = "learners",
out.dir = NULL,
wipe = TRUE,
keep.models = is.null(out.dir),
keep.ximp = FALSE,
save.on.fit = !is.null(out.dir),
save.ood = TRUE,
weight = NULL)
save.impute.learn.rfsrc(object, path, wipe = TRUE, verbose = TRUE)
load.impute.learn.rfsrc(path, targets = NULL, lazy = TRUE, verbose = TRUE)
# S3 method for class 'impute.learn.rfsrc'
predict(object, newdata,
max.predict.iter = 3L,
eps = 1e-3,
targets = NULL,
restore.integer = TRUE,
cache.learners = c("session", "none", "all"),
verbose = TRUE,
...)
impute.ood.rfsrc(object, newdata,
targets = NULL,
max.predict.iter = 3L,
eps = 1e-3,
cache.learners = c("all", "session", "none"),
weight = NULL,
aggregate = c("bounded.product", "weighted.mean",
"weighted.lp", "weighted.lp.log", "top.k"),
aggregate.args = list(),
return.details = FALSE,
return.reconstruction = FALSE,
verbose = TRUE,
...)An optional symbolic model description passed to
impute for the initial training-data imputation when
supervised.formula is not supplied. It follows the same
guidelines as in impute. This formula does not specify the
full-sweep learner bank, the OOD targets, or the supervised
auxiliary forest. When supervised.formula is supplied, the
raw predictor block is taken from the right-hand side of
supervised.formula, and formula is not used.
Training data, converted to a plain data frame before
processing. Matrices, tibbles, and data.table objects can be
supplied. Column names must be unique and nonempty, each column
must be a vector, and numeric values must be finite or missing.
Bare vectors and nested matrix or list columns are not supported.
Variables that are not real-valued are
coerced to factors before fitting when possible; otherwise fitting
stops with an error. Rows and columns that are entirely missing are
dropped before training begins. If supervised.formula is
supplied, data should contain both the supervised response
columns and the raw predictor columns. The learned imputer is then
built on the raw predictor block defined by the right-hand side of
supervised.formula.
Arguments passed to
impute for the initial training-data imputation. The
argument full.sweep is controlled internally and should not
be supplied. fast is also forwarded to the saved target
learners and, unless overridden by supervised.args, to the
supervised forest. verbose controls progress throughout
training, prediction, and scoring. The same verbose, max.predict.iter,
eps, and cache.learners controls are also used by
predict.impute.learn and impute.ood.
Iteration and forest-size controls must be scalar integers in their
permitted ranges; max.iter must be at least one.
Controls the imputation engine used by impute.
With mf.q = 1 and always.use = NULL, targets are
updated one at a time using missForest. Other positive
settings use the multivariate missForest generalization;
a fraction below one controls the proportion of missing-data
variables grouped as responses, and a value above one specifies
their requested group size. A non-NULL always.use
selects the multivariate branch, including an empty vector or a
vector with no matching column names. If mf.q is omitted,
training uses on-the-fly imputation. The selected method is recorded
in manifest$train.imputation.
Finite nonnegative convergence threshold. In impute.learn this
controls the initial training-data imputation. In
predict.impute.learn and impute.ood it controls early
stopping for the prediction-time sweep.
For impute.learn, additional arguments passed to
impute. For predict.impute.learn and
impute.ood, additional arguments are currently ignored.
Optional supervised learning formula used to
augment the learned imputer for improved OOD detection in supervised
settings. The left-hand side defines the supervised response and
the right-hand side defines the raw predictor block to be learned by
impute.learn. The supervised forest is fit internally after
the raw predictor block has been completed. Its training out-of-bag
predicted values and test-time predicted values are
appended internally as auxiliary predictors. Training auxiliary
values use out-of-bag predictions when available, with ordinary
forest predictions and training means used as fallbacks. At
deployment, auxiliary values are computed once from the initialized
raw predictors and held fixed during the imputation passes.
Predictors must be raw column names; non-syntactic names can be
enclosed in backticks. Interactions and transformations are not
supported. These auxiliary predictors are internal and are not expected in
newdata.
Optional named list of arguments passed to the
internal supervised rfsrc fit. Entries for formula,
data, and forest are controlled internally and are
ignored if supplied. All other entries override the defaults used
for this supervised fit, including defaults inherited from
top-level arguments such as fast and internally supplied
defaults such as perf.type = "none".
A named list of options used when fitting
saved target learners after initial training-data imputation.
Defaults are ntree = 100, nodesize = NULL, and
nsplit = 10, independently of the corresponding settings
used for initial imputation. Recognized entries include ntree, nodesize,
nsplit, mtry, splitrule, bootstrap,
sampsize, samptype, perf.type, rfq,
save.memory, importance, and proximity.
Unknown entries are ignored with a warning. Duplicate names produce
a warning, and the last entry for each name is used.
Determines which raw variables receive a saved
full-sweep learner. The default "missing.only" saves
learners only for variables that were missing in the training data.
The option "all" saves a learner for every retained raw
variable. If the training data are complete, target.mode =
"all" must be used. When supervised.formula is supplied,
the default is promoted internally to "all" so that every
retained raw predictor can later be updated and scored. For the
broadest OOD coverage, target.mode = "all" is recommended so
every deployment-time variable can be reconstructed.
Controls which raw predictors are assumed to
be available later when the saved imputer is used on new data. If
NULL, all raw non-target columns are used. If a character
vector, the same raw predictor set is used for all targets. If a
named list, names must be target variables and each entry gives the
raw predictor set for that target; targets omitted from the list
fall back to all raw non-target columns. Unnamed or duplicated list
entries are not allowed. Entries named for non-target variables, and
predictor names not found in the training data, are ignored with a
warning. When supervised.formula is supplied, the internal
auxiliary predicted.* columns are added automatically on top
of these raw predictor sets. In other words,
deployment.xvars restricts the raw predictors only.
If TRUE, uses rfsrc.anonymous when
fitting the saved target-wise full-sweep learner bank. The internal
supervised forest, when requested through supervised.formula,
is fit with rfsrc because it must be retained for later
prediction-time auxiliary variables. Thus anonymous usually
reduces the size of the target-wise learner bank, but not
necessarily the size of the supervised forest.
Names used when writing saved
full-sweep learners to disk. If supervised.formula is
supplied, the internal supervised forest is saved under the same
learner root. learner.root must be a relative directory path
within the imputer directory. learner.prefix must be a single
file-name component. Absolute paths and parent-directory components
are not allowed.
Optional output directory. If supplied and
save.on.fit = TRUE, the manifest and the saved full-sweep
learners are written to this directory during fitting. This requires
the fst package because learners are serialized with
fast.save. If supervised.formula is supplied, the
internal supervised forest is also saved there.
If TRUE, replaces an existing output directory
after the new imputer has been written and verified. If FALSE,
unrelated existing files are retained. Saving uses a temporary
bundle and requires additional disk space.
If TRUE, keeps the fitted full-sweep
learners in memory in the returned object. At least one storage mode
must be enabled: either keep.models = TRUE or
out.dir with save.on.fit = TRUE. If
supervised.formula is supplied, the internal supervised
forest is kept in memory under the same rule.
If TRUE, keeps the completed raw training
predictor table in the returned object. This is not required for
later prediction.
If TRUE and out.dir is supplied,
writes the imputer to disk during fitting.
If TRUE, computes and stores an OOD reference.
The reference is built during training from target-wise out-of-bag
reconstruction discrepancies, their target-wise calibrated training
scores, and a default row-level weighted mean using the saved OOD
target weights. If no weight is supplied at fit time, equal
target weights are used. The saved target-wise training-score matrix
allows impute.ood to rebuild a calibrated row-level
percentile for arbitrary target subsets, test-time weight overrides,
and alternate row aggregates besides the weighted mean.
An object returned by impute.learn or
load.impute.learn.
Directory containing a saved imputer. Use a dedicated
directory rather than a filesystem root, home directory, working
directory, or an ancestor of these. A complete bank can be saved
back to its source path. Otherwise, source and destination directories
must not contain one another. An object loaded with a target subset
must be saved to a different directory. Save and load operations
require the fst package because learners are read and written
with fast.save and fast.load.
Optional subset of target variables to load, update,
or score. Unknown names are ignored with a warning. For
impute.ood, row-level percentile calibration is rebuilt for
the requested target subset from the saved target-wise training OOD
scores whenever those scores are available in the manifest.
A load-time subset limits which learners are available for predictor
completion. It can therefore differ from scoring a subset of a fully
loaded bank when other predictors are missing.
If TRUE, saved learners are loaded only when they
are needed. If FALSE, all saved learners are loaded at once.
If supervised.formula was used at fit time, the internal
supervised forest follows the same lazy versus eager loading rule.
New data to be imputed or scored, converted to a
plain data frame before processing. The column-name and vector-column
requirements for data also apply. Missing columns are added
and extra columns are dropped to match the training schema. A zero-row
table returns correctly typed empty output without calling forests.
Retained numeric values must be finite or missing. Nonmissing values
that cannot be converted to a required numeric type become missing
with a warning; their row indices are recorded in
conversion.issues.
Unseen factor levels are converted to NA for harmonization,
but they are also tracked row-wise. In supervised mode,
newdata should contain only the raw predictor columns learned
by the imputer. The internal auxiliary columns are created
automatically and should not be supplied by the user. If supervised
response columns are supplied in newdata and are not part of
the learned raw predictor block, they are treated as extra columns
and dropped; they are not used to compute auxiliary predictors,
imputations, or OOD scores. In impute.ood,
observed raw target values are used to compute target-wise
discrepancies. Missing raw target values are completed for
predictor-side reconstruction but do not themselves contribute a
finite target-wise OOD score. Any row containing an unseen factor
level is flagged and its row-level OOD score is set to the maximum
value.
Maximum number of full-sweep passes applied to
newdata before the saved learner bank is used for OOD
reconstruction or returned prediction-time imputations. Must be a
nonnegative integer. Zero performs initialization without iterative
target updates; OOD reconstruction is still performed.
If TRUE, generated imputations in
integer-supported variables are rounded to the integer grid. This
includes columns stored as R integers and numeric columns whose
observed finite training values were all integer-valued. Observed
numeric values in newdata are preserved, including fractional
values, both as predictors during imputation and in the returned data.
Restoration applies to initialization values and model-based updates.
Numeric training columns retain numeric storage. Columns stored as
R integers during training are returned as integer vectors only when
every nonmissing completed value is exactly integral and within R's
integer range; otherwise numeric storage preserves observed fractions
and out-of-range values. Set to FALSE to retain unrounded
generated numeric values.
Factor columns are always conformed back to the training schema.
The package operates on real-valued and factor variables;
inputs that are not real-valued are coerced to factors during
preprocessing when possible, otherwise an error is raised.
How saved learners are reused during
prediction or OOD scoring. For predict.impute.learn, the
default "session" loads each needed learner once per call.
The option "none" reloads a learner every time it is needed.
The option "all" loads all requested learners before work
starts. For impute.ood, "all" is the default because
the saved learner bank is typically reused once for predictor-side
completion and again for target reconstruction.
Optional nonnegative target weights used for row-level OOD
aggregation. In impute.learn, these weights define the
default row-level OOD weighting scheme stored in the manifest. In
impute.ood, they define the active row-level weighting scheme
for the current scoring call. If supplied as a named vector, entries
are matched to targets by name; omitted targets are set to zero, and
extra names are ignored. If omitted in impute.ood, the saved
training-time OOD weights are used automatically. Because the fit
stores target-wise training OOD scores, score.percentile can
be recalibrated automatically for test-time weight overrides rather
than being limited to the original training-time weights.
Row-level aggregation metric used by
impute.ood to combine calibrated target-wise OOD scores. The
default "bounded.product" applies a weighted product of
the form \(1-\prod_j \max(1-u_j,\varepsilon)^{\tilde w_j}\),
where \(u_j\) is the calibrated score for target \(j\) and
\(\tilde w_j\) denotes weights normalized over the positively
weighted, scoreable targets in the row. "weighted.mean" is the weighted
average. "weighted.lp" applies a weighted Minkowski \(L_p\)
aggregation to the calibrated target scores.
"weighted.lp.log" first applies the tail-stretching transform
\(-\log\{\max(1-u_j,\varepsilon)\}\) to each calibrated target score
\(u_j\) and then applies the weighted \(L_p\)
aggregation. "top.k" averages only the \(k\) largest
scoreable target scores among the positively weighted targets. When
the fit stores target-wise training OOD scores,
score.percentile is rebuilt for the requested aggregate as
well as the requested targets and weights.
Optional named list of tuning arguments for
aggregate. Recognized entries are p for
"weighted.lp" and "weighted.lp.log", k
(or top.k) for "top.k", and eps for
"weighted.lp.log" and "bounded.product". The
default values are p = 2, k = 1, and
eps = 1e-12.
If TRUE, impute.ood returns the
per-target discrepancy and calibrated target-score matrices and a
richer info list.
If TRUE, impute.ood returns
the target-wise reconstructed values used during OOD scoring. The
reconstructed values are returned on the raw target scale, with one
column per scored target. Typically these values are computed
internally and discarded.
A predictive imputer is calculated in two stages.
The training data are first processed. Variables that are not real-valued are coerced to factors when possible; otherwise fitting stops with an error. Rows and columns that are entirely missing are removed before the training schema is stored.
If the resulting training data contain missing values, the first stage
uses impute as the imputation engine to complete the training
data. Options are specified exactly as in impute. In
particular, mf.q = 1 with always.use = NULL
updates one target at a time, other positive settings give the
multivariate missForest generalization, and if mf.q
is omitted, on the fly imputation
is used if formula is specified, otherwise default unsupervised
imputation is used. If the training data are already complete and
target.mode = "all", this initial imputation step is skipped.
In the second stage, a full forest sweep is fit on the completed
training data. For each target selected by target.mode, rows
where that target was observed are used to fit a forest with that
target on the left-hand side and the predictors selected by
deployment.xvars on the right-hand side. These saved forests
all use the same completed training table; fitting the learner bank
does not further update that table.
By default, deployment.xvars = NULL allows every non-target
column to be used as a predictor. This is convenient, but it can also
introduce leakage if the training data include outcomes, future-only
variables, identifiers, or any fields that will not be available when
the learned imputer is applied to new data. Restrict
deployment.xvars when that is a concern.
If supervised.formula is supplied, the function also fits an
internal supervised forest from the training data. Out-of-bag
predicted values for the training sample, and predicted
values for new data, are appended internally as auxiliary
predictors. These auxiliary variables are not used as imputation
targets, however, they are available as predictors for the saved
forests and can therefore influence both test-time imputation and OOD
scoring. Users who want only the basic unsupervised learned imputer
should leave supervised.formula unspecified.
When supervised.formula is used, deployment.xvars
continues to restrict only the raw predictors. The internally created
auxiliary variables are added automatically to the predictor sets for
the saved raw targets. This lets the learned imputer benefit from
supervised signal without asking the user to prepare and pass those
auxiliary columns manually.
When the imputer is saved to disk, each full-sweep learner is written
separately using fast.save. Loading uses fast.load. In
practice this gives a small manifest plus a directory of saved
learners. The fst package is therefore required for save and
load operations. The explicit save method can write learners either
from memory or by reloading them from an attached saved path. If
supervised.formula is used, the internal supervised forest is
saved alongside the target-wise learner bank. Both training-time and
explicit saving write and read back the new learners before replacing
an existing destination. A failed staging operation leaves the previous
saved imputer unchanged. If replacement fails, restoration of the
previous directory is attempted; an unrecovered backup path is reported.
Training stops if every requested target learner fails. A partial bank
is returned with one summary warning when some learners succeed.
manifest$learners records each target's status and error,
and printing the imputer reports successful and unavailable learners.
Prediction starts by matching newdata to the training schema,
filling missing values with training means or modes, and then applying
one or more full-sweep passes. Only the targets selected by
target.mode are updated by saved learners. In supervised mode,
this same test-time preparation first initializes the raw predictor
table, then computes the internal auxiliary variables from the saved
supervised forest, and then runs the forest test-time sweep. The
auxiliary predictors remain fixed throughout these passes.
Supervised response columns, if present in newdata but not in
the learned raw predictor block, are dropped before this step and play
no role at prediction or scoring time. A target update requires one
valid prediction per requested row. Failed or unavailable predictions
retain the preceding imputed values and are recorded in
target.issues. A pass with no valid model updates is reported
separately from convergence.
If target.mode = "missing.only", a variable that was complete
in training but missing in new data is initialized from the training
fit but does not receive a model-based update. Use
target.mode = "all" if missing values may appear later in any
raw variable. Complete training data also require
target.mode = "all", because otherwise there are no missing
variables from which to determine the saved targets.
If save.ood = TRUE, the fit also stores an OOD reference in the
manifest. For each saved target, the out-of-bag prediction from the
fitted learner is compared with the observed training value to form a
target-wise reconstruction discrepancy. Continuous and integer targets
use absolute reconstruction error. Factor targets use negative log
predictive probabilities. Unavailable or invalid probabilities remain
missing discrepancies; a genuine zero probability is scored using
the probability floor. Each learner entry records n.oob.finite
and n.oob.nonfinite when OOD reference construction is enabled.
The row-level OOD calibration stored at fit time is built by
aggregating the target-wise training scores with a weighted mean using
weight. If weight is omitted at fit time, all saved OOD
targets receive weight 1. If a named vector is supplied, entries are
matched by target name, omitted saved OOD targets receive weight 0,
and the resulting weighting scheme is carried forward for later
deployment-time scoring.
impute.ood first completes the predictor side of newdata
using the same harmonization, initialization, and iterative sweep
logic used by predict.impute.learn. It then reconstructs each
requested raw target directly from its saved learner and compares the
reconstruction with the observed raw target value supplied in
newdata. If a scored raw target is missing in a row, that
target does not contribute to that row's OOD score. Raw target
discrepancies are converted to target-wise OOD scores using the saved
target-specific training references. Discrepancies strictly above a
nonempty reference's maximum, including positive infinity, receive its
largest stored probability. The existing convention for ties is retained,
including equality at the largest quantile. Missing discrepancies and
empty references remain unscored. In supervised mode, auxiliary
variables participate as predictors in the reconstruction, but the
supervised response itself is not used at deployment time.
The row-level OOD score combines those calibrated target-wise scores
over the targets that are both observed and scoreable for that row. By
default, impute.ood uses a bounded product rule, but the row
aggregate can be changed to weighted mean, a weighted \(L_p\) rule,
a log-tail weighted \(L_p\) rule, or a top-\(k\) rule. This makes
it possible to explore row scores that are more sensitive to sparse
but severe coordinate shifts. By default, impute.ood reuses
the same OOD weights saved during impute.learn, so a pipeline
can fix its weighting scheme once upstream and carry it forward
automatically.
A second component, score.percentile, is obtained by rebuilding
the row-level training reference from the saved target-wise training
OOD scores using the requested target subset, the active weight
vector, and the active row aggregate. This means percentile
calibration remains available when the user leaves the saved weights
in place, overrides them at test time, scores only a subset of the
saved OOD targets, or experiments with alternate row aggregates.
Unseen factor levels are tracked row-wise during harmonization.
Because such values are immediate anomalies relative to the training
schema, impute.ood flags those rows and assigns them the
maximum row-level score. If the unseen level occurs in a scored target
itself, the corresponding target-level discrepancy is also treated as
maximal.
impute.learn returns an object of class
c("impute.learn.rfsrc", "impute.learn"). The object
contains a manifest, optionally the fitted full-sweep learners,
optionally the internal supervised forest when
supervised.formula is used, optionally the completed raw
training predictor table, and optionally a path to the saved imputer
on disk. If save.ood = TRUE, the manifest also contains an
ood component storing compact target-wise OOD references, the
saved row-by-target training OOD score matrix used for later
percentile recalibration, and the default OOD aggregation weights.
When supervised mode is active, the manifest also records the
supervised family, response names, and the internally created
auxiliary predicted.* variable names.
load.impute.learn returns an object of the same class.
predict.impute.learn returns a data frame with imputed values
overlaid on the raw predictor table, retaining its row names. An attribute named
"impute.learn.info" contains prediction-time diagnostics such
as the number of sweep passes, pass-difference history, caching mode,
disk-load counts, schema harmonization details, dropped supervised
response columns when present, row-wise unseen-factor flags,
supervised-auxiliary diagnostics when present, and any targets
skipped because a learner was unavailable or a prediction failed.
pass.updated.cells and pass.failed.cells count accepted
updates and failed updates in each pass; converged and
stopping.reason distinguish convergence from initialization-only,
empty-input, iteration-limit, and failed-update stopping.
conversion.issues records the row indices of nonmissing values
that became missing during numeric conversion. In both prediction and
OOD diagnostics, n.disk.loads counts successful target-learner
load operations, including repeated loads with cache.learners = "none".
disk.load.targets lists the distinct targets loaded. Supervised
forest loading is reported separately in info$supervised and is
not included in the target-learner count.
impute.ood returns an object of class
c("impute.ood.rfsrc", "impute.ood"). It is a list with the
following components:
score: the row-level aggregate of calibrated
target-wise OOD scores under the requested aggregate and
weight. Larger values indicate greater
out-of-distribution behavior.
score.percentile: the percentile of score
relative to a row-level training reference rebuilt from the saved
target-wise training OOD scores for the requested targets,
weights, and row aggregate. For legacy fitted objects that do not
contain those saved training scores, the original saved row-level
reference is used when possible; otherwise NA.
targets.used: the number of weighted targets that
contributed to each row-level score.
target.score: optional matrix of target-wise calibrated
OOD scores, returned when return.details = TRUE.
target.delta: optional matrix of raw target-wise
reconstruction discrepancies, returned when
return.details = TRUE.
target.reconstruction: when
return.reconstruction = TRUE, a data frame containing
the saved learners' predictions for the scored targets.
reconstructed.data: when
return.reconstruction = TRUE, the harmonized raw table
with scored targets replaced by their reconstructions. Other
columns retain their harmonized values before schema restoration.
completed.data: when both return.details
and return.reconstruction are TRUE, the raw
predictor table after prediction-time imputation. This is distinct
from the target reconstruction table.
info: a list of diagnostics including harmonization
details, dropped supervised response columns when present,
row-wise unseen-factor flags, learner-loading information,
supervised-auxiliary diagnostics when present, the active row
aggregate and its arguments, whether the saved row-level
calibration was used, and any target-specific issues.
Stekhoven D.J. and Buhlmann P. (2012). MissForest–non-parametric missing value imputation for mixed-type data. Bioinformatics, 28(1):112–118.
Tang F. and Ishwaran H. (2017). Random forest missing data algorithms. Statistical Analysis and Data Mining, 10:363–377.
## ------------------------------------------------------------
## small data example: uses missForest for impute engine
## ------------------------------------------------------------
set.seed(101)
aq <- airquality[, c("Ozone", "Solar.R", "Wind", "Temp", "Month")]
aq$Month <- factor(aq$Month)
id <- sample(1:nrow(aq), 100)
train <- aq[id, ]
test <- aq[-id, ]
## training the imputer
fit <- impute.learn(
data = train,
ntree = 25,
mf.q = 1,
max.iter = 5,
full.sweep.options = list(ntree = 25, nsplit = 5)
)
## test time imputation
test.imp <- predict(fit, test, max.predict.iter = 2, verbose = FALSE)
print(head(test.imp))
# \donttest{
## OOD scoring is most informative when every deployment-time
## variable can be reconstructed, so target.mode = "all" is recommended.
## Optional named OOD weights can also be supplied here. Any omitted
## targets receive weight 0, and the saved weights are reused
## automatically later by impute.ood().
ood.fit <- impute.learn(
data = train,
ntree = 25,
mf.q = 1,
max.iter = 5,
target.mode = "all",
save.ood = TRUE,
full.sweep.options = list(ntree = 25, nsplit = 5),
verbose = FALSE
)
ood <- impute.ood(ood.fit, test, return.details = TRUE, verbose = FALSE)
print(head(ood$score))
print(head(ood$score.percentile))
## try a more spike-sensitive row aggregate
ood.lp <- impute.ood(ood.fit, test,
aggregate = "weighted.lp",
aggregate.args = list(p = 4),
verbose = FALSE)
print(head(ood.lp$score.percentile))
## ------------------------------------------------------------
## supervised OOD example: regression benchmark
## the user supplies only the raw x-columns at test time
## ------------------------------------------------------------
friedman1_sim <- function(n = 150, p = 10, sigma = 1) {
X <- matrix(runif(n * p), nrow = n)
y <- 10 * sin(pi * X[, 1] * X[, 2]) +
20 * (X[, 3] - 0.5)^2 +
10 * X[, 4] + 5 * X[, 5] +
rnorm(n, sd = sigma)
list(X = X, y = y)
}
trn <- data.frame(friedman1_sim())
tst <- data.frame(friedman1_sim())
xvars <- setdiff(names(trn), "y")
## impute data using missForest, construct a supervised forest
## - supervised forests are used to create auxiliary variables
## - improves test time OOD in supervised problems
sup.fit <- impute.learn(
data = trn,
mf.q = 1,
supervised.formula = y ~ .,
supervised.args = list(ntree = 50, nsplit = 5),
full.sweep.options = list(ntree = 25, nsplit = 5),
save.ood = TRUE,
verbose = FALSE
)
## add some missing values to the test data
xnew <- tst[, xvars, drop = FALSE]
xnew[sample(seq_len(nrow(xnew)), 5), xvars[1]] <- NA
xnew[sample(seq_len(nrow(xnew)), 5), xvars[2]] <- NA
## imputation
xnew.imp <- predict(sup.fit, xnew, max.predict.iter = 2, verbose = FALSE)
print(head(xnew.imp))
## OOD score
ood.sup <- impute.ood(sup.fit, xnew, verbose = FALSE)
print(head(ood.sup$score.percentile))
## ------------------------------------------------------------
## Save the learned imputer to disk and load it later.
## This explicit save example writes learners kept in memory.
## Uses missForest for the impute engine.
## ------------------------------------------------------------
bundle.dir <- file.path(tempdir(), "aq.imputer")
fit <- impute.learn(
data = train,
ntree = 25,
mf.q = 1,
max.iter = 5,
full.sweep.options = list(ntree = 25, nsplit = 5),
keep.models = TRUE,
verbose = FALSE
)
save.impute.learn(fit, bundle.dir, verbose = FALSE)
imp <- load.impute.learn(bundle.dir, lazy = TRUE, verbose = FALSE)
test.imp <- predict(imp, test, max.predict.iter = 2, verbose = FALSE)
unlink(bundle.dir, recursive = TRUE)
## ------------------------------------------------------------
## Challenging example with factors, uses save/reload
## ------------------------------------------------------------
## load pbc, convert everything to factors
data(pbc, package = "randomForestSRC")
dta <- data.frame(lapply(pbc, factor))
dta$days <- pbc$days
dta$status <- dta$status
## split the data into unbalanced train/test data (25/75)
## the train/test data have the same levels, but different labels
idx <- sample(1:nrow(dta), round(nrow(dta) * .25))
train <- dta[idx,]
test <- dta[-idx,]
## even harder ... factor level not previously encountered in training
levels(test$stage) <- c(levels(test$stage), "fake")
test$stage[sample(seq_len(nrow(test)), 10)] <- "fake"
## train forest
fit <- suppressWarnings(
impute.learn(Surv(days, status) ~ ., train,
target.mode = "all",
save.ood = TRUE,
keep.models = TRUE)
)
## save/reload
bundle.dir <- file.path(tempdir(), "pbc.imputer")
save.impute.learn(fit, bundle.dir, verbose = FALSE)
imp <- load.impute.learn(bundle.dir, lazy = TRUE, verbose = FALSE)
test.imp <- predict(imp, test, max.predict.iter = 2, verbose = FALSE)
ood <- impute.ood(imp, test, return.details = TRUE, verbose = FALSE)
print(which(ood$info$unseen.rows))
print(summary(test.imp))
unlink(bundle.dir, recursive = TRUE)
# }