Skip to contents

Random forest for classification. Calls randomForest::randomForest() from randomForest.

Dictionary

This Learner can be instantiated via lrn():

lrn("classif.randomForest")

Meta Information

  • Task type: “classif”

  • Predict Types: “response”, “prob”

  • Feature Types: “logical”, “integer”, “numeric”, “factor”, “ordered”

  • Required Packages: mlr3, mlr3extralearners, randomForest

Parameters

IdTypeDefaultLevelsRange
ntreeinteger500\([1, \infty)\)
mtryinteger-\([1, \infty)\)
replacelogicalTRUETRUE, FALSE-
classwtuntypedNULL-
cutoffuntyped--
stratauntyped--
sampsizeuntyped--
nodesizeinteger1\([1, \infty)\)
maxnodesinteger-\([1, \infty)\)
importancecharacterFALSEaccuracy, gini, none-
localImplogicalFALSETRUE, FALSE-
proximitylogicalFALSETRUE, FALSE-
oob.proxlogical-TRUE, FALSE-
norm.voteslogicalTRUETRUE, FALSE-
do.tracelogicalFALSETRUE, FALSE-
keep.forestlogicalTRUETRUE, FALSE-
keep.inbaglogicalFALSETRUE, FALSE-
predict.alllogicalFALSETRUE, FALSE-
nodeslogicalFALSETRUE, FALSE-

References

Breiman, Leo (2001). “Random Forests.” Machine Learning, 45(1), 5–32. ISSN 1573-0565. doi:10.1023/A:1010933404324 .

See also

Author

pat-s

Super classes

mlr3::Learner -> mlr3::LearnerClassif -> LearnerClassifRandomForest

Methods

Inherited methods


LearnerClassifRandomForest$new()

Creates a new instance of this R6 class.


LearnerClassifRandomForest$importance()

The importance scores are extracted from the slot importance. Parameter 'importance' must be set to either "accuracy" or "gini".

Usage

LearnerClassifRandomForest$importance()

Returns

Named numeric().


LearnerClassifRandomForest$oob_error()

OOB errors are extracted from the model slot err.rate.

Usage

LearnerClassifRandomForest$oob_error()

Returns

numeric(1).


LearnerClassifRandomForest$clone()

The objects of this class are cloneable with this method.

Usage

LearnerClassifRandomForest$clone(deep = FALSE)

Arguments

deep

Whether to make a deep clone.

Examples

# Define the Learner
learner = lrn("classif.randomForest", importance = "accuracy")
print(learner)
#> 
#> ── <LearnerClassifRandomForest> (classif.randomForest): Random Forest ──────────
#> • Model: -
#> • Parameters: importance=accuracy
#> • Packages: mlr3, mlr3extralearners, and randomForest
#> • Predict Types: [response] and prob
#> • Feature Types: logical, integer, numeric, factor, and ordered
#> • Encapsulation: none (fallback: -)
#> • Properties: importance, multiclass, oob_error, twoclass, and weights
#> • Other settings: use_weights = 'use', predict_raw = 'FALSE'

# Define a Task
task = tsk("sonar")
# Create train and test set
ids = partition(task)

# Train the learner on the training ids
learner$train(task, row_ids = ids$train)

print(learner$model)
#> 
#> Call:
#>  randomForest(formula = formula, data = data, classwt = classwt,      cutoff = cutoff, importance = TRUE) 
#>                Type of random forest: classification
#>                      Number of trees: 500
#> No. of variables tried at each split: 7
#> 
#>         OOB estimate of  error rate: 17.27%
#> Confusion matrix:
#>    M  R class.error
#> M 62 11   0.1506849
#> R 13 53   0.1969697
print(learner$importance())
#>           V11            V9           V10           V12           V52 
#>  3.432791e-02  2.484085e-02  2.257789e-02  1.689687e-02  8.115034e-03 
#>           V36           V37           V46           V28            V8 
#>  7.889680e-03  5.467282e-03  5.138624e-03  5.017414e-03  4.688276e-03 
#>           V13           V45           V47           V27           V49 
#>  4.617652e-03  4.519573e-03  4.102127e-03  4.072028e-03  4.059520e-03 
#>           V31           V48            V4           V15           V16 
#>  3.831826e-03  3.574100e-03  3.397723e-03  3.359829e-03  3.357933e-03 
#>           V21            V5           V20           V17           V44 
#>  2.985433e-03  2.781510e-03  2.620652e-03  2.428804e-03  2.390239e-03 
#>           V43           V14            V3           V35           V51 
#>  2.384755e-03  2.278591e-03  2.084928e-03  2.017963e-03  1.969917e-03 
#>           V26           V59           V22           V38           V34 
#>  1.961566e-03  1.852687e-03  1.835732e-03  1.820780e-03  1.801656e-03 
#>           V23           V32           V18            V2           V24 
#>  1.769925e-03  1.679042e-03  1.240667e-03  9.992389e-04  8.778612e-04 
#>           V33           V58           V39            V6           V54 
#>  8.748840e-04  8.667787e-04  8.382077e-04  7.823353e-04  7.083503e-04 
#>           V30           V42           V55           V40           V25 
#>  6.915443e-04  5.190632e-04  4.695942e-04  4.290419e-04  4.162841e-04 
#>           V50           V57           V56           V53            V1 
#>  4.093770e-04  2.781657e-04  2.384577e-04 -4.382132e-05 -4.383751e-05 
#>           V19            V7           V60           V29           V41 
#> -8.741086e-05 -3.881665e-04 -4.890518e-04 -5.022002e-04 -6.212179e-04 

# Make predictions for the test rows
predictions = learner$predict(task, row_ids = ids$test)

# Score the predictions
predictions$score()
#> classif.ce 
#>   0.173913