Classification Random Forest Learner
Source:R/learner_randomForest_classif_randomForest.R
mlr_learners_classif.randomForest.RdRandom forest for classification.
Calls randomForest::randomForest() from randomForest.
Meta Information
Task type: “classif”
Predict Types: “response”, “prob”
Feature Types: “logical”, “integer”, “numeric”, “factor”, “ordered”
Required Packages: mlr3, mlr3extralearners, randomForest
Parameters
| Id | Type | Default | Levels | Range |
| ntree | integer | 500 | \([1, \infty)\) | |
| mtry | integer | - | \([1, \infty)\) | |
| replace | logical | TRUE | TRUE, FALSE | - |
| classwt | untyped | NULL | - | |
| cutoff | untyped | - | - | |
| strata | untyped | - | - | |
| sampsize | untyped | - | - | |
| nodesize | integer | 1 | \([1, \infty)\) | |
| maxnodes | integer | - | \([1, \infty)\) | |
| importance | character | FALSE | accuracy, gini, none | - |
| localImp | logical | FALSE | TRUE, FALSE | - |
| proximity | logical | FALSE | TRUE, FALSE | - |
| oob.prox | logical | - | TRUE, FALSE | - |
| norm.votes | logical | TRUE | TRUE, FALSE | - |
| do.trace | logical | FALSE | TRUE, FALSE | - |
| keep.forest | logical | TRUE | TRUE, FALSE | - |
| keep.inbag | logical | FALSE | TRUE, FALSE | - |
| predict.all | logical | FALSE | TRUE, FALSE | - |
| nodes | logical | FALSE | TRUE, FALSE | - |
References
Breiman, Leo (2001). “Random Forests.” Machine Learning, 45(1), 5–32. ISSN 1573-0565. doi:10.1023/A:1010933404324 .
See also
as.data.table(mlr_learners)for a table of available Learners in the running session (depending on the loaded packages).Chapter in the mlr3book: https://mlr3book.mlr-org.com/chapters/chapter2/data_and_basic_modeling.html#sec-learners
mlr3learners for a selection of recommended learners.
mlr3cluster for unsupervised clustering learners.
mlr3pipelines to combine learners with pre- and postprocessing steps.
mlr3tuning for tuning of hyperparameters, mlr3tuningspaces for established default tuning spaces.
Super classes
mlr3::Learner -> mlr3::LearnerClassif -> LearnerClassifRandomForest
Methods
Inherited methods
mlr3::Learner$base_learner()mlr3::Learner$configure()mlr3::Learner$encapsulate()mlr3::Learner$format()mlr3::Learner$help()mlr3::Learner$predict()mlr3::Learner$predict_newdata()mlr3::Learner$print()mlr3::Learner$reset()mlr3::Learner$selected_features()mlr3::Learner$train()mlr3::LearnerClassif$predict_newdata_fast()
LearnerClassifRandomForest$importance()
The importance scores are extracted from the slot importance.
Parameter 'importance' must be set to either "accuracy" or "gini".
Returns
Named numeric().
Examples
# Define the Learner
learner = lrn("classif.randomForest", importance = "accuracy")
print(learner)
#>
#> ── <LearnerClassifRandomForest> (classif.randomForest): Random Forest ──────────
#> • Model: -
#> • Parameters: importance=accuracy
#> • Packages: mlr3, mlr3extralearners, and randomForest
#> • Predict Types: [response] and prob
#> • Feature Types: logical, integer, numeric, factor, and ordered
#> • Encapsulation: none (fallback: -)
#> • Properties: importance, multiclass, oob_error, twoclass, and weights
#> • Other settings: use_weights = 'use', predict_raw = 'FALSE'
# Define a Task
task = tsk("sonar")
# Create train and test set
ids = partition(task)
# Train the learner on the training ids
learner$train(task, row_ids = ids$train)
print(learner$model)
#>
#> Call:
#> randomForest(formula = formula, data = data, classwt = classwt, cutoff = cutoff, importance = TRUE)
#> Type of random forest: classification
#> Number of trees: 500
#> No. of variables tried at each split: 7
#>
#> OOB estimate of error rate: 18.71%
#> Confusion matrix:
#> M R class.error
#> M 69 10 0.1265823
#> R 16 44 0.2666667
print(learner$importance())
#> V11 V12 V10 V9 V45
#> 1.922810e-02 1.809160e-02 1.767879e-02 1.629239e-02 1.330181e-02
#> V48 V36 V46 V47 V49
#> 1.263910e-02 1.206093e-02 1.059678e-02 1.029925e-02 7.283981e-03
#> V17 V8 V37 V18 V16
#> 5.855189e-03 5.752740e-03 5.709647e-03 5.685516e-03 5.657028e-03
#> V44 V52 V4 V27 V13
#> 5.452901e-03 4.809853e-03 4.010577e-03 3.513045e-03 3.509782e-03
#> V23 V51 V5 V28 V15
#> 3.493424e-03 3.130639e-03 2.595155e-03 2.492357e-03 2.471309e-03
#> V39 V6 V21 V35 V58
#> 2.436493e-03 2.259792e-03 2.254900e-03 1.937812e-03 1.782713e-03
#> V14 V34 V25 V20 V19
#> 1.555839e-03 1.506198e-03 1.470779e-03 1.318519e-03 1.291368e-03
#> V2 V50 V7 V43 V24
#> 1.235066e-03 1.216621e-03 1.205468e-03 1.102559e-03 1.094419e-03
#> V29 V22 V26 V32 V42
#> 1.079877e-03 1.031179e-03 8.579462e-04 7.579919e-04 7.558400e-04
#> V53 V31 V40 V55 V54
#> 5.925249e-04 5.280243e-04 2.958856e-04 2.451629e-04 1.511880e-04
#> V57 V30 V56 V38 V59
#> 4.540829e-05 -2.068076e-05 -7.776543e-05 -1.213900e-04 -1.371777e-04
#> V1 V33 V3 V60 V41
#> -2.478404e-04 -2.823621e-04 -3.055303e-04 -4.849227e-04 -7.452704e-04
# Make predictions for the test rows
predictions = learner$predict(task, row_ids = ids$test)
# Score the predictions
predictions$score()
#> classif.ce
#> 0.2318841