Gradient Boosting Classification Learner
Source:R/learner_gbm_classif_gbm.R
mlr_learners_classif.gbm.RdGradient Boosting Classification Algorithm.
Calls gbm::gbm() from gbm.
Meta Information
Task type: “classif”
Predict Types: “response”, “prob”
Feature Types: “integer”, “numeric”, “factor”, “ordered”
Required Packages: mlr3, mlr3extralearners, gbm
Parameters
| Id | Type | Default | Levels | Range |
| distribution | character | bernoulli | bernoulli, adaboost, huberized, multinomial | - |
| n.trees | integer | 100 | \([1, \infty)\) | |
| interaction.depth | integer | 1 | \([1, \infty)\) | |
| n.minobsinnode | integer | 10 | \([1, \infty)\) | |
| shrinkage | numeric | 0.001 | \([0, \infty)\) | |
| bag.fraction | numeric | 0.5 | \([0, 1]\) | |
| train.fraction | numeric | 1 | \([0, 1]\) | |
| cv.folds | integer | 0 | \((-\infty, \infty)\) | |
| keep.data | logical | FALSE | TRUE, FALSE | - |
| verbose | logical | FALSE | TRUE, FALSE | - |
| n.cores | integer | 1 | \((-\infty, \infty)\) | |
| var.monotone | untyped | - | - |
Initial parameter values
keep.datais initialized toFALSEto save memory.n.coresis initialized to 1 to avoid conflicts with parallelization through future.
References
Friedman, H J (2002). “Stochastic gradient boosting.” Computational statistics & data analysis, 38(4), 367–378.
See also
as.data.table(mlr_learners)for a table of available Learners in the running session (depending on the loaded packages).Chapter in the mlr3book: https://mlr3book.mlr-org.com/chapters/chapter2/data_and_basic_modeling.html#sec-learners
mlr3learners for a selection of recommended learners.
mlr3cluster for unsupervised clustering learners.
mlr3pipelines to combine learners with pre- and postprocessing steps.
mlr3tuning for tuning of hyperparameters, mlr3tuningspaces for established default tuning spaces.
Super classes
mlr3::Learner -> mlr3::LearnerClassif -> LearnerClassifGBM
Methods
Inherited methods
mlr3::Learner$base_learner()mlr3::Learner$configure()mlr3::Learner$encapsulate()mlr3::Learner$format()mlr3::Learner$help()mlr3::Learner$predict()mlr3::Learner$predict_newdata()mlr3::Learner$print()mlr3::Learner$reset()mlr3::Learner$selected_features()mlr3::Learner$train()mlr3::LearnerClassif$predict_newdata_fast()
LearnerClassifGBM$importance()
The importance scores are extracted by gbm::relative.influence() from
the model.
Returns
Named numeric().
Examples
# Define the Learner
learner = lrn("classif.gbm")
print(learner)
#>
#> ── <LearnerClassifGBM> (classif.gbm): Gradient Boosting ────────────────────────
#> • Model: -
#> • Parameters: keep.data=FALSE, n.cores=1
#> • Packages: mlr3, mlr3extralearners, and gbm
#> • Predict Types: [response] and prob
#> • Feature Types: integer, numeric, factor, and ordered
#> • Encapsulation: none (fallback: -)
#> • Properties: importance, missings, twoclass, and weights
#> • Other settings: use_weights = 'use', predict_raw = 'FALSE'
# Define a Task
task = tsk("sonar")
# Create train and test set
ids = partition(task)
# Train the learner on the training ids
learner$train(task, row_ids = ids$train)
#> Distribution not specified, assuming bernoulli ...
print(learner$model)
#> gbm::gbm(formula = f, data = data, keep.data = FALSE, n.cores = 1L)
#> A gradient boosted model with bernoulli loss function.
#> 100 iterations were performed.
#> There were 60 predictors of which 40 had non-zero influence.
print(learner$importance())
#> V12 V51 V49 V36 V4 V43 V40
#> 25.0081912 12.2142680 8.4157972 7.2136254 7.0811375 7.0219092 4.4034924
#> V28 V21 V9 V10 V45 V52 V27
#> 4.2378749 4.1898768 4.0774882 3.9714207 3.8251951 3.5570828 3.2370707
#> V37 V54 V44 V20 V23 V34 V1
#> 3.2122744 2.9767788 2.3651913 2.1860596 2.0014896 1.9473183 1.6662144
#> V3 V55 V17 V56 V5 V19 V48
#> 1.6254989 1.5872855 1.4252990 1.3404659 1.1851487 1.1542096 1.0278495
#> V24 V59 V60 V7 V33 V16 V39
#> 0.8757307 0.8167201 0.7874433 0.7105113 0.6386990 0.6034155 0.5463646
#> V31 V15 V29 V41 V11 V13 V14
#> 0.4908469 0.4063133 0.4016186 0.3765382 0.2936844 0.0000000 0.0000000
#> V18 V2 V22 V25 V26 V30 V32
#> 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000
#> V35 V38 V42 V46 V47 V50 V53
#> 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000 0.0000000
#> V57 V58 V6 V8
#> 0.0000000 0.0000000 0.0000000 0.0000000
# Make predictions for the test rows
predictions = learner$predict(task, row_ids = ids$test)
# Score the predictions
predictions$score()
#> classif.ce
#> 0.173913