| Title: | Visualization for Classification Decision Boundaries |
| Version: | 1.0.0 |
| Description: | Exploring, visualizing, and comparing classification decision boundaries. Provides a unified interface for fitting classifiers and rendering 2D decision boundary plots, with support for 2D slice visualization (fixing non-plotted dimensions at reference values) and projection-based visualization for high-dimensional data (including Principal Component Analysis (PCA) and tour projections from the 'tourr' package). Supports native 'R' classifiers, 'tidymodels' workflows, and custom user-supplied models via a flexible adapter system. Includes an interactive 'Shiny' application ('explorapp') for visual exploration, data simulation, drawing, model comparison, probability surfaces, and reproducible exports. |
| License: | GPL-2 | GPL-3 [expanded from: GPL (≥ 2)] |
| URL: | https://github.com/natydasilva/classbound, https://natydasilva.github.io/classbound/ |
| LazyData: | yes |
| Depends: | R (≥ 4.3.0) |
| Imports: | class, ggplot2, rlang, ggnewscale, shiny, DT, RColorBrewer |
| Suggests: | tidymodels, workflowsets, workflows, parsnip, tourr, rpart, roxygen2 (≥ 7.3.0), PPtreeViz, PPtreeExt, ppforest2, randomForest, testthat (≥ 3.0.0), knitr, rmarkdown, palmerpenguins, e1071, MixSim, MASS, jsonlite |
| Encoding: | UTF-8 |
| Config/testthat/edition: | 3 |
| VignetteBuilder: | knitr |
| Config/roxygen2/version: | 8.1.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-09-20 18:35:46 UTC; Manihar |
| Author: | Vaibhav Manihar [aut, cre],
Natalia da Silva |
| Maintainer: | Vaibhav Manihar <vaibhav.manihar@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-09-30 09:10:15 UTC |
classbound: Visualization for Classification Decision Boundaries
Description
The classbound package provides tools for exploring, visualizing, and comparing
classification decision boundaries in R. It supports both two-dimensional data and
high-dimensional data (via 2D slicing or linear projections), and works with native
R classifiers, tidymodels workflows, and user-supplied models.
Details
Two main workflows
Interactive workflow: Launch explorapp() to start the built-in Shiny application.
From there, you can import your own data, simulate datasets, draw data by hand, choose
classifiers, adjust parameters, compare decision boundaries side-by-side, inspect
probability surfaces, inject outliers, and export the results.
Programmatic workflow: Use the modular API directly:
model <- fit_model(data, formula, classifier) model <- boundary_compute(model, feature_range, resolution = 100) plot_boundary(model, obs_data = data, x_col = "x", y_col = "y", true_label = "class")
For a one-step wrapper, use classbound().
High-dimensional data
When a model is trained on more than two features, boundary_compute() supports
two visualization strategies:
-
2D Slice: two features are selected for the axes; all other numeric features are fixed at their median and categorical features at their mode.
-
Projection: a projection matrix maps the high-dimensional feature space to two dimensions (e.g., PCA or a tour basis from the
tourrpackage). The boundary grid is generated in projection space and inverse-projected back for prediction.
Supported classifiers
Any classifier whose predict() method returns a vector or factor of class labels
works automatically. Built-in adapters are provided for rpart, randomForest,
PPtreeViz, PPtreeExt, and ppforest2. For classifiers that return complex
objects (such as lists), use the predfun argument to extract class labels.
Native tidymodels integration is available via boundary_workflow_set().
Author(s)
Maintainer: Vaibhav Manihar vaibhav.manihar@gmail.com
Authors:
Vaibhav Manihar vaibhav.manihar@gmail.com
Natalia da Silva natalia.dasilva@fcea.edu.uy (ORCID)
Ignacio Alvarez-Castro ignacio.lavarez@fcea.edu.uy (ORCID)
See Also
-
classbound()for the all-in-one wrapper -
fit_model()to fit a model -
boundary_compute()to compute a decision boundary -
plot_boundary()to visualize a decision boundary -
explorapp()for the interactive Shiny application -
boundary_workflow_set()for tidymodels multi-model comparison
Convert a fitted model into a classbound object
Description
Wraps a pre-fitted model in a classbound object, adding the feature metadata
(names, types, ranges, imputation values) required by boundary_compute() and
plot_boundary(). This is the entry point for "bring your own model" (BYO) workflows,
including native tidymodels workflows and parsnip model fits.
Usage
as_classbound(model, data, response = NULL, ...)
Arguments
model |
A fitted model object. For |
data |
A data frame of the training data. Used only to extract feature metadata;
the data is not passed through |
response |
Optional string naming the response column in |
... |
Additional arguments passed to methods. |
Details
Unlike fit_model(), as_classbound() does not refit the model or call
preprocess_data(). Feature metadata is extracted solely from the data argument,
which should be the same data used to fit the model. If the model was fitted on
scaled or transformed data, ensure data reflects that transformation.
as_classbound() dispatches on the class of model:
-
Default: wraps any fitted model object (e.g.,
rpart,lda). -
workflow: requires the workflow to already be trained (viafit()). Useboundary_workflow_set()to train and wrap an entire workflow set at once. -
model_fit: wraps a fittedparsnipmodel.
Value
A classbound object ready for use with boundary_compute() and
plot_boundary().
See Also
fit_model(), boundary_workflow_set(), boundary_compute()
Examples
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
# Wrap any pre-fitted model
raw_model <- rpart::rpart(species ~ ., data = peng_data)
cb_model <- as_classbound(raw_model, data = peng_data, response = "species")
# Tidymodels: wrap a fitted workflow
# (boundary_workflow_set() does this automatically for entire workflow sets)
Compute the classification decision boundary
Description
Generates a 2D grid of predicted class labels (and optionally class probabilities)
for a fitted classbound model. The resulting boundary data is stored in the returned
object and consumed directly by plot_boundary().
Usage
boundary_compute(
model,
feature_range = NULL,
resolution = 100,
predfun = NULL,
projection = NULL,
reference = NULL,
...
)
Arguments
model |
A |
feature_range |
A named list of length 2 specifying the axis ranges, e.g.,
|
resolution |
An integer >= 2 specifying the number of grid points per axis. Higher values produce smoother boundaries. Default is 100; use 50 for faster interactive exploration. |
predfun |
An optional custom prediction function for non-standard classifiers.
Must accept |
projection |
An optional named list defining a high-dimensional projection.
Must contain |
reference |
An optional named list of fixed reference values for features not
specified in |
... |
Additional arguments passed to |
Details
2D slice (two-feature visualization)
When projection is NULL, boundary_compute() generates a regular grid over the
two features named in feature_range. If the model was trained on more than two
features, all remaining numeric features are fixed at their training-set median and
categorical features at their training-set mode, unless you supply explicit values
via reference.
This is a 2D slice of the full multivariate decision boundary. Two observations that appear in the same region may still be separated in a dimension that is held fixed. Use this mode when you want to examine how two specific features interact, or when your model was trained on exactly two features.
Projection (high-dimensional visualization)
When projection is provided, the 2D grid is generated in the projected space and
then inverse-projected back to the original feature space before prediction.
This means the model uses all of its training features; only the visualization is
collapsed to two dimensions.
The projection $basis must be a numeric matrix with:
row count equal to the number of training features, row names matching feature names
exactly two columns (the projected axes)
orthonormal columns (
crossprod(basis)must equaldiag(2))
Suitable bases can be obtained from prcomp() or the tourr package.
Multi-model comparison
Pass a named list of classbound objects (all trained on the same features and class
levels) to compute boundaries for multiple models at once. The result contains a
model column, suitable for use with plot_boundary(facet_col = "model").
Value
A modified classbound object with the additional class "classbound_boundary".
The boundary grid is stored in $boundary_data (columns: x, y, prediction,
and per-class probability columns when available). For multi-model input, also
contains a model column.
See Also
fit_model(), plot_boundary(), boundary_workflow_set()
Examples
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
# Fit and compute 2D boundary with auto-detected ranges
m <- fit_model(peng_data, species ~ ., rpart::rpart)
m <- boundary_compute(m, resolution = 50)
head(m$boundary_data)
# Explicit axis ranges
m <- boundary_compute(m,
feature_range = list(bill_length_mm = c(30, 60), bill_depth_mm = c(10, 25)),
resolution = 100
)
# 3-feature model visualized as a 2D slice
peng3d <- na.omit(penguins[, c(
"species", "bill_length_mm",
"bill_depth_mm", "flipper_length_mm"
)])
m3d <- fit_model(peng3d, species ~ ., rpart::rpart)
m3d_slice <- boundary_compute(m3d,
feature_range = list(bill_length_mm = c(30, 60), bill_depth_mm = c(10, 25)),
resolution = 50
)
# 3-feature model visualized via PCA projection
feat_cols <- c("bill_length_mm", "bill_depth_mm", "flipper_length_mm")
pca <- prcomp(peng3d[, feat_cols], scale. = TRUE)
basis <- pca$rotation[, 1:2]
m3d_proj <- boundary_compute(m3d,
feature_range = list(PC1 = c(-4, 4), PC2 = c(-3, 3)),
resolution = 50,
projection = list(basis = basis, center = pca$center, scale = pca$scale)
)
Explore high-dimensional classification boundaries
Description
Generates a high-dimensional uniform domain and identifies points near the decision boundaries of multiple models.
Usage
boundary_explore_nd(models, data, n_points = 10000, threshold = 0.05)
Arguments
models |
A named list of fitted |
data |
A data frame representing the original training data. The bounding box of this data determines the simulation domain. |
n_points |
Integer. Number of points to simulate in the high-dimensional space (default: 10000). |
threshold |
Numeric. A percentile threshold (0 to 1) for determining boundary proximity. Points below this advantage percentile are considered boundary-near. Default: 0.05. |
Value
A data frame containing the simulated coordinates, the predicted classes for each model (prefixed with Y_), and a logical is_boundary column indicating if the point is near the boundary for ANY of the provided models.
Compute classification boundaries for a workflow_set
Description
A dedicated helper to automatically fit and extract classification boundaries
for an entire workflow_set. This avoids the need for manual iteration over models.
Usage
boundary_workflow_set(
wf_set,
data,
feature_range = NULL,
response,
resolution = 100,
...
)
Arguments
wf_set |
A |
data |
A data frame containing the training data. This is required to extract feature metadata and to fit any workflows that are not yet trained. |
feature_range |
An optional named list specifying the minimum and maximum values for each feature,
or a character vector of feature names. If |
response |
A string specifying the name of the response column in |
resolution |
An integer specifying the number of points along each axis (default = 100). |
... |
Additional arguments passed to |
Value
A data frame containing the combined boundary grid for all models, with a model column
indicating the wflow_id.
Examples
library(palmerpenguins)
library(workflowsets)
library(parsnip)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
# Define multiple engines
spec_rpart <- decision_tree() |>
set_engine("rpart") |>
set_mode("classification")
spec_glm <- multinom_reg() |>
set_engine("nnet") |>
set_mode("classification")
# Create a workflow set
wf_set <- workflow_set(
preproc = list(base = species ~ bill_length_mm + bill_depth_mm),
models = list(tree = spec_rpart, log_reg = spec_glm)
)
# Compute 2D boundaries for all models simultaneously (auto-range)
bounds <- boundary_workflow_set(wf_set, peng_data, response = "species", resolution = 30)
Visualize a classification decision boundary
Description
A high-level unified wrapper to fit a model, compute its 2D decision boundary, and plot the results in a single step.
Usage
classbound(
data,
formula,
classifier,
interface = c("formula", "matrix", "custom"),
projection = NULL,
fit_args = list(),
predict_args = list(),
predfun = NULL,
resolution = 100,
...
)
Arguments
data |
A data frame containing the full training dataset. The specific variables
used for modeling and plotting are strictly determined by the |
formula |
A formula specifying the response and predictors. |
classifier |
The classification function to use (e.g., |
interface |
A string specifying how to invoke the classifier: |
projection |
An optional list (e.g., |
fit_args |
A named list of additional arguments passed to the classifier during fitting. |
predict_args |
A named list of additional arguments passed to |
predfun |
A custom function to generate predictions for non-standard models. |
resolution |
An integer specifying the grid resolution for the decision boundary. |
... |
Additional arguments passed to |
Value
A ggplot object visualizing the 2D decision boundary and original observations.
Examples
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
# Quick 2D boundary visualization for an SVM
classbound(peng_data, species ~ bill_length_mm + bill_depth_mm, e1071::svm)
Generate deterministic colors for classification classes
Description
Returns a named character vector mapping class labels to hex color codes. Colors are assigned deterministically: class labels are sorted alphabetically before assignment, so the same class always receives the same color regardless of the order the classes are supplied.
Usage
classbound_palette(classes)
Arguments
classes |
A character vector or factor containing class labels. Duplicates
and |
Details
The first 20 classes are assigned colors from a curated palette of visually
distinct hues designed for use on both light and dark backgrounds. For more than
20 classes, additional colors are generated using the golden-angle method
(hue = (index * phi) %% 1 where phi = 0.618...), which produces perceptually
distinct colors with uniform spacing in HCL space.
plot_boundary() uses this palette by default. If an RColorBrewer palette is
requested via palette = but cannot accommodate the number of classes (e.g.,
"Dark2" supports a maximum of 8), it automatically falls back to
classbound_palette().
Value
A named character vector where names are the unique sorted class labels
and values are hex color codes (e.g., "#E6194B").
See Also
Examples
# 3-class palette
classbound_palette(c("Adelie", "Chinstrap", "Gentoo"))
# Colors are alphabetically ordered, so order of input doesn't matter
identical(
classbound_palette(c("B", "A", "C")),
classbound_palette(c("C", "A", "B"))
)
# Use in plot_boundary() via the palette argument
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
m <- fit_model(peng_data, species ~ ., rpart::rpart)
m <- boundary_compute(m, resolution = 50)
plot_boundary(m,
obs_data = peng_data, x_col = "bill_length_mm",
y_col = "bill_depth_mm", true_label = "species"
)
Render a decision boundary plot for the Shiny app
Description
Render a decision boundary plot for the Shiny app
Usage
create_boundary_plot(
cb_mod,
data,
title,
class_levels = NULL,
class_colors = NULL,
proj_matrix = NULL,
proj_info = NULL,
zoom_x = NULL,
zoom_y = NULL,
show_probs = FALSE,
resolution = 100,
predict_args = list(),
n_outliers = 0,
highlight_outliers = TRUE,
slice_x = NULL,
slice_y = NULL
)
Arguments
cb_mod |
A classbound_model object. |
data |
The training data to overlay. |
title |
The title of the plot. |
class_levels |
The factor levels for the classes. |
proj_matrix |
Optional projection matrix. |
proj_info |
Optional list with center and scale for projection. |
zoom_x |
Optional x-axis limits. |
zoom_y |
Optional y-axis limits. |
show_probs |
Whether to show probability gradients. |
resolution |
The grid resolution. |
slice_x |
Optional. Feature name to use for the X axis in 2D Slice mode. |
slice_y |
Optional. Feature name to use for the Y axis in 2D Slice mode. |
Value
A ggplot object.
Waveform Dataset (UCI)
Description
A subset of the UCI Waveform Database Generator dataset (Version 1). Contains 5000 observations with 21 numeric features and a 3-class response variable. All 21 features contribute to class membership, making this a useful dataset for demonstrating high-dimensional boundary visualization, including 2D slice and projection-based approaches.
Usage
data(data69_1)
Format
A data frame with 5000 rows and 22 columns:
- Y
Class label: a factor with 3 levels (
"1","2","3")- V1
Numeric feature variable
- V2
Numeric feature variable
- V3
Numeric feature variable
- V4
Numeric feature variable
- V5
Numeric feature variable
- V6
Numeric feature variable
- V7
Numeric feature variable
- V8
Numeric feature variable
- V9
Numeric feature variable
- V10
Numeric feature variable
- V11
Numeric feature variable
- V12
Numeric feature variable
- V13
Numeric feature variable
- V14
Numeric feature variable
- V15
Numeric feature variable
- V16
Numeric feature variable
- V17
Numeric feature variable
- V18
Numeric feature variable
- V19
Numeric feature variable
- V20
Numeric feature variable
- V21
Numeric feature variable
Details
Each observation is generated from one of three waveform classes. All features
are numeric; the response variable Y is a factor with levels "1", "2", and
"3". Load the dataset with data(data69_1).
This dataset is included primarily to demonstrate high-dimensional boundary
visualization with tools like PCA projection and tourr. For introductory
2D examples, the palmerpenguins::penguins dataset provides a more
accessible alternative.
Source
https://archive.ics.uci.edu/dataset/107/waveform+database+generator+version+1
Examples
data(data69_1)
dim(data69_1) # 5000 x 22
levels(data69_1$Y) # "1" "2" "3"
head(data69_1[, 1:5])
Shiny app to visually explore and compare classification boundaries for built-in, custom, and tidymodels algorithms in 2D
Description
Shiny app to visually explore and compare classification boundaries for built-in, custom, and tidymodels algorithms in 2D
Usage
explorapp(data = NULL, target_col = NULL, custom_models = list())
Arguments
data |
Optional data frame to import directly into the app. |
target_col |
Optional string specifying the column in |
custom_models |
A list of custom models to inject into the app's comparison UI. Each element should be a named list containing at least |
Value
No return value, called for side effects. Shinyapp is launched.
Examples
if (interactive()) {
# Launch with default models
explorapp()
# Launch with a custom SVM model
explorapp(custom_models = list(
"SVM" = list(
fn = e1071::svm,
args = list(kernel = "linear")
)
))
}
Export current app configuration to JSON
Description
Captures the current state of all Shiny UI inputs as a JSON file for reproducibility. Includes data mode, simulation parameters, grid resolution, and selected classifiers.
Usage
export_config_json(input_list, file)
Arguments
input_list |
A named list of input values. |
file |
The file path to write the JSON. |
Value
The file path (invisibly).
Export data to CSV
Description
Export data to CSV
Usage
export_data_csv(data, file)
Arguments
data |
A data frame to export. |
file |
The file path to write to. |
Value
The file path (invisibly).
Export grid predictions to CSV
Description
Saves the boundary grid data (X, Y, predicted class, probabilities) for a fitted classbound model so results can be reproduced outside the Shiny app.
Usage
export_grid_csv(
cb_mod,
data,
resolution,
proj_matrix = NULL,
proj_info = NULL,
predict_args = list(),
slice_x = NULL,
slice_y = NULL,
file
)
Arguments
cb_mod |
A classbound model object. |
data |
The training data (used to compute axis ranges). |
resolution |
Grid resolution (points per axis). |
proj_matrix |
Optional projection matrix for high-dimensional data. |
proj_info |
Optional list with center/scale for projection. |
predict_args |
Additional arguments passed to predict. |
file |
The file path to write the CSV. |
Value
The file path (invisibly).
Export comparison metrics to CSV
Description
Export comparison metrics to CSV
Usage
export_metrics_csv(metrics, file)
Arguments
metrics |
A data frame of metrics. |
file |
The file path to write to. |
Value
The file path (invisibly).
Export fitted models to RDS
Description
Export fitted models to RDS
Usage
export_models_rds(models, file)
Arguments
models |
A list of fitted model objects. |
file |
The file path to write to. |
Value
The file path (invisibly).
Export ggplot objects to a multi-page PDF
Description
Export ggplot objects to a multi-page PDF
Usage
export_plots_pdf(plots, file, width = 8, height = 8)
Arguments
plots |
A list of ggplot objects. |
file |
The file path to write to. |
width |
Width of the PDF in inches. |
height |
Height of the PDF in inches. |
Value
The file path (invisibly).
Export ggplot objects to individual PNG files in a directory
Description
Export ggplot objects to individual PNG files in a directory
Usage
export_plots_png(plots, dir, dpi = 300, width = 8, height = 8)
Arguments
plots |
A list of ggplot objects (named or unnamed). |
dir |
The directory path to write the PNGs to. |
dpi |
Resolution of the PNG files. |
width |
Width of each PNG in inches. |
height |
Height of each PNG in inches. |
Value
The directory path (invisibly).
Export a reproducibility R script
Description
Writes a self-contained R script that loads the exported data and models and regenerates the boundary plots. Handles both 2D and high-dimensional (projection-based) datasets.
Usage
export_reproduce_script(
model_names,
has_projection = FALSE,
slice_x = NULL,
slice_y = NULL,
resolution = 100,
show_probs = FALSE,
zoom_x = NULL,
zoom_y = NULL,
file
)
Arguments
model_names |
Character vector of model names included in the export. |
has_projection |
Logical, whether a projection.rds file was exported. |
file |
The file path to write the script. |
Value
The file path (invisibly).
Fit a machine learning model
Description
Fits a classification model and wraps it in a classbound object, which carries
the feature metadata needed by boundary_compute() and plot_boundary().
Usage
fit_model(data, formula, classifier, ...)
## Default S3 method:
fit_model(
data,
formula,
classifier,
interface = c("formula", "matrix", "custom"),
fit_args = list(),
...
)
## S3 method for class ''function''
fit_model(
data,
formula,
classifier,
interface = c("formula", "matrix", "custom"),
fit_args = list(),
...
)
## S3 method for class 'character'
fit_model(
data,
formula,
classifier,
interface = c("formula", "matrix", "custom"),
fit_args = list(),
...
)
Arguments
data |
A data frame containing the training features and response variable.
All columns referenced in |
formula |
A formula specifying the response and predictors, e.g.,
|
classifier |
The classification function or model specification to use.
Pass a function (e.g., |
... |
Additional arguments passed to methods. |
interface |
A string specifying how to invoke the classifier: |
fit_args |
A named list of additional arguments forwarded to the classifier
during fitting (e.g., |
Details
Interface modes
fit_model() supports three calling conventions via the interface argument:
-
"formula"(default): passesformulaanddatadirectly to the classifier. Works for the vast majority of R classifiers (e.g.,rpart::rpart,e1071::svm,stats::lda). -
"matrix": constructs a predictor matrixxand response vectoryfrom the formula, then callsclassifier(x, y, ...). Required for classifiers whose primary interface is matrix-based (e.g.,randomForest::randomForest). -
"custom": passes onlyfit_argsto the classifier, giving you full control over the call. Use this for non-standard APIs.
Tidymodels support
If classifier is a parsnip model specification (model_spec), a fitted
model_fit, or a tidymodels workflow, fit_model() dispatches to the
appropriate method automatically. The interface argument is not needed for these
objects.
Preprocessing
fit_model() calls preprocess_data() internally to coerce response labels to
a factor, handle missing values, and extract feature metadata. Do not call
preprocess_data() manually before calling fit_model(); this will corrupt the
stored metadata.
Value
A classbound object (a list of class "classbound") containing:
-
$fit: the raw fitted model object returned by the classifier -
$metadata: a list with$features(names, types, ranges, imputation values) and$class_levels(sorted character vector of class labels) -
$boundary_data:NULLuntilboundary_compute()is called
See Also
boundary_compute(), predict_model(), classbound()
Examples
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
# Formula interface (most classifiers)
m_rpart <- fit_model(peng_data, species ~ ., rpart::rpart)
# Matrix interface (randomForest)
m_rf <- fit_model(peng_data, species ~ ., randomForest::randomForest,
interface = "matrix"
)
# Additional fitting arguments via fit_args
m_rpart_cp <- fit_model(peng_data, species ~ ., rpart::rpart,
fit_args = list(control = rpart::rpart.control(cp = 0.001))
)
Generate a deterministic extreme outlier based on dataset bounds
Description
Generate a deterministic extreme outlier based on dataset bounds
Usage
generate_outlier(data, class_label, magnitude, target_col = "Sim", index = 1)
Arguments
data |
A data frame containing the features and target column. |
class_label |
The class label to assign to the outlier. |
magnitude |
A numeric value indicating how far outside the bounding box to place the outlier. |
target_col |
The name of the target column in the data. |
index |
An integer index used to cycle through bounding box corners. |
Value
A one-row data frame with the generated outlier.
Plot a classbound model boundary
Description
A standard S3 plot method that delegates to plot_boundary().
Usage
## S3 method for class 'classbound'
plot(x, ...)
Arguments
x |
A |
... |
Additional arguments passed to |
Value
A ggplot2 object.
Plot a computed boundary
Description
An S3 plot method that delegates to plot_boundary() for
objects that have had boundaries computed via boundary_compute().
Usage
## S3 method for class 'classbound_boundary'
plot(x, ...)
Arguments
x |
A |
... |
Additional arguments passed to |
Value
A ggplot2 object.
Visualize the classification decision boundary
Description
Renders a 2D ggplot of a classification decision boundary, optionally overlaying
observed training data. Input must be a classbound object that has already been
passed through boundary_compute().
Usage
plot_boundary(
model,
obs_data = NULL,
x_col = NULL,
y_col = NULL,
true_label = NULL,
facet_col = NULL,
type = "2D",
show_gradient = FALSE,
agree_color = "#006666",
disagree_color = "#FF8000",
obs_alpha = 1,
obs_size = 2.5,
render = c("raster", "tile"),
colors = NULL,
palette = NULL,
highlight_outliers = FALSE,
xlim = NULL,
ylim = NULL,
...
)
Arguments
model |
A |
obs_data |
An optional data frame of observations to overlay on the boundary plot.
Typically the training data. If provided, |
x_col |
Column name in |
y_col |
Column name in |
true_label |
Column name in |
facet_col |
Optional string naming a column in the boundary data to facet the
plot by. Set to |
type |
The visualization type. |
show_gradient |
Logical. If |
agree_color |
Color for regions where all models agree (only for
|
disagree_color |
Color for regions where models disagree (only for
|
obs_alpha |
Numeric transparency for overlaid observation points (0.0–1.0). When a projection is active, this is treated as the maximum opacity; actual alpha varies by depth-fading. |
obs_size |
Numeric point size for overlaid observations. |
render |
Rendering method for decision regions: |
colors |
Optional named character vector mapping class labels to colors, e.g.
|
palette |
Optional RColorBrewer palette name (e.g., |
highlight_outliers |
Logical. If |
xlim |
Optional length-2 numeric vector to set x-axis limits. |
ylim |
Optional length-2 numeric vector to set y-axis limits. |
... |
Additional arguments (currently unused). |
Details
Probability surface (gradient)
When show_gradient = TRUE, decision regions are shaded by the predicted class
probability: deep, saturated regions indicate high model confidence, while faded
regions indicate uncertainty near the boundary. Probability shading is only possible
when the underlying classifier returns class probabilities. Classifiers that return
only class labels (e.g., standard SVMs or PPtree models) produce a flat boundary
regardless of show_gradient.
High-dimensional projections and depth fading
When the classbound object contains a projection (from boundary_compute(..., projection = ...)),
plot_boundary() automatically forward-projects any obs_data observations onto the
2D plane. It also computes each point's orthogonal distance from the projection plane
and maps this distance to opacity: points lying exactly on the plane are fully opaque
(alpha = 1.0), while points further away in the original feature space gradually
fade toward alpha = 0.2. This depth-fading provides visual cues about how faithfully
each point's position is captured by the current projection.
Rendering backend and plotly compatibility
The default render = "raster" uses ggplot2::geom_raster(), which is fast and
produces high-quality static output. However, geom_raster() is not supported by
plotly::ggplotly() and will produce a blank interactive plot. To convert a boundary
plot to an interactive plotly figure, use render = "tile" instead:
p <- plot_boundary(model, render = "tile") plotly::ggplotly(p)
Colors and palettes
By default, plot_boundary() uses classbound_palette(), a curated 20-color palette
with deterministic (alphabetical) class-to-color assignment, ensuring consistent colors
across multiple plots. Supply a palette name (e.g., "Dark2") to use an RColorBrewer
palette, or supply colors as a named vector for explicit control. If an RColorBrewer
palette cannot accommodate the number of classes, it falls back to classbound_palette().
Value
A ggplot2 object.
See Also
boundary_compute(), classbound(), classbound_palette()
Examples
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
m <- fit_model(peng_data, species ~ ., rpart::rpart)
m <- boundary_compute(m, resolution = 50)
# Basic boundary plot with observations
plot_boundary(m,
obs_data = peng_data,
x_col = "bill_length_mm",
y_col = "bill_depth_mm",
true_label = "species"
)
# Probability gradient (rpart supports probabilities)
plot_boundary(m,
obs_data = peng_data,
x_col = "bill_length_mm",
y_col = "bill_depth_mm",
true_label = "species",
show_gradient = TRUE
)
# Plotly-compatible rendering
p <- plot_boundary(m,
obs_data = peng_data,
x_col = "bill_length_mm",
y_col = "bill_depth_mm",
true_label = "species",
render = "tile"
)
# plotly::ggplotly(p) # uncomment to convert to interactive
Predict using a classbound model
Description
Generates predictions using a unified interface across all classifiers.
Dispatches natively on objects of class "classbound". Attempting to call
predict() directly on a "classbound_multi" object will result in an error,
as multi-model boundaries are evaluated internally by boundary_compute().
Usage
## S3 method for class 'classbound'
predict(object, newdata, predict_args = list(), predfun = NULL, ...)
## S3 method for class 'classbound_multi'
predict(object, newdata, predict_args = list(), predfun = NULL, ...)
Arguments
object |
A fitted classbound model. This is the object returned by |
newdata |
A data frame of new observations to predict on. |
predict_args |
A named list of additional arguments passed to |
predfun |
A custom function to generate predictions for non-standard models.
The function must accept at least two arguments: |
... |
Additional arguments passed to the specific model adapter. |
Value
A list containing class (a factor of predicted labels) and probs
(a probability matrix, or strictly NULL if the classifier lacks probability support).
Downstream functions like boundary_compute() are designed to handle probs = NULL gracefully.
Internal generic for predicting classifier adapters
Description
This generic is exported for S3 dispatch but is not intended for direct use.
Use predict.classbound or predict_model instead.
Usage
predict_adapter(model, newdata, ...)
Arguments
model |
A fitted native model object. |
newdata |
A data frame of new observations. |
... |
Additional arguments. |
Value
A character vector of predicted class labels, one per row of newdata.
Predict using a fitted PPtreeExt model
Description
Adapter function to generate standardized predictions from a PPtreeExt model.
Usage
## S3 method for class 'PPtreeExtclass'
predict_adapter(model, newdata, ...)
Arguments
model |
A fitted |
newdata |
A data frame of new observations to predict on. |
... |
Additional arguments passed to |
Value
A list containing class (predicted labels) and probs (probabilities, NULL for pptree).
Predict using a fitted PPtreeViz model
Description
Adapter function to generate standardized predictions from a PPtreeViz model.
Usage
## S3 method for class 'PPtreeclass'
predict_adapter(model, newdata, ...)
Arguments
model |
A fitted |
newdata |
A data frame of new observations to predict on. |
... |
Additional arguments passed to |
Value
A list containing class (predicted labels) and probs (probabilities, NULL for PPtreeViz).
Predict using a fitted ppforest2 model
Description
Adapter function to generate standardized predictions from a ppforest2 model.
Developer Note: During development, an issue was observed with
ppforest2 v0.1.3 when training data is contiguous by class but
ordered so that the lowest class does not occur first (e.g.,
Class 3, Class 3, Class 1, Class 1). In this case, the underlying
C++ code can produce Grouping::init: partition must be rooted at row 0.
classbound does not reorder or otherwise modify the training data
specifically to avoid this behavior. If a future version of ppforest2
changes this behavior, no corresponding change to classbound is
expected to be necessary.
Usage
## S3 method for class 'pprf_classification'
predict_adapter(model, newdata, ...)
Arguments
model |
A fitted |
newdata |
A data frame of new observations to predict on. |
... |
Additional arguments passed to |
Value
A list containing class (predicted labels) and probs (probabilities).
Predict using a fitted randomForest model
Description
Adapter function to generate standardized predictions from a randomForest model.
Usage
## S3 method for class 'randomForest'
predict_adapter(model, newdata, ...)
Arguments
model |
A fitted |
newdata |
A data frame of new observations to predict on. |
... |
Additional arguments passed to |
Value
A list containing class (predicted labels) and probs (probabilities).
Predict using a fitted rpart model
Description
Adapter function to generate standardized predictions from an rpart model.
Usage
## S3 method for class 'rpart'
predict_adapter(model, newdata, ...)
Arguments
model |
A fitted |
newdata |
A data frame of new observations to predict on. |
... |
Additional arguments passed to |
Value
A list containing class (predicted labels) and probs (probabilities).
Predict using a fitted classbound model
Description
Generates predictions using a unified interface across all classifiers.
This function is a compatibility wrapper around the standard predict() method for "classbound" objects.
Usage
predict_model(model, newdata, predict_args = list(), predfun = NULL, ...)
Arguments
model |
A fitted classbound model. This corresponds to the |
newdata |
A data frame of new observations to predict on. |
predict_args |
A named list of additional arguments passed to |
predfun |
A custom function to generate predictions for non-standard models.
The function must accept at least two arguments: |
... |
Additional arguments passed to the specific model adapter. |
Value
A list containing class (a factor of predicted labels) and probs
(a probability matrix, or strictly NULL if the classifier lacks probability support).
Examples
library(palmerpenguins)
data(penguins)
peng_data <- na.omit(penguins[, c("species", "bill_length_mm", "bill_depth_mm")])
m_rpart <- fit_model(peng_data, species ~ ., rpart::rpart)
preds <- predict_model(m_rpart, newdata = peng_data[1:5, ])
Preprocess data before model fitting
Description
Applies standard preprocessing steps to a training dataset: validates structure, coerces character columns to factors, drops unused factor levels, rejects missing and infinite values, and converts response labels to a factor.
Usage
preprocess_data(data, labels = NULL, ...)
Arguments
data |
A data frame of raw training data. |
labels |
Optional vector of target labels. Converted to a factor with unused levels dropped. |
... |
Additional arguments (currently unused). |
Details
This function is called automatically by fit_model(). Do not call it manually
before fit_model(), because doing so will result in feature metadata being extracted
from the pre-processed data rather than the original data, which corrupts the
imputation values used by boundary_compute() for 2D slicing.
It is exported for use in custom workflows and testing, but it is not typically
needed in the standard fit_model() → boundary_compute() pipeline.
Value
A list with two elements:
-
$data: the preprocessed data frame -
$labels: the processed factor labels (orNULLif not supplied)
Print a classbound model
Description
Prints a clean summary of a classbound model object, hiding the internal wrapper
structure. For single models, it delegates to the native model's print method. For multi-model
objects (classbound_multi), it prints a comparison summary of the included models.
Usage
## S3 method for class 'classbound'
print(x, ...)
## S3 method for class 'classbound_multi'
print(x, ...)
Arguments
x |
A |
... |
Additional arguments passed to the native model's print method. |
Value
The object, invisibly.
Simulate multivariate normal data for multiple classes
Description
Generates synthetic classification data by drawing independent multivariate normal samples for each class. Useful when you want explicit control over each class's mean and covariance structure.
Usage
simu_n(
means,
covs,
ns,
class_names = NULL,
seed = NULL,
noise_ratio = 0,
test_ratio = 0
)
Arguments
means |
A list of numeric vectors, one per class, specifying the class means. Each vector must have length equal to the number of features. |
covs |
A list of covariance matrices, one per class. |
ns |
A numeric vector of sample sizes, one per class. |
class_names |
Optional character vector of class labels (length equal to
|
seed |
Optional integer for reproducibility. The global random seed is restored after the call. |
noise_ratio |
Numeric in [0, 1). Proportion of |
test_ratio |
Numeric in [0, 1). If greater than 0, generates an additional independent test dataset. This is not a split of the training data. |
Details
Test data
When test_ratio > 0, an additional independent test dataset is generated by
drawing fresh samples of size round(ns * test_ratio) for each class using the
same means and covs. This is not a split of the training data. The training
set has sum(ns) observations; the test set has sum(round(ns * test_ratio))
independently generated observations.
Value
If test_ratio == 0 (default): a data frame with a Sim class column
and feature columns (X1, X2, ...). If test_ratio > 0: a list with
$train and $test data frames.
See Also
Examples
means <- list(c(0, 0), c(3, 3), c(0, 5))
covs <- list(diag(2), diag(2), diag(2))
ns <- c(60, 60, 60)
train_df <- simu_n(means, covs, ns, seed = 1)
head(train_df)
# With independent test data
sim <- simu_n(means, covs, ns, seed = 1, test_ratio = 0.3)
nrow(sim$train) # 180
nrow(sim$test) # 54 (independently generated, not split from train)
Simulate data using Gaussian mixture models
Description
Generates synthetic classification data using Gaussian mixture models constructed by
the MixSim package. Useful for creating controlled datasets to explore decision
boundary behavior when no real dataset is available.
Usage
simulate_mixsim(
n,
K,
p,
MaxOmega,
class_names = NULL,
seed = NULL,
noise_ratio = 0,
test_ratio = 0
)
Arguments
n |
Number of training observations to generate. |
K |
Number of classes. |
p |
Number of numeric features (dimensions). |
MaxOmega |
Maximum pairwise overlap between mixture components (0 to 1). Smaller values create better-separated classes. |
class_names |
Optional character vector of class labels (length |
seed |
Optional integer for reproducibility. The global random seed is restored after the call. |
noise_ratio |
Numeric in [0, 1). Proportion of |
test_ratio |
Numeric in [0, 1). If greater than 0, generates an additional
independent test dataset of size |
Details
Class distributions are randomly generated subject to the MaxOmega overlap
constraint. The simulation is fully reproducible when seed is supplied.
Test data
When test_ratio > 0, an additional independent dataset is generated using the
same mixture parameters but a fresh random draw. This is not a split of the
training data: the training set has n observations and the test set has
round(n * test_ratio) independently generated observations. The two sets are
statistically independent, so the test set is a fair evaluation sample.
Value
If test_ratio == 0 (default): a data frame with a Sim class column and
p feature columns (X1, X2, ...). If test_ratio > 0: a list with $train
and $test data frames.
See Also
Examples
# Generate a 3-class, 2-dimensional training dataset
train_df <- simulate_mixsim(n = 200, K = 3, p = 2, MaxOmega = 0.05, seed = 42)
head(train_df)
# Generate training and independent test data
sim <- simulate_mixsim(
n = 200, K = 3, p = 2, MaxOmega = 0.05,
seed = 42, test_ratio = 0.3
)
nrow(sim$train) # 200
nrow(sim$test) # 60 (independently generated, not split from train)
Summarize a classbound model
Description
For single models, delegates the summary function directly to the native model
wrapped inside the classbound object. For multi-model objects (classbound_multi), it
returns a summary of the contained models.
Usage
## S3 method for class 'classbound'
summary(object, ...)
## S3 method for class 'classbound_multi'
summary(object, ...)
Arguments
object |
A |
... |
Additional arguments passed to the native model's summary method. |
Value
The summary of the native model.