--- title: "SSLfmm: A Practical Workflow" author: "Geoffrey J. McLachlan and Jinran Wu" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{SSLfmm: A Practical Workflow} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") ``` # Overview `SSLfmm` fits semi-supervised Gaussian finite-mixture classifiers when class labels are observed for only part of a sample. The package provides a common interface for complete-case (`cc`), missing-completely-at-random (`mcar`), entropy-dependent missing-at-random (`mar`), and mixed MCAR/MAR analyses. This vignette illustrates the main simulation, fitting, prediction, and assessment workflow using only the public package API. # Simulate partially labelled data ```{r simulation} library(SSLfmm) mu <- matrix(c(-1.5, 1.5), nrow = 1, ncol = 2) sim <- simulate_mixed_missingness( n = 120, pi = c(0.5, 0.5), mu = mu, sigma = matrix(1, 1, 1), alpha = 0.10, mar_rate = 0.25, seed = 2026 ) head(sim$data) table(sim$data$missing_source, useNA = "ifany") ``` The simulated data contain the feature columns (`x1`, ..., `xp`), the observed label (`label`), the complete reference class (`truth`, available because this is a simulation), and information about the missing-label mechanism. # Fit a mixed missingness model The unknown-source mixed model estimates the contribution of MCAR and MAR without requiring the source of each missing label to be supplied. ```{r fit} x <- as.matrix(sim$data["x1"]) fit <- fit_sslfmm( x, sim$data$label, g = 2, method = "mixed", covariance_type = "equal", indicator = "unknown", n_starts = 3, seed = 2027 ) fit summary(fit) ``` For applications in which the source of missingness is recorded, use `indicator = "known"` together with `missing_source`. # Prediction and uncertainty ```{r prediction} pred_class <- predict(fit, x) posterior <- predict(fit, x, type = "posterior") entropy <- predict(fit, x, type = "entropy") head(pred_class) head(posterior) head(entropy) ``` Posterior probabilities quantify class uncertainty. Entropy provides a compact summary of uncertainty and is also the quantity used by the entropy-dependent MAR mechanism implemented in the package. # Classification assessment Because the simulation retains the complete class labels, prediction can be assessed directly. ```{r assessment} perf <- classification_performance( sim$data$truth, pred_class, posterior ) perf$metrics perf$confusion_matrix ``` In genuinely partially labelled data, performance on observations whose labels are missing cannot usually be evaluated without an external reference set. # Included Blood Transfusion example data The package also includes the semi-synthetic `blood_transfusion` data set used in the software-paper application. ```{r data} data("blood_transfusion") head(blood_transfusion) table(blood_transfusion$missing_indicator) ``` The `truth` column is retained for evaluation of the semi-synthetic example, whereas `observed` is the partially observed response supplied to model-fitting functions. # Reproducibility and development The development repository is and issues can be reported at . The package contains automated `testthat` tests for its public API, simulation return contracts, fitting, prediction, input validation, and the included case-study data.