Package {clis}


Type: Package
Title: Conformal Local Influence Screening for Bounded-Response Regression
Version: 0.3.6
Description: Provides fast, statistically calibrated influence diagnostics for zero-or-one inflated beta (BIc) regression models with variable dispersion. The core idea is to use the conformal normal curvature of Poon and Poon (1999) <doi:10.1111/1467-9868.00162> as a non-conformity score within a split-conformal testing procedure, yielding per-observation conformal p-values whose Benjamini-Hochberg adjustment controls the false discovery rate at a user-specified level (Bates and others, 2023) <doi:10.1214/22-AOS2244>. Unlike classical local influence diagnostics, which rely on visual inspection of index plots and do not scale beyond a few hundred observations, 'clis' provides a finite-sample error guarantee and runs in linear time per observation after a single model fit. Methods for four perturbation schemes, block decomposition of influence into the inflation-probability and conditional-mean/precision components, penalised additive (semiparametric) submodels, and a full suite of diagnostic plots are included.
License: GPL-3
URL: https://github.com/Raydonal/clis, https://raydonal.github.io/clis/
BugReports: https://github.com/Raydonal/clis/issues
Encoding: UTF-8
Depends: R (≥ 4.1.0)
Imports: gamlss (≥ 5.4.0), gamlss.dist (≥ 6.0.0), stats, graphics, grDevices, utils
Suggests: testthat (≥ 3.0.0), knitr, rmarkdown, ggplot2, betareg, gamlss.data
VignetteBuilder: knitr
Config/testthat/edition: 3
NeedsCompilation: no
Packaged: 2026-09-19 00:32:12 UTC; raydonal
Author: Raydonal Ospina ORCID iD [aut, cre]
Maintainer: Raydonal Ospina <raydonal@de.ufpe.br>
Config/roxygen2/version: 8.1.0
Repository: CRAN
Date/Publication: 2026-09-29 14:00:17 UTC

clis: Conformal Local Influence Screening for Bounded-Response Regression

Description

Provides fast, statistically calibrated influence diagnostics for zero-or-one inflated beta (BIc) regression models with variable dispersion. The core idea is to use the conformal normal curvature of Poon and Poon (1999) doi:10.1111/1467-9868.00162 as a non-conformity score within a split-conformal testing procedure, yielding per-observation conformal p-values whose Benjamini-Hochberg adjustment controls the false discovery rate at a user-specified level (Bates and others, 2023) doi:10.1214/22-AOS2244. Unlike classical local influence diagnostics, which rely on visual inspection of index plots and do not scale beyond a few hundred observations, 'clis' provides a finite-sample error guarantee and runs in linear time per observation after a single model fit. Methods for four perturbation schemes, block decomposition of influence into the inflation-probability and conditional-mean/precision components, penalised additive (semiparametric) submodels, and a full suite of diagnostic plots are included.

Author(s)

Maintainer: Raydonal Ospina raydonal@de.ufpe.br (ORCID)

Authors:

See Also

Useful links:


Observed or Fisher information matrix for a BIc model

Description

Computes the block-structured information matrix of a fitted zero-or-one inflated beta regression with variable dispersion. By construction the matrix is block diagonal between the inflation parameters and the mean/precision parameters (information orthogonality).

Usage

bic_info(object, use_fisher = TRUE, penalty = NULL)

Arguments

object

A fitted gamlss model of family BEZI or BEOI.

use_fisher

Logical; if TRUE (default) returns the expected (Fisher) information, which is guaranteed positive definite; if FALSE returns the observed information.

penalty

Optional penalty matrix S of dimension ⁠(M+m+p) x (M+m+p)⁠ for penalised additive submodels. When supplied, the returned information is the penalised information J + S (or I + S), following the semiparametric extension. The penalty must be block diagonal across the gamma, beta, delta groups so that separability is preserved; see bic_penalty().

Value

A list with the information matrix (info), its inverse (info_inv), the parameter-block indices (idx), the weight sequences (weights), and (if a penalty was supplied) the effective degrees of freedom (edf) and their block split (edf_blocks).

References

Ospina, R. and Ferrari, S. L. P. (2012). A general class of zero-or-one inflated beta regression models. Computational Statistics & Data Analysis, 56(6), 1609-1623.


Description

Constructs a list with the link function, its inverse, and the first and second derivatives of the inverse link, used throughout the influence computations.

Usage

bic_link(link)

Arguments

link

Character string naming the link. One of "logit", "probit", "cloglog", or "log".

Value

A list with components name, linkfun, linkinv, mu.eta (first derivative of the inverse link), and mu.eta2 (second derivative).

Examples

lk <- bic_link("logit")
lk$mu.eta(0.3)   # d mu / d eta at mu = 0.3


Extract the penalty matrix from a fitted additive BIc model

Description

Assembles the block-diagonal penalty matrix S(lambda) from the smooth terms of a fitted gamlss model whose submodels use penalised additive terms (for example pb() P-spline terms). The result is suitable as the penalty argument of bic_info().

Usage

bic_penalty(object)

Arguments

object

A fitted gamlss model of family BEZI or BEOI with penalised additive terms in one or more submodels.

Details

The function reads the smoothing structure stored by gamlss for each parameter (mu, sigma, nu) and places the corresponding penalty contributions on the diagonal blocks. Terms without a penalty contribute a zero block, so a model with a mix of linear and smooth terms is handled transparently. By construction the returned matrix is block diagonal across the three coefficient groups, so it preserves the separability required by the semiparametric theory.

Value

A block-diagonal penalty matrix of dimension ⁠(M+m+p) x (M+m+p)⁠, block diagonal across the inflation, mean, and precision coefficient groups.

See Also

bic_info()


Randomized quantile residuals for BEZI/BEOI models

Description

Computes randomized quantile residuals for a fitted zero- or one-inflated beta model. For the continuous observations the beta CDF is scaled by the non-inflation probability; for the boundary observations the residual is randomized uniformly over the probability atom, which is the standard construction for a mixed discrete-continuous distribution. This is provided because the generic residuals() method does not always handle the inflation atom correctly, which can make every residual share the same sign.

Usage

bic_quantile_residuals(object, seed = NULL)

Arguments

object

A fitted gamlss model of family BEZI or BEOI.

seed

Optional integer seed for the randomization at the boundary.

Value

A numeric vector of randomized quantile residuals.


Conformal Local Influence Screening

Description

Performs scalable, FDR-controlled influence screening for a fitted zero-or-one inflated beta regression model. The conformal normal curvature score of each observation is used as a non-conformity score within a split-conformal procedure: the data are partitioned into a calibration set (presumed clean) and a screening set; conformal p-values are computed for the screening set; and Benjamini-Hochberg adjustment declares an influential subset with false discovery rate controlled at level alpha.

Usage

clis_screen(
  object,
  scheme = "caseweights",
  p = 1L,
  alpha = 0.1,
  calib_frac = 0.5,
  calib_idx = NULL,
  use_fisher = TRUE,
  r_max = 4L,
  score = c("B_Et", "m_r"),
  penalised = FALSE,
  penalty = NULL,
  seed = NULL
)

Arguments

object

A fitted gamlss model of family BEZI or BEOI.

scheme

Character; the perturbation scheme. One of "caseweights" (default), "disccovar", "meancovar", or "preccovar".

p

Integer; covariate index for the covariate-perturbation schemes.

alpha

Numeric in (0, 1); target false discovery rate. Default 0.1.

calib_frac

Numeric in (0, 1); fraction of observations used for conformal calibration when calib_idx is not supplied. Default 0.5.

calib_idx

Optional integer vector of observation indices to use as the calibration set. The conformal false discovery rate guarantee requires the calibration set to be (nearly) free of influential points; when a trusted clean subset is known, supply it here. If NULL (default) a random subset of size floor(calib_frac * n) is drawn, which is appropriate only when influential points are rare.

use_fisher

Logical; use the Fisher information (default TRUE).

r_max

Integer; order for aggregate contributions used as the score.

score

Character; which CNC quantity to use as the non-conformity score: "B_Et" (basic perturbation, default) or "m_r" (aggregate contribution at order r_max).

penalised

Logical; if TRUE, use the penalised information J + S obtained from the smooth terms of an additive gamlss fit, implementing the semiparametric extension. The penalty is extracted with bic_penalty(). Default FALSE.

penalty

Optional penalty matrix, of the dimension of the full coefficient vector, added to the information before inversion. Use this when the smooth basis is supplied explicitly as design columns.

seed

Optional integer for reproducible calibration splitting.

Details

Classical local influence diagnostics rank observations by a curvature measure and rely on visual inspection of an index plot against an ad-hoc threshold. This neither scales to large n nor provides any control of the type-I error rate. CLIS addresses both limitations: the conformal wrapper supplies a finite-sample FDR guarantee, and the per-observation score is computed in linear time after a single model fit.

The calibration set is assumed to be predominantly free of influential observations. Because influential points are typically rare, a random split satisfies this approximately; for adversarial settings, a robust pre-filter can be applied before calibration.

Value

An object of class clis with components:

influential

Integer indices (into the screening set) declared influential at FDR level alpha.

influential_global

Integer indices into the original data.

pvalues

Conformal p-values for the screening set.

padj

Benjamini-Hochberg adjusted p-values.

scores

Non-conformity scores for all observations.

calib_idx, screen_idx

Calibration and screening indices.

alpha, scheme

Inputs echoed back.

decomp

Block decomposition (case-weights scheme only).

Examples


if (requireNamespace("gamlss", quietly = TRUE) &&
    requireNamespace("betareg", quietly = TRUE)) {
  data("ReadingSkills", package = "betareg")
  ReadingSkills$dys <- as.numeric(ReadingSkills$dyslexia == "yes")
  fit <- gamlss::gamlss(accuracy1 ~ dys * iq,
                        sigma.formula = ~ dys, nu.formula = ~ iq,
                        family = gamlss.dist::BEOI, data = ReadingSkills,
                        control = gamlss::gamlss.control(trace = FALSE))
  res <- clis_screen(fit, alpha = 0.1, seed = 1)
  print(res)
}



Block decomposition of conformal normal curvature

Description

Decomposes the per-observation CNC scores into a contribution from the inflation-probability submodel and a contribution from the conditional-mean/precision submodel, exploiting the information orthogonality of the BIc model. The two contributions sum exactly to the total, with no cross terms.

Usage

cnc_block_decomp(delta_out, info_out)

Arguments

delta_out

A perturbation-matrix list (e.g. from delta_caseweights()) with a full Delta component.

info_out

An information list from bic_info() with info_inv and idx.

Value

A list with the discrete contribution B_gamma, the continuous contribution B_betadelta, their sum B_total, and the discrete fraction ratio_gamma.


Conformal normal curvature matrix

Description

Computes the Frobenius-normalised conformal normal curvature (CNC) matrix of Poon and Poon (1999) for a given perturbation matrix and information matrix inverse.

Usage

cnc_matrix(Delta, info_inv)

Arguments

Delta

A perturbation matrix of dimension ⁠(M+m+p) x n⁠, e.g. from delta_caseweights().

info_inv

The inverse information matrix from bic_info().

Value

A list with the ⁠n x n⁠ CNC matrix B, its eigenvalues lambda (normalised so that the sum of squares is one), eigenvectors vectors, and the Frobenius norm normF of the unnormalised curvature.

References

Poon, W.-Y. and Poon, Y. S. (1999). Conformal normal curvature and assessment of local influence. Journal of the Royal Statistical Society: Series B, 61(1), 51-61.


Per-observation conformal normal curvature scores

Description

Computes the basic-perturbation CNC scores B_{E_t} and the aggregate contribution measures m[r]_t from a fitted CNC matrix.

Usage

cnc_scores(cnc, r_max = 4L)

Arguments

cnc

A list returned by cnc_matrix().

r_max

Integer; the maximum order of r-influential eigenvectors to consider for the aggregate contributions.

Value

A list with per-observation scores B_Et, the matrix of aggregate contributions m_r (one column per r), the eigenvalue thresholds threshold_r, the aggregate-contribution thresholds threshold_mt, the number of r-influential eigenvectors k_r, and the cutoff b2 for B_Et.


Linear-time per-observation CNC scores

Description

Computes the basic-perturbation CNC scores B_{E_t} without ever forming the n \times n curvature matrix. The score is the scaled diagonal of F_0 = \Delta^\top \mathcal{I}^{-1}\Delta, and both the diagonal and the Frobenius norm \|F_0\|_F are obtained from quantities of dimension (M+m+p), so the cost is linear in n and the memory footprint is negligible. This is the scalable path used for large samples, where the dense n\times n eigenproblem of cnc_matrix() is infeasible.

Usage

cnc_scores_linear(Delta, info_inv)

Arguments

Delta

A perturbation matrix of dimension ⁠(M+m+p) x n⁠.

info_inv

The inverse information matrix from bic_info().

Value

A list with per-observation scores B_Et, the cutoff b2, the Frobenius norm normF, and n.


Conformal p-values from non-conformity scores

Description

Given calibration scores (assumed to come from non-influential, "clean" observations) and test scores, computes marginal conformal p-values in the sense of Bates and others (2023). Larger scores indicate stronger evidence of influence, so the p-value for a test point is the calibrated rank of its score among the calibration scores.

Usage

conformal_pvalues(cal_scores, test_scores)

Arguments

cal_scores

Numeric vector of calibration non-conformity scores.

test_scores

Numeric vector of test non-conformity scores.

Details

For a test score s and calibration scores c_1, \ldots, c_n, the conformal p-value is

p = \frac{1 + \#\{i : c_i \ge s\}}{n + 1}.

These p-values are marginally valid (super-uniform under the null that the test point is exchangeable with the calibration set) and, by the positive-dependence result of Bates and others (2023), permit Benjamini-Hochberg FDR control.

Value

A numeric vector of conformal p-values, one per test score.

References

Bates, S., Candes, E., Lei, L., Romano, Y. and Sesia, M. (2023). Testing for outliers with conformal p-values. The Annals of Statistics, 51(1), 149-178.

Examples

set.seed(1)
cal  <- rnorm(100)
test <- c(rnorm(8), 4, 5)   # last two are outliers
conformal_pvalues(cal, test)


Perturbation matrix for the case-weights scheme

Description

Perturbation matrix for the case-weights scheme

Usage

delta_caseweights(object)

Arguments

object

A fitted BEZI/BEOI gamlss model.

Value

A list with the full Delta matrix and its gamma, beta, delta blocks.


Perturbation matrix for the discrete-covariate scheme

Description

Perturbation matrix for the discrete-covariate scheme

Usage

delta_disccovar(object, p = 1L)

Arguments

object

A fitted BEZI/BEOI gamlss model.

p

Index of the perturbed covariate in the inflation design matrix.

Value

A list with the full Delta and its blocks.


Perturbation matrix for the mean-covariate scheme

Description

Perturbation matrix for the mean-covariate scheme

Usage

delta_meancovar(object, p = 1L)

Arguments

object

A fitted BEZI/BEOI gamlss model.

p

Index of the perturbed covariate in the mean design matrix.

Value

A list with the full Delta and its blocks.


Perturbation matrix for the precision-covariate scheme

Description

Perturbation matrix for the precision-covariate scheme

Usage

delta_preccovar(object, p = 1L)

Arguments

object

A fitted BEZI/BEOI gamlss model.

p

Index of the perturbed covariate in the precision design matrix.

Value

A list with the full Delta and its blocks.


Simulated envelope for quantile residuals

Description

Simulated envelope for quantile residuals

Usage

envelope_bic(object, B = 200L, level = 0.95, seed = NULL)

Arguments

object

A fitted BEZI/BEOI gamlss model.

B

Integer; number of simulated samples.

level

Numeric; envelope coverage (default 0.95).

seed

Optional integer seed.

Value

Invisibly returns the observed residuals and envelope bands.


National DTP3 vaccination coverage (2022)

Description

Reads country-level proportions of diphtheria-tetanus-pertussis (DTP3) vaccine coverage for 2022, with development covariates, from the CSV shipped in the package's extdata directory. The response places a point mass at one (countries achieving complete coverage), making it a natural application of one-inflated beta (BEOI) regression.

Usage

load_vaccination()

Value

A data frame with one row per country and the columns

iso3c

ISO 3166-1 alpha-3 country code.

dtp3

DTP3 coverage proportion in [0,1] (response).

ln_gdp

Natural log of GDP per capita (PPP, constant USD).

urb

Urbanisation rate (proportion of urban population).

ln_pop

Natural log of total population.

hdi

Human Development Index in [0,1].

Source

WHO/UNICEF Joint Monitoring Programme (coverage); World Bank (GDP, urbanisation, population); UNDP Human Development Report (HDI).

Examples

vaccination <- load_vaccination()
summary(vaccination$dtp3)
mean(vaccination$dtp3 == 1)   # fraction at the boundary


Plot a CLIS screening result

Description

Produces a two-panel diagnostic plot for a clis_screen() result: a plot of conformal p-values (with the Benjamini-Hochberg rejection boundary) and a plot of the non-conformity scores with the declared influential observations highlighted.

Usage

plot_clis(x, label_top = 8L, ...)

Arguments

x

A clis object from clis_screen().

label_top

Integer; how many top influential points to label.

...

Currently ignored.

Value

Invisibly returns x.


Plot the classical CNC diagnostic panels

Description

Plot the classical CNC diagnostic panels

Usage

plot_cnc_panels(
  cnc,
  scores,
  decomp = NULL,
  r_plot = 3L,
  label_top = 5L,
  suffix = ""
)

Arguments

cnc

A list from cnc_matrix().

scores

A list from cnc_scores().

decomp

Optional list from cnc_block_decomp() for a third panel.

r_plot

Integer; which aggregate-contribution order to plot.

label_top

Integer; number of points to label.

suffix

Character; appended to panel titles.

Value

Invisibly returns scores.


Classical local-influence index plot

Description

Reproduces the standard influence display of the beta-regression diagnostics literature: the per-observation conformal normal curvature B_{E_t} (or the aggregate contribution m[r]_t) plotted against the observation index, with the reference cutoff drawn and the most influential observations labelled. This is the plot practitioners of local influence expect to see; the conformal screening in clis_screen() and plot_clis() complements it with an error-controlled declaration.

Usage

plot_influence(
  object,
  scheme = "caseweights",
  p = 1L,
  measure = c("B_Et", "m_r"),
  r_plot = 3L,
  use_fisher = TRUE,
  penalised = FALSE,
  label_top = 5L,
  labels = NULL
)

Arguments

object

A fitted BEZI/BEOI gamlss model.

scheme

Perturbation scheme: "caseweights" (default), "disccovar", "meancovar", or "preccovar".

p

Covariate index for the covariate-perturbation schemes.

measure

Which quantity to plot: "B_Et" (default) or "m_r".

r_plot

Aggregate-contribution order when measure = "m_r".

use_fisher

Logical; use the Fisher information (default TRUE).

penalised

Logical; use the penalised information (default FALSE).

label_top

Integer; number of most-influential points to label.

labels

Optional character vector of observation labels (for example country codes); defaults to the observation index.

Value

Invisibly, a list with the plotted scores and the cutoff.


Diagnostic residual plots for a BIc model

Description

Diagnostic residual plots for a BIc model

Usage

plot_residuals(object, label_top = 3L)

Arguments

object

A fitted BEZI/BEOI gamlss model.

label_top

Integer; number of extreme points to label.

Value

Invisibly returns a list of residual vectors.