| Type: | Package |
| Title: | Experimental Design and Inference |
| Version: | 1.0.2 |
| Description: | Implements a comprehensive suite of experimental designs, both fixed (e.g., block, stratified, matched-pair, cluster, factorial, and mixed-integer-programming-based optimal designs) and sequential (including matching-on-the-fly designs, biased coin designs, and covariate-adaptive urn designs that assign treatment one subject at a time while maintaining covariate balance), for continuous, incidence, count, proportion, survival, and ordinal response types. For each design and response type combination, provides the corresponding inference procedures, including exact, asymptotic, distribution-free, and resampling-based (bootstrap, jackknife, and randomization) methods, so that estimation and testing are always matched to how the data were generated. An 'InferenceSuite' facility runs all applicable inference procedures for a given design and response type at once and reports a single Cauchy-combined p-value summarizing their evidence. A built-in simulation framework supports power analysis and operating-characteristic studies across designs, response types, and inference procedures, with optional parallelization via 'mirai'. Missing covariate data is handled automatically via built-in imputation. Core numerical routines are implemented in C++ via 'Rcpp' for speed on large designs and simulation studies. Machine-specific tuning for optimization is included. |
| License: | GPL-3 |
| Encoding: | UTF-8 |
| Language: | en-US |
| Depends: | R (≥ 3.5.0) |
| LinkingTo: | Rcpp, RcppEigen, RcppNumerical |
| Imports: | R6, checkmate, Rcpp, data.table, survival, missRanger, missForest, numDeriv, digest, methods, MASS, RhpcBLASctl, randomizr |
| Suggests: | anticlust, blockTools, clinfun, coin, dplyr, mirai, aftgee, betareg, copula, doParallel, ggplot2, fixest, gamlss.dist, geepack, glmmTMB, glmnet, icenReg, interval, jsonlite, knitr, lme4, multgee, nbpMatching, nlme, nnet, ordinal, pscl, qpdf, quantreg, R.utils, RcppXPtrUtils, Rfit, rmarkdown, testthat, VGAM, ompr, ompr.roi, ROI.plugin.glpk, sandwich, withr |
| VignetteBuilder: | knitr |
| Collate: | 'design_abstract.R' 'design_blocking_abstract.R' 'design_matching_abstract.R' 'design_fixed_abstract.R' 'design_component_registry.R' 'design_class_factory.R' 'design_fixed_greedy.R' 'helper_optimal_shared.R' 'helper_optimal_milp_solvers.R' 'helper_user_compiled_fn.R' 'helper_optimal_annealing.R' 'design_fixed_greedy_d_optimal.R' 'design_fixed_bernoulli.R' 'design_fixed_binary_match.R' 'design_fixed_blocked_cluster.R' 'design_fixed_blocking.R' 'design_fixed_cluster.R' 'design_fixed_factorial.R' 'design_fixed_ibcrd.R' 'design_fixed_matching_greedy_pair_switching.R' 'design_fixed_optimal.R' 'design_fixed_optimal_blocks.R' 'design_fixed_rerandomization.R' 'design_observational.R' 'design_observational_blocks.R' 'design_observational_matching.R' 'design_seq_one_by_one_abstract.R' 'design_custom_extensions.R' 'design_seq_one_by_one_atkinson.R' 'design_seq_one_by_one_bernoulli.R' 'design_seq_one_by_one_efron.R' 'design_seq_one_by_one_ibcrd.R' 'design_seq_one_by_one_KK14.R' 'design_seq_one_by_one_KK21_stepwise.R' 'design_seq_one_by_one_KK21.R' 'design_seq_one_by_one_pocock_simon.R' 'design_seq_one_by_one_random_block_size.R' 'design_seq_one_by_one_spbr.R' 'design_seq_one_by_one_urn.R' 'design_class_registry.R' 'EDI.R' 'helper_glm_fit.R' 'globals.R' 'comprehensive_slow_paths.R' 'helper_additional_asserts.R' 'contracts_mixins.R' 'helper_robust_sandwich.R' 'inference_all_abstract_mle_or_KM_summary_table.R' 'inference_ext_ci_inversion.R' 'inference_ext_information_matrix.R' 'inference_ext_likelihood_test_memoization.R' 'inference_mixin_off_optimum_likelihood_eval.R' 'inference_all_abstract_asymp_lik.R' 'inference_all_abstract_marginal_estimand.R' 'inference_ext_bartlett_approx.R' 'inference_mixin_cordeiro_ferrari_approx.R' 'inference_mixin_lemonte_gradient_approx.R' 'inference_mixin_kk_gee_shared.R' 'inference_mixin_kk_glmm_shared.R' 'inference_mixin_kk_passthrough.R' 'inference_mixin_kk_passthrough_compound.R' 'helper_bootstrap_ci.R' 'contracts_resampling_draws.R' 'inference_ext_bca_bootstrap_ci.R' 'inference_ext_exchangeable_resampling_units.R' 'inference_ext_minimum_volatility_selector.R' 'inference_ext_m_out_of_n_bootstrap.R' 'inference_ext_prw_subsampling.R' 'inference_all_abstract_non_param_boot.R' 'inference_all_abstract_rand_bootstrap.R' 'inference_all_abstract_rand_bootstrap_ci.R' 'inference_all_abstract_bayesian_bootstrap.R' 'inference_all_abstract_jackknife.R' 'inference_all_abstract_exact.R' 'inference_all_abstract_asymp.R' 'inference_ext_param_bootstrap_estimate.R' 'inference_all_abstract_param_boot.R' 'inference_all_abstract_asymp_lik_std_mod_cache.R' 'inference_all_abstract_count_likelihood.R' 'inference_ext_quantile_rand_ci.R' 'inference_all_abstract_rand_ci.R' 'inference_ext_sequential_mc_pval.R' 'inference_all_abstract_rand.R' 'inference_all_abstract.R' 'inference_all_abstract_KK_passthrough_compound.R' 'inference_all_abstract_quantile_rand_ci.R' 'inference_all_KK_mean_diff_IVWC.R' 'inference_all_KK_quantile_regr_ivwc_abstract.R' 'inference_all_KK_quantile_regr_one_lik_abstract.R' 'inference_all_KK_wilcox_ivwc.R' 'inference_all_average_diff.R' 'inference_all_simple_mean_diff_pooled_var.R' 'inference_all_simple_wilcox.R' 'inference_continuous_KK_bai_abstract.R' 'inference_continuous_KK_glmm.R' 'inference_continuous_KK_ols_ivwc.R' 'inference_continuous_KK_ols_one_lik.R' 'inference_continuous_KK_quantile_regr_ivwc.R' 'inference_continuous_KK_quantile_regr_one_lik.R' 'inference_continuous_KK_robust_regr_ivwc.R' 'inference_continuous_KK_robust_regr_one_lik.R' 'inference_continuous_KK14_bai.R' 'inference_continuous_KK21_bai.R' 'inference_continuous_lin.R' 'inference_continuous_ols.R' 'inference_continuous_quantile_regr.R' 'inference_continuous_robust_regr.R' 'inference_count_composite_likelihood.R' 'inference_count_KK_combined.R' 'inference_count_KK_cond_poisson.R' 'inference_count_KK_gee.R' 'inference_count_negbin.R' 'inference_count_poisson.R' 'inference_count_quasipoisson.R' 'inference_count_robust_poisson.R' 'inference_count_zero_augmented_poisson_abstract.R' 'inference_count_hurdle.R' 'inference_count_zero_inflated.R' 'inference_custom_extensions.R' 'inference_rand_custom.R' 'inference_helpers_zhang.R' 'inference_incidence_binomial_identity.R' 'inference_incidence_cmh.R' 'inference_incidence_exact_binomial.R' 'inference_incidence_exact_zhang.R' 'inference_incidence_extended_robins.R' 'helper_gcomp.R' 'helper_marginal_estimand.R' 'inference_incidence_gcomp_abstract.R' 'inference_incidence_gcomp.R' 'inference_incidence_KK_cond_logit_glmm_abstract.R' 'inference_incidence_KK_cond_logit_glmm.R' 'inference_incidence_KK_cond_logit.R' 'inference_incidence_KK_combined.R' 'inference_incidence_KK_gcomp_abstract.R' 'inference_incidence_KK_marginal_abstract.R' 'inference_incidence_KK_marginal.R' 'inference_incidence_KK_newcombe_ivwc_univ.R' 'inference_incidence_log_binomial.R' 'inference_incidence_logit.R' 'inference_incidence_miettinen_nurminen_univ.R' 'inference_incidence_modified_poisson.R' 'inference_incidence_newcombe_univ.R' 'inference_incidence_probit.R' 'inference_incidence_risk_diff.R' 'inference_incid_wald.R' 'inference_indicidence_exact_fisher.R' 'inference_ordinal_adj_cat_logit.R' 'inference_ordinal_cauchit.R' 'inference_ordinal_cloglog.R' 'inference_ordinal_gcomp.R' 'inference_ordinal_jonckheere_terpstra_test.R' 'inference_ordinal_KK_clmm_abstract.R' 'inference_ordinal_KK_combined.R' 'inference_ordinal_KK_cond_adj_cat_logit.R' 'inference_ordinal_KK_cond_logit_abstract.R' 'inference_ordinal_ordered_probit.R' 'inference_ordinal_paired_sign_test.R' 'inference_ordinal_partial_proportional_odds.R' 'inference_ordinal_proportional_odds.R' 'inference_ordinal_ridit.R' 'inference_ordinal_stereotype_logit.R' 'inference_proportion_beta.R' 'inference_proportion_fractional_logit.R' 'inference_proportion_gcomp.R' 'inference_proportion_KK_combined.R' 'inference_proportion_KK_quantile_regr_ivwc.R' 'inference_proportion_KK_quantile_regr_one_lik.R' 'inference_proportion_quantile_regr.R' 'inference_proportion_zero_one_inflated_beta.R' 'inference_suite.R' 'inference_survival_coxph.R' 'inference_survival_dep_cens_transform.R' 'inference_survival_gehan_wilcox.R' 'inference_survival_GLMM_weibull_frailty_loggamma.R' 'inference_survival_KK_lwa_cox_ivwc_abstract.R' 'inference_survival_KK_lwa_cox_one_lik_abstract.R' 'inference_survival_KK_lwa_cox.R' 'inference_survival_KK_rank_regr_ivwc_abstract.R' 'inference_survival_KK_rank_regr.R' 'inference_survival_KK_strat_cox.R' 'inference_survival_GLMM_weibull_frailty_normal.R' 'inference_survival_KK_weibull_marginal.R' 'inference_survival_km_diff.R' 'inference_survival_log_rank.R' 'inference_survival_rmst.R' 'inference_survival_strat_cox.R' 'inference_survival_weibull.R' 'inference_class_registry.R' 'helper_package_checks.R' 'helper_math.R' 'helper_model_matrix.R' 'helper_response_asserts.R' 'helper_robust_regression.R' 'helper_rcpp_doc_stubs.R' 'helper_matching.R' 'helper_survival_fits.R' 'helper_zoib.R' 'helper_inference_survival_turnbull.R' 'RcppExports.R' 'simulations_framework.R' 'simulation_framework_report.R' 'local_machine_tuning_synthetic_fixtures.R' 'local_machine_tuning_harness.R' 'local_machine_tuning_axes.R' 'local_machine_tuning_persistence.R' 'local_machine_tuning_correctness.R' 'local_machine_tuning.R' 'zzz.R' |
| URL: | https://github.com/kapelner/EDI, https://kapelner.github.io/EDI/, https://pypi.org/project/edi-kernels/ |
| BugReports: | https://github.com/kapelner/EDI/issues |
| Config/roxygen2/version: | 8.0.0.9000 |
| NeedsCompilation: | yes |
| Packaged: | 2026-09-25 14:50:44 UTC; kapelner |
| Author: | Adam Kapelner |
| Maintainer: | Adam Kapelner <adam.kapelner@mail.huji.ac.il> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-07 09:00:15 UTC |
Experimental Design and Inference
Description
EDI
Details
Provides comprehensive support for many fixed and sequential experimental designs and many infererential methods (parametric, nonparametric, exact) for response types continuous, incidence, count, proportion, survival (with censoring) and ordinal. Supports automatic missing data imputation, parallelization, provides robustness fallbacks and is optimized with C++.
Author(s)
Adam Kapelner kapelner@qc.cuny.edu
References
Adam Kapelner and Abba Krieger A Matching Procedure for Sequential Experiments that Iteratively Learns which Covariates Improve Power, Arxiv 2010.05980
See Also
Useful links:
Report bugs at https://github.com/kapelner/EDI/issues
Abstract Quantile Regression Compound Estimator for KK Matching-on-the-Fly Designs
Description
An abstract base class providing shared quantile regression logic for KK matching-on-the-fly
designs. Subclasses override the transform_y_fn private$m field to apply a response
transformation before quantile regression (e.g., identity for continuous,
qlogis for proportion outcomes).
Usage
.init_kk_quantile_regr_ivwc(
self,
private,
super,
des_obj,
model_formula,
tau,
transform_y_fn,
verbose,
smart_cold_start_default
)
Abstract Quantile Regression Combined-Likelihood Compound Estimator for KK Designs
Description
Fits a single joint quantile regression over all KK design data by stacking matched-pair differences and reservoir observations into one design matrix.
Usage
.init_kk_quantile_regr_one_lik(
self,
private,
super,
des_obj,
model_formula,
tau,
transform_y_fn,
verbose
)
Details
Column layout of X_stack: [beta_0 | beta_T | beta_xs (p cols)] Pair rows: [0 | 1 | Xd_k] -> Q_tau(yd_k) = beta_T + Xd_k' beta_xs Reservoir rows: [1 | w_i | X_i] -> Q_tau(y_i) = beta_0 + w_i*beta_T + X_i'*beta_xs Fitting a single rq() on the stacked dataset minimises the combined check-function loss.
Special cases: Pairs only: beta_0 column is all-zero and dropped; layout [beta_T | beta_xs]. Reservoir only: standard quantile regression layout [beta_0 | beta_T | beta_xs].
Standard errors use Powell's "nid" sandwich estimator, falling back to "iid".
Normalize and Validate an Optimizer Algorithm Name for the fast_* C++ Backends
Description
Internal helper shared by the package's fast_* GLM/survival/ordinal fitting
wrappers (e.g. fast_logistic_regression,
fast_coxph_regression) to resolve a user-supplied
optimization_alg argument to one of the fixed set of optimizer names the
underlying C++ backends actually implement, applying a model-specific default when
none is supplied and rejecting anything else. This centralizes the
default/validation logic so each fast_* wrapper does not have to repeat it.
Usage
.normalize_optimizer_algorithm(
optimization_alg,
allow_irls = FALSE,
default = if (allow_irls) "irls" else "lbfgs"
)
Arguments
optimization_alg |
Character string (possibly abbreviated) naming the desired
optimizer, |
allow_irls |
Logical. Whether |
default |
Character string used when |
Details
The three possible optimizer names, when supported by a given model, correspond to
distinct fitting algorithms in the C++ backends: "newton_raphson" (full
Newton-Raphson using the analytic Hessian), "lbfgs" (limited-memory
quasi-Newton, avoiding an explicit Hessian), and "irls" (iteratively
reweighted least squares, the classical GLM-fitting algorithm — only meaningful,
and only offered, for exponential-family GLMs, hence gated by allow_irls).
Which optimizers a given fast_* function actually accepts (and which is its
default) varies by model; this function only encodes the generic
irls-vs-not-irls split, not per-model specifics.
optimization_alg is matched against the allowed set via
match.arg, so unambiguous partial string matches (e.g.
"newton") are accepted; an unmatched or ambiguous value raises
match.arg's standard error rather than silently falling back to the default.
missing(optimization_alg) or an explicit NULL both resolve to
default before matching.
Value
A validated, unabbreviated character string: one of
"newton_raphson", "lbfgs", or (only when allow_irls = TRUE)
"irls".
See Also
match.arg, which performs the validation/partial-matching.
Inference based on Maximum Likelihood for KK designs
Description
Initialize Bai adjusted-t inference for a completed KK continuous-response design, including the optional convex combination of matched-pair and reservoir estimates.
Computes the appropriate estimate for compound mean difference across pairs and reservoir
Computes a 1-alpha level frequentist confidence interval
Here we use the theory that MLE's computed for GLM's are asymptotically normal
(except in the case
of estimat_type "median difference" where a nonparametric bootstrap confidence
interval (see the controlTest::quantileControlTest method)
is employed. Hence these confidence intervals are asymptotically valid and thus
approximate for any sample size.
Compute the Bai-adjusted two-sided p-value for the treatment
effect using the matched-design adjusted statistic. See related
InferenceBaiAdjustedTKK14
methods.
Usage
BaiAdjustedTSource
Details
Inference for mean difference. Note that warm starts are disabled for this class as the Bai adjusted t-test is a closed-form estimator and does not benefit from initialization.
This class requires the nbpMatching package, which is listed in Suggests
and is not installed automatically with EDI. Install it manually with
install.packages("nbpMatching") before using this class.
Value
The setting-appropriate (see description) numeric estimate of the treatment effect
A (1 - alpha)-sized frequentist confidence interval for the treatment effect
The approximate frequentist p-value
Examples
# (loading the nbpMatching package alone takes a few seconds)
if (requireNamespace("nbpMatching", quietly = TRUE)) {
seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceBaiAdjustedTKK14$new(seq_des)
seq_des_inf$compute_estimate()
}
# (loading the nbpMatching package alone takes a few seconds)
if (requireNamespace("nbpMatching", quietly = TRUE)) {
seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceBaiAdjustedTKK14$new(seq_des)
seq_des_inf$compute_asymp_confidence_interval()
}
# (loading the nbpMatching package alone takes a few seconds)
if (requireNamespace("nbpMatching", quietly = TRUE)) {
seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceBaiAdjustedTKK14$new(seq_des)
seq_des_inf$compute_asymp_two_sided_pval()
}
Robust-Regression IVWC Compound Inference for KK Designs
Description
Initialize KK inverse-variance combined robust-regression
inference and prepare the matched/reservoir components used by
InferenceContinKKRobustRegrIVWC.
Point estimate of the treatment effect combining a matched-pair robust
regression (MASS::rlm on within-pair differences) and a reservoir robust
regression (treatment vs. control), combined by inverse-variance weighting: the two
component estimates \hat\beta_{T,\text{matched}} and
\hat\beta_{T,\text{reservoir}} (each with its own estimated variance) are
combined as \hat\beta_T = \sum_k w_k \hat\beta_{T,k} / \sum_k w_k with
w_k = 1/\widehat{\mathrm{Var}}(\hat\beta_{T,k}); a component is dropped from the
combination if it is not estimable (e.g. too few observations). See
Inference for the general estimate contract.
Wald confidence interval for the inverse-variance-combined treatment
effect: \hat\beta_T \pm t_{1-\alpha/2,\,df}\cdot \hat{se}(\hat\beta_T), using the
combined variance 1/\sum_k w_k from compute_estimate()'s inverse-variance
weighting. likelihood_tier = "quasi" for this class (M/MM-estimator objective,
not a normalized likelihood), so only the Wald testing type is supported. See
InferenceAsymp for the shared contract.
Two-sided Wald p-value for H_0: \beta_T = \code{delta} vs.
H_1: \beta_T \neq \code{delta}, using the inverse-variance-combined estimate and
standard error from compute_estimate(). See
InferenceAsymp for the shared Wald/asymptotic
semantics.
Duplicate the robust-regression inference object while
preserving the selected match-specific formulas and clearing cached fit
results; see Inference for the common
duplication contract.
Usage
ContinKKRobustRegrIVWCSource
Details
Fits a variance-weighted compound estimator for KK matching-on-the-fly designs with continuous responses using robust linear regression ('MASS::rlm') for the matched-pair and reservoir components separately.
Model. Two robust M/MM-estimator regressions are fit independently:
one on the within-pair outcome differences for matched pairs (with covariate
differences as predictors, no intercept), and one on the reservoir subjects
(raw covariates plus a treatment indicator). Each yields a treatment-effect
estimate \hat\beta_{T,k} and estimated variance. The two are combined
by inverse-variance weighting into a single \hat\beta_T, the same
combination rule used by InferenceContinKKOLSIVWC
but with robust rather than OLS component fits. A component missing enough
data (e.g. no matched pairs) is dropped from the combination.
Robust fitting. Uses MASS::rlm (or an internal Rcpp IRLS
kernel when use_rcpp = TRUE, the default) with method = "M" or
"MM"; "MM" (the default) uses an LQS-based high-breakdown
start, while "M" can optionally warm-start from OLS
(start_with_ols = TRUE). likelihood_tier = "quasi": the
M/MM objective is not a normalized likelihood, so only Wald-type asymptotic
inference is available (no score/gradient/likelihood-ratio testing types).
Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring; a KK matching-on-the-fly design. Robust regression down-weights outlying residuals, trading some efficiency under exactly-Gaussian errors for resistance to heavy tails and contamination.
Value
Numeric scalar treatment-effect estimate on the outcome's natural scale.
A length-2 numeric vector c(lower, upper), or NA bounds if
nonestimable.
Numeric scalar p-value in [0, 1], or NA_real_ if nonestimable.
References
Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential
allocation with higher power and efficiency. Biometrics, 70(2),
378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)
See Also
Analogous Python API for robust linear models: statsmodels RLM. Robust regression (orientation).
Legacy class. Not fully tested in comprehensive_tests.R.
Count Composite Likelihood Inference Base
Description
Computes the treatment estimate.
Usage
CountCompositeLikelihoodSource
Details
Shared branch for count models whose reported estimator is robust or quasi-likelihood based.
Conditional-Poisson Inference for KK Designs with Combined Likelihood
Description
Initialize conditional-Poisson one-likelihood inference for
KK count designs and prepare the combined likelihood used by
InferenceCountKKCondPoissonOneLik.
Compute the conditional-Poisson one-likelihood treatment estimate by fitting the combined matched/reservoir likelihood and caching the treatment log-rate coefficient for related likelihood-test methods.
Recomputes the combined conditional-Poisson estimate under Bayesian-bootstrap weights.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Uses the shared Wald confidence-interval contract; see
InferenceAsymp.
Computes a Wald two-sided p-value.
Computes a design-adjusted score confidence interval.
Computes a design-adjusted likelihood-ratio confidence interval.
Computes a design-adjusted gradient confidence interval.
Computes a design-adjusted score p-value.
Computes a design-adjusted likelihood-ratio p-value.
Computes a design-adjusted gradient p-value.
Usage
CountKKCondPoissonOneLikLikelihoodSource
KK Hurdle Poisson Combined-Likelihood Inference for Count Responses
Description
Initialize KK hurdle-Poisson one-likelihood inference for
count responses and prepare the combined matched/reservoir likelihood.
See InferenceCountKKHurdlePoissonOneLik
and InferenceParamBootstrap
for related likelihood and bootstrap methods.
Compute the one-likelihood hurdle-Poisson treatment-effect estimate by fitting the combined count likelihood and caching the treatment log-rate coefficient for related p-value and interval methods.
Recomputes the combined hurdle-Poisson estimate under Bayesian-bootstrap weights.
Compute the configured asymptotic confidence interval for the
one-likelihood hurdle-Poisson treatment coefficient, delegating to Wald,
score, likelihood-ratio, or gradient paths as documented in
InferenceCountLikelihood.
Computes a design-conservative score confidence interval.
Computes a design-conservative likelihood-ratio confidence interval.
Computes a design-conservative gradient confidence interval.
Compute the configured asymptotic two-sided p-value for the
one-likelihood hurdle-Poisson treatment coefficient, delegating to Wald,
score, likelihood-ratio, or gradient paths as documented in
InferenceCountLikelihood.
Computes a design-conservative score p-value.
Computes a design-conservative likelihood-ratio p-value.
Computes a design-conservative gradient p-value.
Compute the Wald confidence interval for the one-likelihood
hurdle-Poisson treatment coefficient, falling back to bootstrap when the
model standard error is unavailable. See
InferenceAsymp.
Compute the Wald two-sided p-value for the one-likelihood
hurdle-Poisson treatment coefficient, falling back to the Bayesian-bootstrap
p-value when the model standard error is unavailable. See
InferenceAsymp.
Usage
CountKKHurdlePoissonOneLikLikelihoodSource
An Abstract Experimental Design
Description
Internal method. An abstract R6 Class encapsulating the data and functionality for an experimental design. This class takes care of data storage and response handling.
Details
Throughout the package, treatment assignment vectors w use the
\{0, 1\} encoding: 1 indicates a treated subject and 0
a control subject. All public methods that return or accept w
(e.g. get_w(), draw_ws_according_to_design()) use this
convention. A handful of variance estimators (e.g. InferenceIncidCMH,
InferenceIncidExtendedRobins) recode to a signed \{-1,+1\}
contrast internally where their formulas require it; that recoding is
local to those classes and does not affect this public convention.
Saving and loading
Design (and its DesignSeqOneByOne subclasses) is the unit of
persistence for a trial. Persist a des_obj with base R's
saveRDS()/readRDS() – there is no dedicated
save_edi_design()/load_edi_design() wrapper, and none is
planned: the audit behind this section found nothing that needs
transformation on load beyond what is documented here. Inference*
objects are disposable, cheaply reconstructed from a Design object
on demand (see each class's $new()), and must never be
saveRDS()'d directly – nothing currently prevents it (they
serialize "successfully" like any R6 object), but the result is a frozen
snapshot a user could easily mistake for something that stays live against
the design, and re-running inference from a reloaded Design is both
cheap and the only tested path.
Worked example (mirrors the round-trip tests in
R/EDI/tests/testthat/test-save-load-design.R):
des_obj = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "continuous")
for (i in 1:10) {
des_obj$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
des_obj$add_one_subject_response(i, y = rnorm(1))
}
saveRDS(des_obj, "trial.rds", version = 2)
# ...new R session...
des_obj = readRDS("trial.rds")
for (i in 11:20) {
des_obj$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
des_obj$add_one_subject_response(i, y = rnorm(1))
}
inf_obj = InferenceContinOLS$new(des_obj) # reconstructed fresh, never persisted
inf_obj$compute_estimate()
Passing version = 2 to saveRDS() is recommended, matching the
one existing internal precedent for RDS serialization in this package
(SimulationFramework's replication cache); it is not required for a
same-R-version round trip.
Version stamp. Every Design object records the package
version it was constructed under (get_edi_version_created()). This
is stamped once at construction and is not refreshed by
readRDS() – it reflects the version that originally built the
object, not whatever version is currently loaded. The first "resume the
trial" call after a reload (draw_ws_according_to_design() for fixed
designs, add_one_subject_to_experiment_and_assign() for sequential
designs) compares the stamped version's major component against the
currently loaded package's major component and emits a one-time
warning() on a mismatch; minor/patch differences are silent, since
most field additions are additive under this class's
lock_objects = FALSE R6 fields and do not warrant nagging on every
routine upgrade. Objects saved before this field existed self-initialize
it to the currently loaded version the first time it is read, rather than
erroring on the missing field.
RNG/reproducibility caveat. private$seed is consumed only
once, inside maybe_set_seed() at construction time, and is
not re-applied on readRDS(). Continuing to enroll subjects
after a reload therefore draws from whatever the global
.Random.seed happens to be in the new session, not a deterministic
continuation of the original stream. This is almost certainly the right
behavior for a real trial (bit-for-bit-reproducible continuation across a
process restart is not a property a production trial should have), but it
means a same-seed reload-and-continue is not expected to reproduce
the same draws as an uninterrupted run with that seed – do not rely on
that for testing.
Known non-serializable case. A DesignFixedOptimal
constructed with objective = "custom" from a raw
RcppXPtrUtils::cppXPtr() external pointer (rather than a C++ source
string) cannot be safely reloaded: compiled function pointers do not
survive a saveRDS()/readRDS() round trip, and there is no
retained source to recompile from. This is detected on first use after
reload and raises a clear error rather than failing silently; supply
custom_objective as a C++ source string instead of a pre-built
cppXPtr() object if you need this design to survive a save/reload
cycle – that form recompiles itself automatically the first time it is
used post-reload. Every other audited private cache on Design and
its components (all_subject_data_cache, permutations_cache,
lin_centered_covariates, matching/blocking/cluster component state
such as m, xm_structural, boot_pair_rows) was traced
to its originating C++ return type and confirmed to hold only plain
R matrices/vectors/lists, not external pointers or other
non-serializable values.
Active bindings
num_coresCurrent number of cores in the global budget.
Methods
Public methods
-
Design$unavailable_inference_classes_due_to_missing_packages() -
Design$incompatible_inference_classes_due_to_design_structure()
Design$is_blocking_design()
Check whether this design currently has blocking structure.
The base implementation returns FALSE. Designs that compose
BlockingStructure override this method with the structural check.
Usage
Design$is_blocking_design()
Returns
FALSE for designs without BlockingStructure.
Design$is_matching_design()
Check whether this design currently has matching structure.
The base implementation returns FALSE. Designs that compose
MatchingStructure override this method with the structural check.
Usage
Design$is_matching_design()
Returns
FALSE for designs without MatchingStructure.
Design$is_a_kk_matching_capable()
Characterization: is this a KK matching-on-the-fly-capable
design (sequential KK or its fixed binary-match equivalent)? Default
FALSE; overridden to TRUE on
DesignSeqOneByOneKK14 and DesignFixedBinaryMatch.
Usage
Design$is_a_kk_matching_capable()
Design$is_a_cluster_capable()
Characterization: is this a cluster-structured design?
Default FALSE; overridden to TRUE on
DesignFixedCluster and DesignFixedBlockedCluster.
Usage
Design$is_a_cluster_capable()
Design$is_a_bernoulli_capable()
Characterization: is this a Bernoulli-randomized design?
Default FALSE; overridden to TRUE on
DesignSeqOneByOneBernoulli and DesignFixedBernoulli.
Usage
Design$is_a_bernoulli_capable()
Design$new()
Initialize an experimental design
Usage
Design$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = FALSE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., ordinal_levels = NULL, seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TProbability of treatment assignment.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size (if fixed).
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates when building the model matrix for inference. One of:
"impute"(default)Missing values are filled in using random-forest imputation (
missRanger, falling back tomissForeston failure). The response vector is included as an auxiliary predictor when available. This preserves all covariates and all subjects but introduces imputed values that influence inference."drop_column"Any covariate column that contains at least one missing value is dropped entirely from the model matrix before inference. No values are invented; the remaining complete columns are used as-is. This is conservative but transparent.
"error"An error is thrown as soon as any missing value is detected in the covariate matrix. Use this when you want to guarantee that inference runs on exactly the data you supplied, with no silent modification.
design_formulaA formula object used to create the design matrix from covariates. Default is
~ ..ordinal_levelsIf the response type is "ordinal", the labels for the levels.
seedInteger seed for reproducibility.
Returns
A new 'Design' object
Design$add_one_subject_response()
For CARA designs, add a single subject response.
Usage
Design$add_one_subject_response(t, y = NULL, y_L = NULL, y_R = NULL)
Arguments
tThe subject index.
yThe exact response value. Supply this XOR both
y_Landy_R– never together, never just one of the two.y_LFor a censored survival response, the lower bound of the event-time interval. Right-censored: the last known event-free time (pair with
y_R = Inf). Left-censored:0, which must be stated explicitly rather than defaulted. Interval-censored: the interval's lower bound. Storage accepts any well-formed left-/interval-censored value; whether a givenInferenceclass can actually consume it depends on that class (most survivalInferenceclasses still only accept exact/right-censored data and will reject construction with a clear error otherwise – see individual class docs).y_RFor a censored survival response, the upper bound of the event-time interval. Right-censored:
Inf. Left-/ interval-censored: the confirmed-by time / interval upper bound.
Design$add_all_subject_responses()
For non-CARA designs, add all subject responses.
Usage
Design$add_all_subject_responses(ys = NULL, y_Ls = NULL, y_Rs = NULL)
Arguments
ysThe exact responses as a numeric vector,
NAfor any subject whose response is censored (supplyy_Ls/y_Rsfor those instead).y_LsThe censored-response lower bounds,
NAfor any subject with an exact response inys. Right-censored: the last known event-free time (pair withy_Rs = Inf). Left-censored:0, stated explicitly. Interval-censored: the interval's lower bound. Storage accepts any well-formed left-/interval-censored value; whether a givenInferenceclass can actually consume it depends on that class (most survivalInferenceclasses still only accept exact/ right-censored data and will reject construction with a clear error otherwise – see individual class docs).y_RsThe censored-response upper bounds,
NAfor any subject with an exact response inys. Right-censored:Inf. Left-/interval-censored: the confirmed-by time / interval upper bound.
Design$overwrite_all_subject_assignments()
For analysis on already-completed experimental data
Usage
Design$overwrite_all_subject_assignments(w)
Arguments
wA {0,1} vector of subject assignments (1 = treated, 0 = control).
Design$is_fixed_sample_size()
Check if this design was initialized with a fixed sample size n
Usage
Design$is_fixed_sample_size()
Returns
TRUE if fixed.
Design$assert_all_subjects_arrived()
Asserts if all subjects arrived.
Usage
Design$assert_all_subjects_arrived()
Design$assert_all_responses_recorded()
Asserts if all responses are recorded.
Usage
Design$assert_all_responses_recorded()
Design$check_experiment_completed()
Checks if the experiment is completed.
Usage
Design$check_experiment_completed()
Returns
TRUE if experiment is complete, FALSE otherwise.
Design$assert_even_allocation()
Checks if the experiment has a 50-50 allocation.
Usage
Design$assert_even_allocation()
Design$assert_fixed_sample()
Checks if the experiment has a fixed sample size.
Usage
Design$assert_fixed_sample()
Design$any_censoring()
Checks if the experiment has any censored responses
Usage
Design$any_censoring()
Returns
TRUE if any censored.
Design$has_general_censoring()
Checks if the experiment has any left- or
interval-censored survival responses – i.e. any subject whose
y_R is finite (right-censored subjects have
y_R = Inf, which is excluded). Most survival
Inference classes cannot yet consume this shape of data
(see get_effective_time()/get_effective_dead());
this is the check Inference$initialize() uses to reject
construction cleanly for those classes.
Usage
Design$has_general_censoring()
Returns
TRUE if any subject is left- or interval-censored.
Design$get_t()
Get t
Usage
Design$get_t()
Returns
The current number of subjects.
Design$get_X_raw()
Get raw X information
Usage
Design$get_X_raw()
Returns
A data frame of subject data.
Design$get_X_imp()
Get imputed X information
Usage
Design$get_X_imp()
Returns
Same as Xraw except with imputations.
Design$get_X()
Get X matrix
Usage
Design$get_X()
Returns
A numeric matrix of subject data.
Design$get_y()
Get y
Usage
Design$get_y()
Returns
A numeric vector of subject responses.
Design$get_y_original()
Get y_original
Usage
Design$get_y_original()
Returns
A numeric vector of the original subject responses.
Design$get_w()
Get w
Usage
Design$get_w()
Returns
A {0,1} vector of subject assignments (1 = treated, 0 = control).
Design$draw_ws_according_to_design()
Draw treatment assignment vectors according to the design.
Usage
Design$draw_ws_according_to_design(r = 1L)
Arguments
rNumber of vectors to draw. Default is 1.
Returns
A matrix of size n x r with {0,1} entries (1 = treated, 0 = control).
Design$capabilities()
Returns the capabilities this design instance exposes
(see fix_design_hierarchy.md, "Capability Model").
Deliberately instance-level, not a class-registry read (fix_design_hierarchy.md,
TODO-28): is_blocking_design()/is_matching_design() depend on
real construction-time state (e.g. private$m/private$blocking_capable),
not just which components a class composes – DesignFixediBCRD
constructed with an unknown n, for instance, composes
BlockingStructure but is not blocking-capable for that particular
instance. A class-registry-only answer (this function briefly unioned in
get_effective_design_capabilities(), a purely class-level, component-
composition-based check) would silently report "blocking" for every instance
of such a class regardless of its actual construction state – confirmed as a
real, reproducible false positive during this TODO's implementation, not a
hypothetical. get_effective_design_capabilities()/
design_class_registry.R's direct_components still exist and are
correct – they're the right tool for a generator-only query with no
instance in hand (see design_class_generator_supports_batch_w_pregeneration()),
just not for this instance-level method.
Usage
Design$capabilities()
Returns
A character vector of capability names.
Design$supports()
Returns whether this design object supports a capability.
See capabilities().
Usage
Design$supports(capability)
Arguments
capabilityA capability name, e.g.
"blocking","matching", or"batch_w_pregeneration".
Returns
TRUE if the capability is present, FALSE otherwise.
Design$applicable_inference_class_names()
Returns the sorted character vector of concrete, exported
Inference class names legal for this design object under
default constructor arguments, derived purely from this design's own
normalized metadata (response type, KK-matching capability, blocking,
and both censoring axes) filtered through the registry's
compatibility predicates – the same normalization and predicate
logic InferenceSuite uses for
discovery (see normalize_inference_design_metadata() and
is_inference_class_compatible_with_design_metadata() in
inference_suite.R). No candidate class is constructed to
determine applicability, so this has no side effects and cannot be
influenced by a constructor failure or a missing optional package
(see unavailable_inference_classes_due_to_missing_packages()
for that case, reported separately). A class whose censoring
tolerance depends on non-default constructor arguments (e.g.
InferenceSurvivalCoxPHRegr only tolerates general censoring
with testing_type = "wald") is listed here when its
default configuration is compatible; a construction-time
error for an incompatible non-default argument combination remains
the documented behavior of that class's initialize().
Usage
Design$applicable_inference_class_names()
Returns
A sorted character vector of applicable Inference class
names.
Design$unavailable_inference_classes_due_to_missing_packages()
Companion to applicable_inference_class_names():
returns the subset of otherwise design-compatible Inference
classes that are excluded solely because a registered
required_packages entry is not installed, as a named list
(class name -> character vector of missing package names) – kept
separate from plain design incompatibility so callers can tell "not
applicable to this design" apart from "applicable, but an optional
dependency isn't installed."
Usage
Design$unavailable_inference_classes_due_to_missing_packages()
Returns
A named list, class name -> missing package names; empty list if none.
Design$incompatible_inference_classes_due_to_design_structure()
Companion to applicable_inference_class_names():
returns the subset of otherwise design-compatible Inference
classes that are excluded because they declared a
design_compatibility_reason predicate (a design-*structure*
requirement, e.g. even treatment allocation or equal block sizes,
beyond what response type/KK/blocking/censoring metadata alone can
express) and this design object fails it, as a named list (class
name -> one-line reason string) – kept separate from plain design
incompatibility and from a missing package for the same reason
unavailable_inference_classes_due_to_missing_packages() is
kept separate: so callers can tell exactly why a class is missing
from applicable_inference_class_names() instead of only
discovering it as a construction-time error.
Usage
Design$incompatible_inference_classes_due_to_design_structure()
Returns
A named list, class name -> reason string; empty list if none.
Design$randomization_family()
Returns this design object's registry-backed randomization
family (see fix_design_hierarchy.md, "Class Metadata"), e.g.
"kk14", "bernoulli", "rerandomization". Replaces
class-identity (inherits()/is()) dispatch at call sites that
need to distinguish design variants (see "Class-Identity Dispatch
Replacement"). Returns NA_character_ if the class is not registered
or is one of the unsplit/timing-root abstract bases.
Usage
Design$randomization_family()
Returns
A single character string (or NA_character_).
Design$supports_resampling()
Check if the design supports resampling at all – FALSE
only for the abstract timing-family bases themselves (DesignFixed,
DesignSeqOneByOne, and their custom-extension abstract bases)
instantiated directly; TRUE for every concrete subclass,
including ObservationalDesign. This is the general check
for resampling methods that never need the design's own randomization
mechanism – plain nonparametric bootstrap, Bayesian bootstrap,
m-out-of-n bootstrap, PRW subsampling – which only resample already-observed
units/rows and their fixed, observed assignment, so they remain valid and
available even for a design with no randomization mechanism at all (see
ObservationalDesign's class documentation: "resampling subjects with
their observed, fixed assignment does not require a known randomization
probability"). Contrast with supports_randomization_draw()/
supports_resampling_replay() below, which gate the narrower set of
methods that actually do need to invoke the design's mechanism (a plain
randomization test/CI, or a bootstrap randomization test that re-randomizes
resampled data) and are therefore FALSE for ObservationalDesign
specifically – see fix_design_hierarchy.md, "Observational Design
Migration" for the live bug that split fixes.
Usage
Design$supports_resampling()
Returns
TRUE if supported.
Design$supports_randomization_draw()
Check if this design can draw a fresh treatment assignment
from its own randomization mechanism – the eligibility condition for
permutation-style randomization tests/CIs (compute_rand_two_sided_pval()
and friends), which redraw w directly. FALSE for the abstract
timing-family bases themselves (same as supports_resampling()) and,
unlike supports_resampling(), also FALSE for
ObservationalDesign (no draw mechanism at all – w is supplied
by the user, so there is nothing to redraw); TRUE for every other
concrete subclass. See supports_resampling()'s documentation for why
this is a narrower, separate capability rather than reusing that one, and
"Observational Design Migration" for the live bug this fixes
(ObservationalDesign previously answered the old, unsplit
supports_resampling() TRUE, silently passing the
randomization-test eligibility assert before failing later and deeper,
inside draw_ws_raw()'s throwing stub).
Usage
Design$supports_randomization_draw()
Returns
TRUE if a fresh randomization draw is supported.
Design$supports_resampling_replay()
Check if this design's mechanism can be faithfully replayed
against resampled data – the eligibility condition specifically for the
bootstrap randomization test (BRT), which resamples units and then
re-randomizes each resample using the design's own mechanism (see
inference_all_abstract_rand_bootstrap.R's repeated
draw_ws_according_to_design() calls). Not the eligibility
condition for plain nonparametric/Bayesian/m-out-of-n/PRW-subsampling
bootstrap – those never redraw w at all (they resample already-observed
units and their fixed, observed assignment) and are gated by the broader
supports_resampling() instead, which stays TRUE for
ObservationalDesign. FALSE for the same abstract timing-family
bases as supports_randomization_draw() and for
ObservationalDesign (no randomization mechanism to replay); TRUE
for every other concrete subclass. See supports_randomization_draw()'s
documentation for why this is a separate capability rather than the same flag
reused.
Usage
Design$supports_resampling_replay()
Returns
TRUE if bootstrap-randomization-test-style replay is supported.
Design$prepare_for_resampling_replay()
Hook invoked by the bootstrap-randomization-test machinery
on a design object whose assignment mechanism is about to be replayed
against resampled data (once per replicate draw site, ahead of
draw_ws_according_to_design(1L)). The base implementation is a
no-op; designs whose replay is a full re-optimization
(DesignFixedOptimal) override it to switch to their
per-replicate solver profile (solver_args$brt_*). Idempotent.
Usage
Design$prepare_for_resampling_replay()
Returns
invisible(NULL).
Design$warm_all_subject_data_cache()
Warm the per-subject assignment-data cache, when this design uses covariates. This is an internal optimization hook for randomization inference; it keeps cache mutation inside the Design object instead of exposing its private environment to callers.
Usage
Design$warm_all_subject_data_cache()
Returns
TRUE invisibly when a cache warm-up was attempted, or
FALSE invisibly when the design does not use covariates.
Design$get_n()
Get n, the sample size
Usage
Design$get_n()
Returns
The number of subjects.
Design$get_y_L()
Get y_L
Usage
Design$get_y_L()
Returns
A numeric vector of censored-response lower bounds
(NA for exact-response subjects).
Design$get_y_R()
Get y_R
Usage
Design$get_y_R()
Returns
A numeric vector of censored-response upper bounds
(NA for exact-response subjects).
Design$get_effective_time()
Get the effective response time per subject: the exact
value y where recorded, or the lower bound y_L
for a censored subject. This reconstructs "the one informative
number" every response type other than left-/interval-censored
survival data has always had, for code that needs a single
numeric value per subject rather than the y/y_L/
y_R triple directly.
Usage
Design$get_effective_time()
Returns
A numeric vector, one value per subject.
Design$get_effective_dead()
Get the effective event indicator per subject: 1
for an exact response, 0 for a censored one. This
reconstructs today's dead semantics for right-censored
survival data (and is trivially all-1 for every other
response type, which never has censoring). It is only valid
for exact/right-censored data – a left- or interval-censored
subject also returns 0 here, which is not meaningful
right-censoring status, so callers must confirm (e.g. via
any_censoring() plus their own censoring-shape checks)
that no such rows are present before relying on this value.
Usage
Design$get_effective_dead()
Returns
An integer vector, one value per subject.
Design$get_prob_T()
Get probability of treatment
Usage
Design$get_prob_T()
Returns
The specified probability.
Design$get_response_type()
Get response type
Usage
Design$get_response_type()
Returns
The specified response type.
Design$get_response_type_original()
Get the original response type
Usage
Design$get_response_type_original()
Returns
The original specified response type.
Design$get_ordinal_levels()
Get ordinal levels
Usage
Design$get_ordinal_levels()
Returns
The levels of the ordinal response.
Design$get_original_ordinal_levels()
Get original ordinal levels
Usage
Design$get_original_ordinal_levels()
Returns
The labels for the levels of the original ordinal response.
Design$get_missingness_method()
Get the missingness method
Usage
Design$get_missingness_method()
Returns
The missingness handling method: "impute", "drop_column",
or "error".
Design$get_edi_version_created()
Get the EDI package version this object was created under.
Stamped once, at construction time, from
utils::packageVersion("EDI"); never re-stamped on
readRDS() reload, so it reflects the version that originally
built the object rather than whatever version is currently loaded.
Objects saved before this field existed self-initialize it to the
currently loaded version the first time it is read (there is
no way to recover the true original version for those objects),
rather than erroring on the missing field.
Usage
Design$get_edi_version_created()
Returns
A character string, e.g. "1.0.0".
Design$transform_y()
Transform the response vector y
Usage
Design$transform_y( transform_fun, transformed_response_type, ordinal_levels = NULL )
Arguments
transform_funA function that takes y_original and returns a new y.
transformed_response_typeThe response type of the transformed y.
ordinal_levelsIf the transformed response type is "ordinal", the labels for the levels.
Design$get_design_formula()
Get the model formula
Usage
Design$get_design_formula()
Returns
The model formula.
Design$duplicate()
Duplicate this design object
Usage
Design$duplicate(verbose = FALSE)
Arguments
verboseA flag for verbosity.
Returns
A new 'Design' object with the same data
Design$clone()
The objects of this class are cloneable with this method.
Usage
Design$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
# Design is abstract and cannot be instantiated directly; construct a
# concrete subclass instead, e.g.:
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
Internal base for user-defined sequential-design extensions
Description
DesignCustomSequential is intentionally not exported. Subclasses
implement assignment_rule() and return a scalar 0/1 assignment for the
current subject. EDI handles subject storage, responses, and redraws through
DesignSeqOneByOne.
Super classes
Design -> DesignSeqOneByOne -> DesignCustomSequential
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignCustomSequential$assignment_rule()
User-defined assignment rule.
Usage
DesignCustomSequential$assignment_rule()
Returns
A binary treatment assignment.
DesignCustomSequential$assign_wt()
Standard internal assignment entry point.
Usage
DesignCustomSequential$assign_wt()
Returns
A binary treatment assignment.
DesignCustomSequential$clone()
The objects of this class are cloneable with this method.
Usage
DesignCustomSequential$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
A Fixed Design
Description
An abstract R6 Class encapsulating the data and functionality for a fixed experimental design. This class takes care of whole-experiment randomization.
Super class
Design -> DesignFixed
Methods
Public methods
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixed$new()
Initialize a fixed experimental design
Usage
DesignFixed$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL, ... )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TProbability of treatment assignment.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
...Extra arguments passed to the
Designsuperclass.
Returns
A new 'DesignFixed' object
DesignFixed$assign_w_to_all_subjects()
Assign treatment to all subjects in the fixed experiment.
Usage
DesignFixed$assign_w_to_all_subjects(w_precomputed = NULL)
Arguments
w_precomputedOptional {0,1} numeric vector of length n. If supplied the allocation is used directly and
draw_ws_according_to_designis not called (avoids e.g. the Java round-trip forDesignFixedGreedy).
DesignFixed$add_all_subjects_to_experiment()
Add all subjects' covariates to a fixed design at once.
Usage
DesignFixed$add_all_subjects_to_experiment(X_all)
Arguments
X_allA data frame containing the full covariate matrix.
Returns
Invisibly returns the design object.
DesignFixed$add_all_subject_responses()
Add all subject responses for a fixed design.
Usage
DesignFixed$add_all_subject_responses(ys = NULL, y_Ls = NULL, y_Rs = NULL)
Arguments
ysThe exact responses as a numeric vector,
NAfor any subject whose response is censored (supplyy_Ls/y_Rsfor those instead).y_LsThe censored-response lower bounds,
NAfor any subject with an exact response inys. Right-censored: the last known event-free time (pair withy_Rs = Inf). Left-censored:0, stated explicitly. Interval-censored: the interval's lower bound. Storage accepts any well-formed left-/interval-censored value; whether a givenInferenceclass can actually consume it depends on that class (most survivalInferenceclasses still only accept exact/ right-censored data and will reject construction with a clear error otherwise – see individual class docs).y_RsThe censored-response upper bounds,
NAfor any subject with an exact response inys. Right-censored:Inf. Left-/interval-censored: the confirmed-by time / interval upper bound.
DesignFixed$overwrite_all_subject_assignments()
Overwrite all subject assignments for a fixed design.
Usage
DesignFixed$overwrite_all_subject_assignments(w)
Arguments
wA {0,1} vector of subject assignments (1 = treated, 0 = control).
DesignFixed$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixed$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
# DesignFixed is abstract and cannot be instantiated directly; construct a
# concrete subclass instead, e.g.:
des = DesignFixedBernoulli$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()
A Fixed-Sample-Size Bernoulli (Independent-Coin-Flip) Randomized Design
Description
A fixed-sample-size DesignFixed in which each subject's
treatment assignment w_i is drawn independently as
w_i \stackrel{iid}{\sim} \mathrm{Bernoulli}(p), i = 1, \dots, n, where
p is prob_T. This is the classical Bernoulli (independent-coin-flip)
randomized design: unlike
DesignFixediBCRD (complete randomization), which
fixes the number of treated subjects at exactly \mathrm{round}(np), a Bernoulli
design leaves the realized number of treated subjects
n_T = \sum_i w_i \sim \mathrm{Binomial}(n, p) random; the trade-off is
independence across subjects (useful for some asymptotic/martingale arguments) at the
cost of not guaranteeing exact balance, which can matter for small n or for
inference procedures (e.g. exact permutation tests over a fixed number of treated) that
assume a fixed n_T.
Draw mechanism. draw_ws_raw(r) delegates to
generate_permutations_bernoulli_cpp(), which fills an n \times r matrix of
independent \mathrm{Bernoulli}(p) draws (one column per requested replicate,
via a Mersenne Twister RNG seeded once per call from R's RNG state), so
r replicate allocation vectors are generated with a single C++ call rather than
r separate calls into R's own random-number generation.
assign_w_to_all_subjects() draws a single such allocation
(r = 1) and applies it to all subjects at once.
No exchange/balance search. Because subjects are treated independently, there
is no optimization step analogous to
DesignFixedGreedyDOptimal: covariates, if supplied, do
not influence the assignment probabilities or realized allocation at all.
Super classes
Design -> DesignFixed -> DesignFixedBernoulli
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedBernoulli$is_a_bernoulli_capable()
Characterization: this design draws each subject's treatment
assignment as an independent \mathrm{Bernoulli}(p) coin flip (see class
documentation), so it is Bernoulli-capable by construction.
Usage
DesignFixedBernoulli$is_a_bernoulli_capable()
Returns
Always TRUE for this class.
DesignFixedBernoulli$new()
Initialize a fixed Bernoulli (independent-coin-flip) experimental
design. Unlike DesignFixediBCRD, the
realized number of treated subjects is not fixed at n * prob_T; it is
random (\mathrm{Binomial}(n, prob\_T)) because each subject's assignment
is an independent coin flip (see class documentation).
Usage
DesignFixedBernoulli$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TPer-subject probability
pthat a given subject is assigned to treatment; need not be0.5(unlikeDesignFixedGreedyDOptimal).include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedBernoulli' object
DesignFixedBernoulli$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedBernoulli$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Neyman, J. (1923, transl. 1990). "On the Application of Probability Theory to Agricultural Experiments." Statistical Science, 5(4), 465-472, for the potential-outcomes framework under which Bernoulli and complete randomization are compared; see also randomized experiment for orientation on Bernoulli vs. complete (restricted) randomization.
Examples
des = DesignFixedBernoulli$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()
A Fixed, Non-Bipartite-Matched-Pair Design with Within-Pair Randomization
Description
A fixed-sample-size DesignFixed that (1) partitions the
n subjects into n/2 disjoint matched pairs by solving a non-bipartite
(optimal) pairwise-matching problem on a covariate distance matrix, minimizing the
total within-pair distance across all pairs, then (2) randomizes treatment
within each pair independently: for pair k, one of its two members is
assigned to treatment and the other to control with probability 1/2 each,
independently across pairs. This is the classical matched-pair randomized design (a
special case of blocking with block size 2), which guarantees exact covariate balance
within every pair (up to the matching algorithm's distance metric) while preserving
randomization-based inference validity, in contrast to
DesignFixedGreedyDOptimal, which balance the
covariates in aggregate via an optimality criterion but do not guarantee pairwise
closeness of any two subjects.
Matching algorithm. Pairing is computed once (lazily, on first call to
draw_ws_raw()/assign_w_to_all_subjects(), via
private$ensure_matching_structure_computed()) by
compute_binary_match_structure(), which forms an n \times n pairwise
distance matrix — squared Mahalanobis distance if mahal_match = TRUE (the
default; using the sample covariance of the covariates, ridge-regularized if singular),
or squared Euclidean distance otherwise — and solves the minimum-weight non-bipartite
perfect matching on that distance matrix via nonbimatch
(nbpMatching, a Suggests-only dependency; loading is deferred until matching is
actually needed, so pre-computed w vectors injected via m never require
it). For a single covariate (p = 1), pairing instead reduces to simply sorting
subjects by that covariate and pairing consecutive subjects, since the non-bipartite
matching problem is trivial in one dimension. The resulting pairing is cached in
private$bms/private$m for the lifetime of the design object (or until
explicitly reset via set_m()) and is not recomputed per draw.
Within-pair randomization. Given the fixed pairing, each replicate allocation
(see draw_binary_match_assignments_cpp()) independently flips, for every pair,
which of its two members is treated (an independent fair coin flip per pair per
replicate, using a splitmix64-seeded Mersenne Twister per replicate column for
reproducible parallel draws); this guarantees exactly n/2 treated subjects
overall (only prob_T = 0.5 is supported; the constructor errors otherwise).
Passing a pre-computed m to the constructor supplies the matched-pair structure
directly (each pair ID occurring in exactly 2 rows), bypassing the matching
computation entirely while keeping the same within-pair randomization.
No-covariate fallback. If no covariates are available at draw time
(private$m is NULL, e.g. matching hasn't run and no explicit m was
supplied), draw_ws_raw() falls back to an unmatched balanced complete
randomization (a uniformly random permutation of n/2 ones and n/2 zeros),
since there is no covariate information to match on.
Batch pregeneration. draw_binary_match_assignments_cpp()'s output is
trusted unvalidated – it guarantees exactly n x r valid \{0,1\} columns
with n/2 treated subjects per column by construction (see
fix_design_hierarchy.md, "AllocationMatrixValidation"). supports_batch_w_pregeneration() returns TRUE so that the
calling framework generates all replicate w vectors for a simulation cell in one
batch (amortizing the one-time nbpMatching matching cost across all replicates
of that cell) rather than recomputing the matching structure per replicate.
Super classes
Design -> DesignFixed -> DesignFixedBinaryMatch
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedBinaryMatch$is_a_kk_matching_capable()
Characterization: this design computes its matched-pair structure on the fly from covariates (see class documentation), so it is KK matching-capable by construction.
Usage
DesignFixedBinaryMatch$is_a_kk_matching_capable()
Returns
Always TRUE for this class.
DesignFixedBinaryMatch$supports_batch_w_pregeneration()
Returns TRUE so the calling framework pre-generates all
replicate w vectors for a simulation cell in one batch, paying the
one-time nbpMatching non-bipartite matching cost once per cell and
reusing the resulting pairing across all replicates, rather than
recomputing it per replicate.
Usage
DesignFixedBinaryMatch$supports_batch_w_pregeneration()
Returns
Always TRUE for this class.
DesignFixedBinaryMatch$new()
Initialize a binary (non-bipartite) matched-pair fixed
experimental design. Matching itself is deferred until the first draw (see
class documentation); this constructor only records configuration and,
if m is supplied, installs the explicit pairing immediately.
Usage
DesignFixedBinaryMatch$new( response_type, prob_T = 0.5, mahal_match = TRUE, include_is_missing_as_a_new_feature = TRUE, n = NULL, m = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values.
prob_TThe probability of the treatment assignment. Must be 0.5, since within-pair randomization only supports an even 1-treated/1-control split per pair.
mahal_matchMatch using squared Mahalanobis distance (accounting for covariate correlation/scale) if
TRUE(default), else squared Euclidean distance on the raw covariate matrix.include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
mOptional integer vector of explicit matched-pair identifiers, one per subject. If supplied, 'n' must also be supplied, 'length(m)' must equal 'n', all values must be positive, and each pair ID must occur exactly twice. This bypasses the package-computed matching step.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedBinaryMatch' object
DesignFixedBinaryMatch$assign_w_to_all_subjects()
Assign treatment to all subjects (see
DesignFixed$assign_w_to_all_subjects() for the
general contract). Before delegating, this override ensures the matched-pair
structure is computed (private$ensure_matching_structure_computed())
even when w_precomputed is supplied and
draw_ws_according_to_design() is therefore never called — downstream
code (e.g. blocked/matched-pair inference) still needs private$m to be
populated regardless of how w was obtained.
Usage
DesignFixedBinaryMatch$assign_w_to_all_subjects(w_precomputed = NULL)
Arguments
w_precomputedOptional {0,1} numeric vector of length
n. If supplied, it is used directly as the treatment allocation instead of drawing a fresh within-pair-randomized allocation.
DesignFixedBinaryMatch$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedBinaryMatch$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Greevy, R., Lu, B., Silber, J. H., and Rosenbaum, P. (2004). "Optimal multivariate matching before randomization." Biostatistics, 5(2), 263-275, doi:10.1093/biostatistics/5.2.263, for optimal non-bipartite matched-pair designs prior to randomization. See also matched pair and Mahalanobis distance for orientation.
Examples
des = DesignFixedBinaryMatch$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()
A Fixed, Blocked-and-Clustered Randomized Design
Description
A fixed-sample-size DesignFixed in which the unit of
randomization is the cluster, not the individual subject: within each block
(stratum, formed from strata_cols), whole clusters (identified by
cluster_col) are jointly randomized to treatment or control, so all subjects in
the same cluster always receive the same assignment. This is the design used when
individual-level randomization is infeasible or invalid (e.g. clusters are classrooms,
clinics, or households where within-cluster interference/spillover would violate
SUTVA under individual randomization), combined with blocking to improve precision by
comparing clusters only to other clusters in the same stratum.
Randomization mechanism. Blocking keys are computed per subject via
private$get_strata_keys() (shared with other
blocking-structure designs): categorical columns in
strata_cols are used as-is, continuous columns are discretized into
preferred_num_bins_for_continuous_covariate quantile-based bins, and multiple
strata_cols are combined into a single composite block key. Within each
resulting block, whole clusters (by cluster_col) are randomized to treatment
with probability prob_T via block_and_cluster_ra
(randomizr), which performs blocked-and-clustered complete random assignment:
within each block, clusters (not subjects) are permuted so that, subject to rounding,
the target proportion prob_T of clusters in that block is treated, and every
subject in a treated cluster receives w = 1. r independent replicate
allocation columns are generated via replicate() (one randomizr call per
replicate; there is no batch/vectorized draw path for this design, unlike
DesignFixedBinaryMatch).
Cluster-aware bootstrap. draw_bootstrap_indices() overrides the default
subject-level bootstrap to resample at the cluster level via
resample_group_rows_cpp(): with bootstrap_type = "within_blocks"
(the default when bootstrap_type is NULL), clusters are resampled with
replacement within each block, preserving the block structure; otherwise, whole
blocks (strata) are themselves resampled with replacement. This mirrors the standard
cluster-robust bootstrap principle that resampling must occur at the level of the
randomization unit (clusters), not individual subjects, to yield a valid
variance/interval estimate under cluster-correlated outcomes.
Super classes
Design -> DesignFixed -> DesignFixedBlockedCluster
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedBlockedCluster$is_a_cluster_capable()
Characterization: this design randomizes whole clusters (see class documentation), so it is cluster-structured by construction.
Usage
DesignFixedBlockedCluster$is_a_cluster_capable()
Returns
Always TRUE for this class.
DesignFixedBlockedCluster$new()
Initialize a blocked and cluster randomized fixed experimental design.
Usage
DesignFixedBlockedCluster$new( strata_cols, cluster_col, response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, preferred_num_bins_for_continuous_covariate = 2, num_bins_for_continuous_covariate = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
strata_colsA character vector of column names to use for stratification (blocks).
cluster_colThe column name in the data that identifies the cluster for each subject.
response_typeThe data type of response values.
prob_TThe target probability that a given cluster within a block is assigned to treatment (subjects inherit their cluster's assignment).
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
preferred_num_bins_for_continuous_covariateThe number of quantile bins to use for continuous strata. Default is 2.
num_bins_for_continuous_covariateDeprecated alias for 'preferred_num_bins_for_continuous_covariate'.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedBlockedCluster' object
DesignFixedBlockedCluster$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedBlockedCluster$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Middleton, J. A., and Aronow, P. M. (2015). "Unbiased estimation of the average treatment effect in cluster-randomized experiments." Statistics, Politics and Policy, 6(1-2), 39-75, doi:10.1515/spp-2013-0002, for blocked/clustered randomized-assignment inference; see also the randomizr package vignette for the assignment-generation conventions this design relies on, and cluster randomized controlled trial for orientation.
Examples
des = DesignFixedBlockedCluster$new(n = 20, response_type = 'continuous',
strata_cols = 'x2', cluster_col = 'cl')
X = data.frame(x1 = rnorm(20), x2 = factor(rep(1:2, each = 10)), cl = factor(rep(1:10, each = 2)))
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()
A Fixed, Stratified-Block Randomized Design
Description
A fixed-sample-size DesignFixed that first partitions
subjects into blocks (strata) formed from covariates, then randomizes treatment
independently within each block at probability prob_T (via
block_ra when randomizr is installed, else an internal
generate_permutations_blocking_cpp() fallback). Blocking on a covariate removes
its between-block variation from the treatment-effect comparison (comparisons are
always within-block), improving precision relative to unblocked randomization whenever
the blocking covariate(s) are prognostic of the outcome, at the cost of requiring the
analysis to account for the blocking structure (e.g. via a block/stratum fixed effect
or a CMH-type test). This differs from
DesignFixedBlockedCluster, which
randomizes whole clusters of subjects together within each block rather than
subjects individually.
Block construction. Blocking keys are computed by
private$get_strata_keys() (shared across
blocking-structure designs): each column in
strata_cols contributes a categorical key (continuous columns are discretized
into preferred_num_bins_for_continuous_covariate quantile bins), and multiple
columns are combined into one composite block key per subject; if strata_cols
is NULL, all available covariate columns are used. B_target caps the
number of resulting blocks by greedily adding strata_cols in order only while
the running block count stays at or below the target (earlier columns take priority);
exact_num_blocks = TRUE instead hard-fails if the greedy construction does not
land on exactly B_target blocks. equal_block_sizes = TRUE (the default)
additionally requires every block to have the same subject count, checked once at
construction (via n %% B_target) if n and B_target are both
already known, and again once covariates arrive; some downstream inference classes
(InferenceIncidCMH, InferenceIncidExtendedRobins) require equal block
sizes unconditionally, regardless of this flag. An explicit m (one block ID per
subject) bypasses covariate-derived block construction entirely.
Within-block randomization and bootstrap. Within each block, treatment is
assigned independently via block_ra's complete random
assignment (subject to rounding, prob_T of each block's subjects are treated);
the internal C++ fallback (generate_permutations_blocking_cpp()) is used only
if randomizr is not installed. draw_bootstrap_indices() resamples
within each block by default (bootstrap_type = "within_blocks" or
NULL, via stratified_bootstrap_indices_cpp()), or resamples whole blocks
with replacement otherwise (via resample_group_rows_cpp()) — mirroring the
block structure in the resampling scheme, analogous to the cluster-level bootstrap in
DesignFixedBlockedCluster.
Super classes
Design -> DesignFixed -> DesignFixedBlocking
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedBlocking$new()
Initialize a fixed stratified-block randomized experimental
design. Block construction and validation follow the rules described in
the class documentation; see the parameter descriptions below for the
greedy B_target/exact_num_blocks/equal_block_sizes
contract.
Usage
DesignFixedBlocking$new( strata_cols = NULL, response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, preferred_num_bins_for_continuous_covariate = 2, B_target = NULL, exact_num_blocks = FALSE, equal_block_sizes = TRUE, m = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
strata_colsA character vector of column names to use for stratification. If 'NULL' (the default), all available covariate columns are used.
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TProbability of treatment assignment.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
preferred_num_bins_for_continuous_covariateThe number of quantile bins to use for continuous strata. Default is 2.
B_targetThe target number of blocks. Columns from 'strata_cols' are added greedily in order, each column being included only if it does not push the total number of unique blocks beyond this target. For categorical covariates their natural levels are used; for continuous covariates 'preferred_num_bins_for_continuous_covariate' quantile bins are used. Earlier columns are always preferred over later ones. When 'n' is known at construction time, the default is the largest divisor of 'n' that is at most 'floor(sqrt(n))' (so the default always satisfies 'equal_block_sizes = TRUE'); if 'n' is not yet known, it is resolved to 'floor(sqrt(n))' when subjects are added. Set 'B_target = NULL' to use all columns unconditionally. An explicitly supplied 'B_target' that does not divide 'n' still errors immediately when 'equal_block_sizes = TRUE'. Set 'exact_num_blocks = TRUE' to hard fail if the final key construction does not produce exactly 'B_target' blocks.
exact_num_blocksWhether to require the greedy key construction to produce exactly 'B_target' blocks. Default 'FALSE'.
equal_block_sizesWhether to require all blocks to have the same number of subjects. Default 'TRUE'. When 'TRUE' and both 'n' and 'B_target' are known at construction time, an error is raised immediately if 'n' is not divisible by 'B_target'. A second check fires when subjects are added: if the covariate-based strata produce unequal block counts the design errors at that point. Set to 'FALSE' to allow unequal blocks (note that 'InferenceIncidCMH' and 'InferenceIncidExtendedRobins' still require equal block sizes regardless).
mOptional integer vector of explicit block identifiers, one per subject. If supplied, 'n' must also be supplied and 'length(m)' must equal 'n'. The constructor then records this blocking structure immediately via 'set_m()', bypassing covariate-derived strata construction.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedBlocking' object
DesignFixedBlocking$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedBlocking$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd, for the original rationale for blocking in randomized experiments; Cochran, W. G., and Cox, G. M. (1957). Experimental Designs (2nd ed.), Wiley, for stratified (randomized block) design theory. See also randomized block design for orientation.
Examples
des = DesignFixedBlocking$new(n = 20, response_type = 'continuous',
strata_cols = 'x2', equal_block_sizes = FALSE)
X = data.frame(x1 = rnorm(20), x2 = factor(rep(1:2, 10)))
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()
A Fixed, Unblocked Cluster Randomized Design
Description
A fixed-sample-size DesignFixed in which whole clusters
of subjects (identified by cluster_col), rather than individual subjects, are
the unit of randomization: every subject in a given cluster always receives the same
treatment assignment. This is the unblocked analog of
DesignFixedBlockedCluster — there is no
stratification step here, so clusters are randomized to treatment as a single pool
rather than within strata. Cluster-level randomization is required whenever
individual-level randomization would create within-cluster interference/spillover
that violates SUTVA (e.g. clusters are classrooms, clinics, villages, or households),
at the cost of an effective sample size driven by the number of clusters, not
subjects, and a corresponding need for cluster-aware inference.
Randomization mechanism. draw_ws_raw(r) extracts each subject's
cluster ID from cluster_col (erroring if any are missing) and calls
cluster_ra (randomizr) once per replicate, which
performs complete random assignment at the cluster level: subject to rounding,
prob_T of clusters are assigned to treatment, and all subjects sharing a
cluster inherit that cluster's assignment. r independent replicate columns are
generated via replicate() (one randomizr call per replicate).
Cluster-aware bootstrap. draw_bootstrap_indices() overrides the
default subject-level bootstrap to resample whole clusters with replacement (via
resample_group_rows_cpp()) rather than individual rows, since outcomes are
correlated within a cluster (shared assignment plus, typically, shared context) and
the exchangeable resampling unit for a valid bootstrap variance/interval estimate is
therefore the cluster, not the subject.
Super classes
Design -> DesignFixed -> DesignFixedCluster
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedCluster$is_a_cluster_capable()
Characterization: this design randomizes whole clusters (see class documentation), so it is cluster-structured by construction.
Usage
DesignFixedCluster$is_a_cluster_capable()
Returns
Always TRUE for this class.
DesignFixedCluster$new()
Initialize a cluster randomized fixed experimental design (no
blocking/stratification; see
DesignFixedBlockedCluster if
stratification is also needed).
Usage
DesignFixedCluster$new( cluster_col, response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
cluster_colThe column name in the data that identifies the cluster for each subject.
response_typeThe data type of response values.
prob_TThe target probability that a given cluster is assigned to treatment (subjects inherit their cluster's assignment).
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedCluster' object
DesignFixedCluster$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedCluster$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Middleton, J. A., and Aronow, P. M. (2015). "Unbiased estimation of the
average treatment effect in cluster-randomized experiments." Statistics,
Politics and Policy, 6(1-2), 39-75, doi:10.1515/spp-2013-0002. See also
cluster
randomized controlled trial for orientation, and
DesignFixedBlockedCluster for the
blocked variant of this design.
Examples
des = DesignFixedCluster$new(n = 20, response_type = 'continuous', cluster_col = 'cl')
X = data.frame(x = rnorm(20), cl = factor(rep(1:5, each = 4)))
des$add_all_subjects_to_experiment(X)
des$assign_w_to_all_subjects()
Internal base for user-defined fixed-design extensions
Description
DesignFixedCustom is intentionally not exported. Subclasses implement
draw_assignments(r = 1) and return an n x r 0/1 assignment
matrix. EDI handles subject storage, responses, and validation through
DesignFixed.
Super classes
Design -> DesignFixed -> DesignFixedCustom
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedCustom$draw_assignments()
Draw assignments from the custom design.
Usage
DesignFixedCustom$draw_assignments(r = 1)
Arguments
rNumber of assignment vectors to draw.
Returns
An n x r matrix of 0/1 assignments.
DesignFixedCustom$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedCustom$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
A Fixed, Balanced Two-Arm Factorial Design
Description
A fixed-sample-size DesignFixed for factorial
treatment structures: subjects are assigned to one of the cells of a factorial
combination of one or more named factors (e.g. list(treatment = 2),
or, once multi-arm support lands, list(drug = 2, dose = 2) for a
2 \times 2 design), with assignment counts balanced as evenly as possible
across cells within each replicate draw.
Currently restricted to exactly two total factor-level combinations (i.e.
two arms), e.g. a single two-level factor — the product of levels across all
factors entries must equal exactly 2; the constructor errors otherwise. In this
two-arm regime, the design reduces to a balanced complete randomization between cell 1
(w = 0) and cell 2 (w = 1), and w follows the same {0,1}
internal / {-1,+1} public convention as every other Design subclass, so
DesignFixedFactorial inherits assign_w_to_all_subjects(),
draw_ws_according_to_design(), and get_w() unmodified from
DesignFixed/Design and works
unmodified with every Inference class; only
draw_ws_raw() (the low-level allocation-vector generator) and
get_w_factorial() (an additional factor-level accessor, see below) are
specific to this class. Support for more than two combinations (true multi-factor,
multi-arm designs) is tracked separately — see
package_metadata/new_feature_plans/multi_arm_designs.md.
Allocation generation. draw_ws_raw(r) builds a base allocation vector
by repeating the sequence of cell indices 0:(num_combinations - 1) out to
length n (so cells are as close to equally represented as possible, off by at
most one subject when n is not a multiple of the number of cells), then
independently permutes ( sample) that base vector once per
replicate column to produce r balanced-but-randomized allocations.
Super classes
Design -> DesignFixed -> DesignFixedFactorial
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedFactorial$new()
Initialize a factorial fixed experimental design. The product
of levels implied by factors must currently equal exactly 2 (see
class documentation); any other total raises an error.
Usage
DesignFixedFactorial$new( factors, response_type, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
factorsA list where names are factor names and values are number of levels (e.g. list(treatment = 2)). The product of levels across all factors must currently equal exactly 2 (two-arm only).
response_typeThe data type of response values.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedFactorial' object
DesignFixedFactorial$get_w_factorial()
Decode each subject's scalar cell index (private$w, in
0:(num_combinations - 1)) back into its per-factor level
assignments, using the same expand.grid enumeration
(private$combinations) established at construction. This is the
inverse of the encoding draw_ws_raw() produces, and is the only way
to recover individual factor levels once support for more than two total
combinations lands, since get_w() (inherited, see class
documentation) only ever returns the scalar 0/1 (or -1/+1) cell index.
Usage
DesignFixedFactorial$get_w_factorial()
Returns
A data frame with n rows and one column per entry of
factors, giving each subject's level (an integer in
1:levels) for that factor; NULL if treatment has not yet
been assigned to all subjects (i.e. private$w is empty or contains
NA).
DesignFixedFactorial$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedFactorial$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
des = DesignFixedFactorial$new(n = 12, response_type = 'continuous', factors = list(treatment = 2))
des$add_all_subjects_to_experiment(data.frame(x=1:12))
des$assign_w_to_all_subjects()
A Fixed, Covariate-Balanced Design via Greedy Pairwise-Swap Search
Description
A fixed-sample-size DesignFixed that searches, among
balanced (n/2-treated) allocations, for one that directly minimizes a covariate
imbalance criterion f(d), d = M(2w - 1), via a native C++ (RcppEigen +
OpenMP) greedy pairwise-swap search (greedy_design_search_cpp()). Unlike
DesignFixedGreedyDOptimal, which optimizes a
model-based information-matrix criterion (D-/A-optimality) implied by an assumed
linear model, this design optimizes a direct covariate-distance criterion between the
treated and control group means/covariance, with no linear-model assumption: the two
supported objectives are
-
"mahal_dist"(default):d = L^{-1} X^\top (2w-1) / nwhere\Sigma = LL^\topis the Cholesky factor of the covariate covariance matrix, andf(d) = \lVert d \rVert_2^2— the squared Mahalanobis distance between the treated and control covariate means, accounting for covariate correlation/scale. Falls back to"abs_sum_diff"if the covariate covariance matrix is (numerically) singular, since the Cholesky factorization then fails. -
"abs_sum_diff":d = X_{\mathrm{std}}^\top (2w-1) / n(covariates standardized to unit variance, no correlation adjustment) andf(d) = \lVert d \rVert_1— the sum of absolute standardized mean differences across covariates.
Search algorithm. Starting from a random balanced (Fisher-Yates) allocation,
the search runs in one of two modes selected by n_iter: Inf (default)
runs exhaustive best-improvement search — each round scans every
(treated, control) pair, applies the single globally best improving swap, and repeats
until no swap improves f(d), guaranteeing convergence to a strict local optimum;
a positive integer instead runs exactly that many stochastic steps, each
picking a uniformly random (treated, control) pair and accepting the swap only if it
improves f(d) (with patience-based early stopping). r independent design
searches (one per requested replicate) run in parallel via OpenMP, each with its own
std::mt19937 generator. This search's
randomization is reproducible via the constructor's seed argument: per-thread
RNGs are seeded from R's own RNG state (GetRNGstate()/unif_rand())
before the parallel region begins, so private$maybe_set_seed() does govern the
resulting allocation, independent of the number of OpenMP threads used.
Constraints and fallbacks. Only exactly balanced allocation
(prob_T = 0.5, n even) is supported; the constructor errors otherwise. If
no covariates are available, the search degenerates to pure balanced Fisher-Yates
randomization (no swap search, since there is nothing to balance on).
greedy_design_search_cpp()'s output is trusted unvalidated – it guarantees
exactly n x r valid \{0,1\} columns with n/2 treated subjects per
column by construction (see fix_design_hierarchy.md, "AllocationMatrixValidation").
Super classes
Design -> DesignFixed -> DesignFixedGreedy
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedGreedy$new()
Initialize a greedy pairwise-swap search fixed experimental
design. Only prob_T = 0.5 is supported (see class documentation).
Usage
DesignFixedGreedy$new( response_type, prob_T = 0.5, objective = "mahal_dist", n_iter = Inf, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values.
prob_TThe probability of the treatment assignment. Must be 0.5.
objectiveThe covariate-imbalance objective to minimize: either
"mahal_dist"(default, squared Mahalanobis distance between treated and control covariate means) or"abs_sum_diff"(sum of absolute standardized mean differences); see class documentation for the exact criteria.n_iterNumber of swap iterations.
Inf(default) uses exhaustive best-improvement search guaranteed to reach a strict local optimum. A positive integer runs that many stochastic random-pair iterations with patience-based early stopping.include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility. This design's search is reproducible via
seed(see class documentation).
Returns
A new 'DesignFixedGreedy' object
DesignFixedGreedy$supports_batch_w_pregeneration()
Returns TRUE so the calling framework pre-generates all
replicate w vectors for a simulation cell in a single batched call
to greedy_design_search_cpp() (which parallelizes the r
independent searches over OpenMP threads internally), rather than issuing
r separate single-replicate C++ calls.
Usage
DesignFixedGreedy$supports_batch_w_pregeneration()
Returns
Always TRUE for this class.
DesignFixedGreedy$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedGreedy$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Krieger, A. M., Azriel, D., and Kapelner, A. (2019). "Nearly random designs with greatly improved balance." Biometrika, 106(3), 695-701, doi:10.1093/biomet/asz026, for the greedy-swap balance-optimization approach this class implements. See also Mahalanobis distance for orientation on the default objective.
Examples
des = DesignFixedGreedy$new(n = 10, response_type = 'continuous')
A Fixed, Model-Based Optimal Design via Greedy Pairwise-Exchange Search
Description
A fixed-sample-size DesignFixed that searches, among
allocations with exactly n_T = \mathrm{round}(n \cdot \mathrm{prob}_T) treated
subjects, for allocations optimizing a model-based information-matrix
criterion implied by the linear model y = \beta_T w + Z_0 \gamma + \epsilon
with Z_0 = [1\ X], via a native C++ greedy pairwise-exchange (Fedorov/
DETMAX-style) local search. This class is the merger of the former
DesignFixedDOptimal and DesignFixedAOptimal classes; the criterion is
selected by the objective and interest constructor arguments.
The optimality-criterion family and its argument mapping. Write
M(w) = [w\ Z_0]^\top [w\ Z_0] for the information (moment) matrix,
P = Z_0 (Z_0^\top Z_0)^{-1} Z_0^\top, and
s(w) = n_T - w^\top P w (the treatment-coefficient information given the
covariate block). The classical criteria map to constructor arguments as
follows:
D_M– full-matrix determinant optimality (maximize\det M(w), i.e.|M|)objective = "D", interest = "all". By the Schur-complement identity\det M(w) = \det(Z_0^\top Z_0) \cdot s(w)(withw^\top w = n_Tfixed andZ_0not depending onw), the covariate block factors out, soD_Mreduces to maximizings(w).D_s– subset determinant optimality (minimize\det(K^\top M(w)^{-1} K)for a coordinate-selectionK: the treatment coefficient plus a chosen covariate subset)the default
objective = "D", interest = "treatment"isD_swith the interest set = {treatment};interest = ~ x1 + x2(a one-sided formula) orinterest = c("x1", "x2")(model-matrix column names) selects treatment + those covariates. Becausewenters only the treatment row/column ofM(w), every suchD_scriterion factorizes as\det(V_{SS})/s(w)with\det(V_{SS})constant inw– so all determinant-type settings (D_Mand everyD_s) select identical allocations and share the same search kernel. For the single treatment contrast,D_s-, c-, and per-parameter A-optimality coincide as well, which is whyobjective = "A", interest = "treatment"is silently equivalent to the default (allowed by design; no message is emitted).D_A– general contrast optimality (minimize\det(A^\top M(w)^{-1} A)for an arbitrary contrast matrixA)interest = <contrast matrix>– arrives with Stage 2 of the merge plan (the generalized-criterion kernel) and currently raises an informative error, as do interest sets excluding the treatment coefficient.D_B(andA_B) – Bayesian optimality (criteria computed on the posterior informationM(w) + R)prior_precision =a scalar\tauor a matrixR_0, combined with eitherobjective; see the Bayesian section below for exactly which coefficients a scalar\taupenalizes.- A – trace optimality (minimize
\mathrm{tr}(K^\top M(w)^{-1} K)) -
objective = "A"withinterest = "all"(all parameters: objective(w^\top H w + 1)/s(w),H = Z_0 (Z_0^\top Z_0)^{-2} Z_0^\top), or withinterest =formula/names (A_s: same kernel with the subset-restrictedH_S; see below). Unlike the determinant family, trace criteria over different interest sets generally select different allocations.
Bayesian variants. Supplying prior_precision replaces
(Z_0^\top Z_0)^{-1} with the ridge-regularized (Z_0^\top Z_0 + R_0)^{-1}
in the construction of P (and H), yielding Bayesian
D_B/A_B-optimality. A scalar \tau penalizes the
covariate coefficients only – the treatment coefficient and the intercept
are unpenalized (R_0 = \tau \cdot \mathrm{diag}(0, 1, \ldots, 1) over
Z_0's columns) – and, when standardize_covariates = TRUE (the
default), the covariates are centered and scaled to unit variance first so
\tau is interpretable per standardized coefficient. A full matrix
prior_precision is used as R_0 verbatim (dimensions
(1+p) \times (1+p) over [\mathrm{intercept}, \mathrm{covariates}] of
the design's model matrix; standardize_covariates is ignored).
Search algorithm. For each of the r requested allocations
independently: start from a uniformly random balanced-count allocation (a BCRD
draw with exactly n_T treated), then repeatedly apply the single best
improving treated/control pairwise exchange until no exchange improves the
criterion (a strict local optimum). The returned allocations therefore form a
restricted-randomization distribution over locally optimal allocations, which is
what makes randomization inference possible for this design. The search is
reproducible via the constructor's seed argument: the C++ kernels seed a
local generator from R's own RNG stream, so a fixed seed yields identical
draws (this corrects the former classes' documentation, which predated the RNG
migration).
Covariate-subset criteria (D_s/A_s) via interest = a formula or
names. interest also accepts a one-sided formula (e.g.
~ x1 + x2) or a character vector of model-matrix column names, meaning
the treatment coefficient plus the named covariate coefficients (the
treatment is always in the interest set; the intercept never is). Both reduce to
the existing kernels with no new machinery: under objective = "D",
because w only enters the treatment row/column of M(w), the
subset determinant factorizes as
\det(K^\top M(w)^{-1} K) = \det(V_{SS}) / s(w) with
\det(V_{SS}) constant in w – so subset-D selects allocations
identical to the default treatment-focused criterion (allowed silently,
like objective = "A", interest = "treatment"); under
objective = "A", the subset trace criterion is
(w^\top H_S w + 1) / s(w) with
H_S = (Z_0 V S)(Z_0 V S)^\top built from the selected columns – the
same trace kernel with a subset-restricted H. Formula terms are
expanded against the design's model matrix, so factor covariates must be
referred to by their expanded model-matrix column names. Note that restricting
the design's model matrix itself via design_formula also changes the
default covariate set downstream inference adjusts for
(Inference$initialize() inherits the design's formula), whereas
interest affects the allocation criterion only. General contrast
matrices (D_A), and interest sets excluding the treatment coefficient, arrive
with Stage 2 of the merge plan (the generalized-criterion kernel; see
package_metadata/finished_features/fix_design_hierarchy.md).
Constraints and fallbacks. prob_T may be any value in (0, 1)
for which 1 \le \mathrm{round}(n \cdot \mathrm{prob}_T) \le n - 1. If no
covariates are available, the search degenerates to pure random allocation with
n_T treated (there is no criterion to optimize).
Super classes
Design -> DesignFixed -> DesignFixedGreedyDOptimal
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedGreedyDOptimal$new()
Initialize a model-based optimal-search fixed experimental
design. Covariates, if any, are supplied later via
add_all_subjects_to_experiment(); the optimality search itself does
not run until assign_w_to_all_subjects() (or
draw_ws_according_to_design()) is called.
Usage
DesignFixedGreedyDOptimal$new( response_type, prob_T = 0.5, objective = "D", interest = "treatment", prior_precision = NULL, standardize_covariates = TRUE, n_iter = Inf, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal". Determines only which downstream inference/response machinery this design is paired with; it does not affect the optimality search itself.
prob_TProbability of treatment assignment, in
(0, 1). The search fixes the treated count at\mathrm{round}(n \cdot prob_T).objectiveThe optimality criterion:
"D"(default, determinant) or"A"(trace). See the class documentation for the exact criteria and for whyobjective = "A"withinterest = "treatment"is equivalent to the default.interestWhich parameters the criterion targets:
"treatment"(default),"all", a one-sided formula (e.g.~ x1 + x2), a single formula string (e.g."x1 * x2 + x7", promoted to~ x1 * x2 + x7), or a character vector of model-matrix column names – all but"all"meaning the treatment coefficient plus the named covariate coefficients (D_s/A_s; see class documentation, including why subset-D selects the same allocations as the default). Formula terms (including interactions likex1:x2) must correspond to columns of the design's model matrix: to target an interaction coefficient, the interaction must be indesign_formulatoo – you cannot be "interested in" a coefficient the working model does not contain. Contrast matrices (general D_A) arrive with Stage 2 of the merge plan and currently raise an error.prior_precisionNULL(default, non-Bayesian), a single positive scalar\tau(ridge prior precision on the covariate coefficients only; treatment and intercept unpenalized), or a full(1+p) \times (1+p)symmetric prior-precision matrixR_0over[\mathrm{intercept}, \mathrm{covariates}].standardize_covariatesIf
TRUE(default) andprior_precisionis a scalar, covariates are centered and scaled to unit variance before the penalized criterion matrices are built. Ignored otherwise.n_iterNumber of exchange iterations.
Inf(default) runs the exhaustive best-improvement search to a strict local optimum. Finite values (the stochastic swap mode shared withDesignFixedGreedy) arrive with the Stage-2 shared search engine and currently raise an error.include_is_missing_as_a_new_featureFlag for missingness indicators.
nSample size (if fixed).
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility. Unlike the former
DesignFixedDOptimal/DesignFixedAOptimaldocumentation claimed, the optimality search is reproducible viaseed(see class documentation).
Returns
A new 'DesignFixedGreedyDOptimal' object
DesignFixedGreedyDOptimal$get_objective()
The optimality criterion this design was constructed with.
Usage
DesignFixedGreedyDOptimal$get_objective()
Returns
"D" or "A".
DesignFixedGreedyDOptimal$get_interest()
The parameter-interest setting this design was constructed with.
Usage
DesignFixedGreedyDOptimal$get_interest()
Returns
"treatment" or "all".
DesignFixedGreedyDOptimal$get_prior_precision()
The Bayesian prior precision this design was constructed with.
Usage
DesignFixedGreedyDOptimal$get_prior_precision()
Returns
NULL, a positive scalar, or a symmetric matrix.
DesignFixedGreedyDOptimal$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedGreedyDOptimal$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Atkinson, A. C., Donev, A. N., and Tobias, R. D. (2007). Optimum Experimental Designs, with SAS. Oxford University Press, for the D-/A-optimality criteria and exchange algorithms for constrained design search. See also optimal design for orientation.
Examples
des = DesignFixedGreedyDOptimal$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()
A Fixed, Matched-Pair Design with Greedy Which-Member-Treated Optimization
Description
A fixed-sample-size DesignFixed that combines
DesignFixedBinaryMatch's non-bipartite
matched-pair structure with DesignFixedGreedy's
greedy imbalance-minimization search, restricted so every move respects the pairing:
subjects are first paired by covariate closeness (as in
DesignFixedBinaryMatch), guaranteeing exactly one treated and one control
subject per pair; then, rather than assigning within-pair treatment status by a coin
flip, the greedy search (greedy_design_search_cpp(), pair-constrained mode)
chooses which member of each pair is treated so as to directly minimize the
same aggregate covariate-imbalance objective as
DesignFixedGreedy (squared Mahalanobis distance
or sum of absolute standardized mean differences between the treated and control
group means) across the whole sample, not just within each pair. This targets
both close within-pair matches (from the matching step) and low aggregate covariate
imbalance (from the greedy refinement) simultaneously — a strictly more constrained
search than plain DesignFixedGreedy, since only the 2^{n/2}
which-member-treated assignments consistent with the fixed pairing are considered,
rather than all \binom{n}{n/2} balanced allocations.
Search algorithm. Pairing is computed by
compute_binary_match_structure() exactly as in
DesignFixedBinaryMatch (Mahalanobis or
Euclidean distance per objective), lazily on first draw and cached in
private$bms. Given the pairing, each replicate search initializes with a
random coin flip per pair (which member starts treated), then in exhaustive mode
(n_iter = Inf, default) repeatedly finds and applies the single pair-flip that
most decreases the imbalance objective, stopping at a strict local optimum (or runs
exactly n_iter random-pair stochastic flip-if-improving steps otherwise, with
patience-based early stopping) — the same two search modes as
DesignFixedGreedy, but with moves restricted to
"flip which side of a given pair is treated" rather than "swap any treated/control
pair of subjects." Random initialization and swap selection are seeded from R's own
RNG (via greedy_design_search_cpp()'s per-thread seeding), so seed does
govern reproducibility here.
Pair-preserving bootstrap. draw_bootstrap_indices() resamples whole
matched pairs (via draw_matching_bootstrap_sample_cpp()) rather than
individual subjects, since the greedy search only ever flips which member of a pair
is treated (never crosses pairs), so w always has exactly one treated subject
per pair — the pair, not the subject, is the exchangeable resampling unit.
Constraints. Only prob_T = 0.5 is supported (the constructor errors
otherwise), and n must be divisible by 4 (draw_ws_raw() errors
otherwise); n/2 matched pairs are formed regardless of parity, but the
additional divisible-by-4 requirement is enforced by this class specifically (unlike
DesignFixedBinaryMatch, which only requires
even n).
Super classes
Design -> DesignFixed -> DesignFixedMatchingGreedyPairSwitching
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedMatchingGreedyPairSwitching$new()
Initialize a fixed design that performs binary matching followed
by greedy which-member-treated optimization (see class documentation).
Only prob_T = 0.5 is supported, and n must be divisible by 4.
Usage
DesignFixedMatchingGreedyPairSwitching$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n, verbose = FALSE, objective = "mahal_dist", n_iter = Inf, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values.
prob_TThe probability of treatment assignment. Must be
0.5.include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size; must be divisible by 4.
verboseA flag for verbosity.
objectiveThe covariate-imbalance objective to minimize when choosing which pair member is treated: either
"mahal_dist"(default, squared Mahalanobis distance between treated/control means, also used as the matching distance) or"abs_sum_diff"(sum of absolute standardized mean differences); see class documentation for the exact criteria.n_iterNumber of swap iterations.
Inf(default) uses exhaustive best-improvement search guaranteed to reach a strict local optimum. A positive integer runs that many stochastic random-pair iterations with patience-based early stopping.missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new DesignFixedMatchingGreedyPairSwitching object.
DesignFixedMatchingGreedyPairSwitching$supports_batch_w_pregeneration()
Returns TRUE so the calling framework pre-generates all
replicate w vectors for a simulation cell in one batched call to
greedy_design_search_cpp(), paying the one-time nbpMatching
pairing cost once per cell (cached in private$bms) and reusing it
across all replicates and the OpenMP-parallelized greedy searches, rather
than recomputing the pairing per replicate.
Usage
DesignFixedMatchingGreedyPairSwitching$supports_batch_w_pregeneration()
Returns
Always TRUE for this class.
DesignFixedMatchingGreedyPairSwitching$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedMatchingGreedyPairSwitching$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Krieger, A. M., Azriel, D., and Kapelner, A. (2019). "Nearly random designs with greatly improved balance." Biometrika, 106(3), 695-701, doi:10.1093/biomet/asz026; Greevy, R., Lu, B., Silber, J. H., and Rosenbaum, P. (2004). "Optimal multivariate matching before randomization." Biostatistics, 5(2), 263-275, doi:10.1093/biostatistics/5.2.263, for the matched-pair design this class refines.
Examples
des = DesignFixedMatchingGreedyPairSwitching$new(n = 10, response_type = 'continuous')
A Fixed, Deterministic Single-Allocation Optimal Design
Description
A fixed-sample-size DesignFixed that computes
exactly one allocation w^* – the minimizer of a chosen
covariate-imbalance or information objective over all allocations with
n_T = \mathrm{round}(n \cdot \mathrm{prob}_T) treated subjects – by
numerical optimization, rather than drawing from a restricted-randomization
distribution the way DesignFixedGreedy/
DesignFixedGreedyDOptimal do.
The objective family and its argument mapping. Write
Z_0 = [1\ X], P = Z_0 (Z_0^\top Z_0)^{-1} Z_0^\top, and
s(w) = n_T - w^\top P w. The objectives map to constructor arguments
and solved forms as follows:
"D"(determinant /D_M,D_s) and"A"withinterest = "treatment"maximize
s(w), solved as the binary quadratic program\min_w w^\top P w– identical criteria andinterest/prior_precisionsemantics toDesignFixedGreedyDOptimal(the same shared construction machinery is used, so the two classes optimize literally the same matrices)."A"withinterest = "all"or a covariate subset (A,A_s)minimize
(w^\top H w + 1)/s(w)with the sibling class'sH/H_S, solved exactly via Dinkelbach's algorithm (Dinkelbach 1967) over product-linearized MILP subproblems.- Bayesian
D_B/A_B prior_precision =a scalar\tau(covariates only; intercept and treatment unpenalized) or a full matrixR_0, replacing(Z_0^\top Z_0)^{-1}with the ridge-regularized inverse – identical to the sibling class."mahal_dist"/"abs_sum_diff"-
DesignFixedGreedy's covariate-imbalance criteria, definitionally identical (column-centeredX, the same standardization and singular-covariance fallback), translated to exactly solvable forms: the Mahalanobis criterion is the pure quadraticw^\top Q wwithQ = 4 X \Sigma^{-1} X^\top / n^2, and the absolute-sum criterion is an l1 objective solved by the standard linear MILP. "custom"a user-compiled black box under the
user_compiled_fns.hcalling convention (double f(const Eigen::MatrixXd& X, const Eigen::VectorXd& w), minimized), supplied viacustom_objective; always solved by the annealing path (no structure to linearize).
Solvers and certificates. solver = "auto" (default) uses
the exact "ompr" MILP path (optimum_certificate = "global",
a certified global optimum) wherever tractable – always for
"abs_sum_diff" (pure linear MILP); up to
solver_args$linearization_max_n (default 20, set by a GLPK
benchmark: the product linearization adds n(n-1)/2 auxiliaries and
branch-and-bound cost climbs steeply past n \approx 20) for the
quadratic and Dinkelbach criteria – and the native simulated-annealing
solver beyond it, or always for "custom". The annealing solver is a
formal method, not a heuristic: Metropolis acceptance over
treated/control swaps with a configurable cooling schedule, for which
Hajek (1988) proves convergence in probability to the global optimum
under a slow-enough (logarithmic) schedule; the practical geometric
schedule used by default is asymptotically motivated only, so its
certificate is always "annealing_converged", never
"global". solver = "ompr"/"annealing" force a path.
Commercial backends extend the exact range via
solver_args$roi_solver; see that parameter's wiring guides.
Inference. There is no usable randomization distribution
conditional on the observed data (given X there is exactly one
w^* up to the mirror coin), so permutation-style randomization
tests/CIs are unavailable (supports_randomization_draw() is
FALSE); the bootstrap randomization test IS available (the
mechanism – "optimize this dataset" – is replayed on each resampled
covariate matrix), as is all model-based and plain-resampling inference.
The mirror coin. At prob_T = 0.5, whenever the mirror
1 - w^* is a verified co-optimum (checked numerically by evaluating
the objective, never by a symmetry table), a fair seeded coin picks between
w^* and its mirror (mirror_coin = TRUE, the default). This
restores exact treated/control label symmetry – and estimator
unbiasedness – at zero cost to balance. A mirror that evaluates strictly
better than the solver's answer raises an error (it would be a solver bug).
Super classes
Design -> DesignFixed -> DesignFixedOptimal
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$randomization_family()Design$supports()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedOptimal$new()
Initialize a deterministic single-allocation optimal fixed
experimental design. The optimization itself does not run until
assign_w_to_all_subjects() (or
draw_ws_according_to_design(r = 1)) is called.
Usage
DesignFixedOptimal$new( response_type, prob_T = 0.5, objective = "D", interest = "treatment", prior_precision = NULL, standardize_covariates = TRUE, custom_objective = NULL, solver = "auto", solver_args = list(), mirror_coin = TRUE, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values.
prob_TProbability of treatment assignment, in
(0, 1); the solve fixes the treated count at\mathrm{round}(n \cdot prob_T).objective"D"(default),"A","mahal_dist","abs_sum_diff", or"custom"; see the class documentation.interestFor
objective = "D"/"A"only:"treatment"(default),"all", a one-sided formula, a formula string, or model-matrix column names – identical semantics toDesignFixedGreedyDOptimal.prior_precisionFor
objective = "D"/"A"only:NULL(default), a positive scalar\tau, or a symmetric prior-precision matrixR_0– identical semantics toDesignFixedGreedyDOptimal.standardize_covariatesIf
TRUE(default) andprior_precisionis a scalar, covariates are standardized before the penalized criterion matrices are built. ("mahal_dist"/"abs_sum_diff"standardize internally per their definitions regardless.)custom_objectiveRequired iff
objective = "custom"(and forbidden otherwise): anRcppXPtrUtils::cppXPtr()external pointer or a C++ source string under theuser_compiled_fns.hcalling convention (double f(const Eigen::MatrixXd& X, const Eigen::VectorXd& w);Xis the design's model matrix,wa candidate 0/1 allocation; the returned value is minimized). A source string is compiled through the samecppXPtrmechanism and retained so parallel workers can recompile locally. A plain R function is not accepted, and cannot be: the annealing solver evaluates the objective once per candidate swap – typically thousands of times per chain, timesn_chains, and again per BRT replicate – and an R-call round-trip on every one of those evaluations is orders of magnitude too slow to be usable, not merely slower. Since"custom"is always solved by annealing (never the MILP path), there is no lower-frequency code path where an R closure would be merely inconvenient; the restriction is a hard performance requirement of this objective's only execution path. See the class examples for a workedcppXPtr()construction. Save/ reload note: a rawcppXPtr()object does not survivesaveRDS()/readRDS()(seeDesign's "Saving and loading" section); pass a C++ source string instead if this design needs to be reloadable.solver"auto"(default),"ompr", or"annealing".solver_argsA named list of solver tuning arguments. Supported:
roi_solver("glpk"/"gurobi"/"cplex"– a closed set; arbitrary ROI plugin names are rejected),linearization_max_n,max_dinkelbach_iter,n_chains,max_iter,initial_temp,cooling_rate, and (consumed by the BRT replicate path)brt_max_iter,brt_n_chains,brt_solver.Wiring up Gurobi (
roi_solver = "gurobi"): (1) obtain a Gurobi license (free academic licenses are available) and install the Gurobi Optimizer itself – this sets upGUROBI_HOMEand the license file, entirely outside this package's control; (2) install Gurobi's own R package, which is not on CRAN – it ships inside the Gurobi installation:R CMD INSTALL "$GUROBI_HOME/R/gurobi_<version>_R_<Rmajor.minor>.tar.gz"(exact filename depends on your Gurobi version and platform); (3)install.packages("ROI.plugin.gurobi")from CRAN; (4) verify"gurobi" %in% ROI::ROI_registered_solvers()after loading the plugin; (5) passsolver_args = list(roi_solver = "gurobi").Wiring up CPLEX (
roi_solver = "cplex"): (1) obtain an IBM CPLEX license (free academic licenses are available) and install IBM ILOG CPLEX Optimization Studio; (2) installRcplex(CRAN) – unlike the Gurobi bridge, it compiles from source against your local CPLEX SDK and must be pointed at your CPLEX version's include/lib directories at install time; followRcplex's own INSTALL instructions for your CPLEX version rather than a fixed command, since the flags change across CPLEX releases; (3)install.packages("ROI.plugin.cplex")from CRAN; (4) verify"cplex" %in% ROI::ROI_registered_solvers(); (5) passsolver_args = list(roi_solver = "cplex").ROI.plugin.gurobi/ROI.plugin.cplex/Rcplexare deliberately never listed in this package'sSuggests: declaring them would misrepresent the dependency as somethinginstall.packages()could satisfy, when the vendor installation/license underneath cannot be. Availability is checked lazily at solve time; if the plugin loads but the solve fails, the likely cause is a missing vendor installation or license.mirror_coinIf
TRUE(default), flip a fair seeded coin betweenw^*and a verified co-optimal mirror1 - w^*after every solve (only possible atprob_T = 0.5); see the class documentation.include_is_missing_as_a_new_featureFlag for missingness indicators.
nSample size (if fixed).
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility (consumed by the annealing solver and the mirror coin; MILP solves are deterministic up to the seeded label flip).
Returns
A new 'DesignFixedOptimal' object
DesignFixedOptimal$get_objective()
The objective this design was constructed with.
Usage
DesignFixedOptimal$get_objective()
Returns
One of "D", "A", "mahal_dist",
"abs_sum_diff", "custom".
DesignFixedOptimal$get_interest()
The parameter-interest setting ("D"/"A" only).
Usage
DesignFixedOptimal$get_interest()
Returns
The interest construction argument.
DesignFixedOptimal$get_prior_precision()
The Bayesian prior precision this design was constructed with.
Usage
DesignFixedOptimal$get_prior_precision()
Returns
NULL, a positive scalar, or a symmetric matrix.
DesignFixedOptimal$get_solver()
The solver setting this design was constructed with.
Usage
DesignFixedOptimal$get_solver()
Returns
"auto", "ompr", or "annealing".
DesignFixedOptimal$get_mirror_coin()
The mirror-coin setting this design was constructed with.
Usage
DesignFixedOptimal$get_mirror_coin()
Returns
TRUE or FALSE.
DesignFixedOptimal$get_optimization_diagnostics()
Diagnostics cached by the most recent solve: the solver
used, optimum_certificate ("global" for exact
"ompr" solves, "annealing_converged" otherwise), the
achieved objective value, mirror-coin outcome
(mirror_feasible/mirror_tied/mirror_flipped),
elapsed time, and the solver's own detail fields.
Usage
DesignFixedOptimal$get_optimization_diagnostics()
Returns
A named list, or NULL if no solve has run yet.
DesignFixedOptimal$supports_randomization_draw()
Characterization: FALSE – given the observed data
there is exactly one w^* (up to the vacuous 2-atom mirror pair),
so there is no randomization distribution to draw from and
permutation-style randomization tests/CIs are unavailable. The
bootstrap randomization test remains available via
supports_resampling_replay() (the deterministic mechanism is
replayed on each resample).
Usage
DesignFixedOptimal$supports_randomization_draw()
Returns
Always FALSE for this class.
DesignFixedOptimal$prepare_for_resampling_replay()
BRT replicate-mode switch (called by the
bootstrap-randomization-test machinery ahead of each replayed draw;
see Design$prepare_for_resampling_replay()). Subsequent solves
use the per-replicate solver profile: solver_args$brt_solver
(default "annealing" with the reduced
brt_max_iter/brt_n_chains schedule – replicate
assignments need to be faithful applications of the mechanism, not
individually re-verified to the observed solve's convergence
standard; "ompr" buys exact per-replicate solves at the
user's expense). Idempotent; the mirror coin still applies per
replicate (the BRT replays the coin-inclusive mechanism).
Usage
DesignFixedOptimal$prepare_for_resampling_replay()
Returns
invisible(NULL).
DesignFixedOptimal$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedOptimal$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Dinkelbach, W. (1967). On nonlinear fractional programming. Management Science 13(7):492-498, for the exact A-optimality reduction. Hajek, B. (1988). Cooling schedules for optimal annealing. Mathematics of Operations Research 13(2):311-329, for the annealing solver's formal convergence property. Atkinson, A. C., Donev, A. N., and Tobias, R. D. (2007). Optimum Experimental Designs, with SAS. Oxford University Press, for the D-/A-optimality criteria.
Examples
# objective = "mahal_dist" (the default MILP path) needs a MILP solver --
# ompr + ompr.roi + a ROI plugin (ROI.plugin.glpk by default), all Suggests,
# not installed automatically with EDI. Confirmed 2026-09-25: ungated, this
# example errored outright on R CMD check's --no-suggests leg (and would do
# the same for any real user without the optional MILP-solver stack), not
# just a CI-specific issue.
if (requireNamespace("ompr", quietly = TRUE) &&
requireNamespace("ompr.roi", quietly = TRUE) &&
requireNamespace("ROI.plugin.glpk", quietly = TRUE)) {
des = DesignFixedOptimal$new(n = 14, response_type = 'continuous', objective = "mahal_dist")
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(14)))
des$assign_w_to_all_subjects()
des$get_optimization_diagnostics()
}
# A custom compiled objective (the user_compiled_fns.h calling convention),
# built with RcppXPtrUtils::cppXPtr() -- here, squared imbalance of the
# centered covariate sums. Compiling it needs a C++ toolchain and takes
# several seconds:
if (requireNamespace("RcppXPtrUtils", quietly = TRUE)) {
fobj = RcppXPtrUtils::cppXPtr(
"double f(const Eigen::MatrixXd& X, const Eigen::VectorXd& w) {
Eigen::RowVectorXd mu = X.colwise().mean();
Eigen::MatrixXd Xc = X.rowwise() - mu;
Eigen::VectorXd s = 2.0 * w - Eigen::VectorXd::Ones(X.rows());
return (Xc.transpose() * s).squaredNorm();
}", depends = "RcppEigen")
des2 = DesignFixedOptimal$new(n = 14, response_type = 'continuous',
objective = "custom", custom_objective = fobj)
des2$add_all_subjects_to_experiment(data.frame(x1 = rnorm(14)))
des2$assign_w_to_all_subjects()
}
A Fixed, Covariate-Homogeneous-Block Randomized Design
Description
A fixed-sample-size DesignFixed that first partitions
subjects into B approximately equal-sized, covariate-homogeneous blocks by
(approximately or exactly) minimizing total within-block pairwise covariate distance
\sum_{k=1}^{B} \sum_{i, j \in \text{block } k, i < j} D(x_i, x_j), then
randomizes treatment independently within each resulting block at probability
prob_T (via block_ra, the same mechanism as
DesignFixedBlocking). Unlike
DesignFixedBlocking, which forms blocks from user-specified column-wise
strata (categorical levels / quantile bins), here blocks are formed directly from a
multivariate distance/clustering criterion over all covariates jointly, so this
design generalizes matched-pair designs
(DesignFixedBinaryMatch) from block size 2
to arbitrary block size B. When B is omitted and n is known
at initialization, the default is floor(sqrt(n)), truncated below at 1 (the
usual heuristic block count for balancing within-block homogeneity against
within-block sample size).
Block-formation algorithms. method selects among three block
construction strategies with different exactness/scalability trade-offs (see the
new() argument documentation for details): "K-way" (default, balanced
k-means-style anticlustering via anticlust, fast and empirically close to
optimal), "greedy" (nearest-neighbor greedy matching via blockTools,
fast even for large n), and "ompr" (an exact mixed-integer program
solved with GLPK via ompr/ompr.roi, globally optimal but scaling as
O(n^2 B) in the number of decision variables — practical only for small
n). Block membership is computed lazily (on first draw, via
get_or_compute_block_ids()) and cached in private$block_ids for reuse
across replicates.
Distance specification ("ompr" only). dist selects the
pairwise distance D(x_i, x_j) the exact solver minimizes:
"euclidean", "sum_abs_diff" (sum of absolute coordinate differences),
"mahal" (default, Mahalanobis distance accounting for covariate
correlation/scale), or a user-supplied distance function. The "K-way" and
"greedy" methods use their own respective packages' built-in distance
conventions and do not consult dist.
No-covariate fallback. If the covariate matrix has zero columns, block
membership is assigned by simple round-robin (rep(seq_len(B), length.out = n))
rather than by any of the three clustering algorithms, since there is no covariate
information to cluster on.
Solver backend (method = "ompr" only). roi_solver selects the
MILP backend ompr.roi dispatches to, a closed set
c("glpk", "gurobi", "cplex") (default "glpk") – validated against this
set, not passed through to ROI::ROI_registered_solvers() unchecked, so a typo
or an unsupported solver name fails fast with a clear message rather than an opaque
ompr error three layers down. Gurobi and CPLEX are supported because they extend
which block-formation problems stay practical, not just which are expressible: GLPK's
branch-and-bound is single-threaded with no commercial-grade presolve/cutting-plane
machinery, while Gurobi/CPLEX are typically an order of magnitude faster on the same
MILP and solve in parallel, meaningfully extending the n for which the exact
"ompr" method stays practical. Scope is deliberately closed to these two for
now – not because other ROI plugins (CBC, SYMPHONY, ...) wouldn't work mechanically
(the dispatch is generic), but because Gurobi and CPLEX are the two most widely used
commercial solvers and the only ones worth a maintained step-by-step guide at this
point; extending the closed set is a small, low-risk addition later if a real need for
a third backend appears, not a reason to leave the set open-ended now.
Wiring up Gurobi:
Obtain a Gurobi license (a free academic license is available from Gurobi for non-commercial use) and install the Gurobi Optimizer itself. This sets up
GUROBI_HOMEand the license file (gurobi.lic, discoverable via theGRB_LICENSE_FILEenvironment variable or Gurobi's default search path) – entirely outside this package's control or dependency graph.Install Gurobi's own R package. Not available via CRAN – it ships inside the Gurobi installation itself:
R CMD INSTALL "$GUROBI_HOME/R/gurobi_<version>_R_<Rmajor.minor>.tar.gz"(exact filename/path depends on your Gurobi version and platform; see theR/subdirectory of your Gurobi install). This is the vendor interfaceROI.plugin.gurobiwraps – required even though it's not what you call directly.Install the ROI bridge package from CRAN:
install.packages("ROI.plugin.gurobi"). This package is on CRAN (it only depends on ROI + thegurobiR package from step 2 being present at load time) and is the only new artifact this class's own dependency graph ever touches.Verify: after
library(ROI.plugin.gurobi),"gurobi" %in% ROI::ROI_registered_solvers()should beTRUE.Pass
roi_solver = "gurobi"to the constructor.
Wiring up CPLEX:
Obtain an IBM CPLEX license (a free academic license is available from IBM) and install IBM ILOG CPLEX Optimization Studio.
Install Rcplex (CRAN), CPLEX's R interface. Unlike
ROI.plugin.gurobi, Rcplex is a source package that compiles against your local CPLEX installation – it needs to be pointed at your CPLEX SDK's include/lib directories at install time (typically viaconfigure.argstoinstall.packages(), naming your CPLEX version'scplex/include/cplex/lib/<platform>paths). The exact flag names and paths are CPLEX-version- and platform-specific – follow Rcplex's ownINSTALL/README instructions for your installed CPLEX version rather than a fixed command copied from here, since this changes across CPLEX releases.Install the ROI bridge package from CRAN:
install.packages("ROI.plugin.cplex")(depends on Rcplex from step 2 being present and working).Verify: after
library(ROI.plugin.cplex),"cplex" %in% ROI::ROI_registered_solvers()should beTRUE.Pass
roi_solver = "cplex"to the constructor.
Dependency-graph consequence: ROI.plugin.gurobi/ROI.plugin.cplex
(and Rcplex) are never added to Suggests – they're free/CRAN-
available themselves, but declaring them would misrepresent the dependency as
something install.packages("EDI", dependencies = TRUE) could satisfy, when the
vendor package/license underneath cannot be. The lazy-check pattern already used for
ompr/ompr.roi/ROI.plugin.glpk extends naturally: check
requireNamespace("ROI.plugin.gurobi"/"ROI.plugin.cplex") at solve time (not at
package load or class-definition time) for whichever roi_solver was requested,
and error informatively – naming the missing package and, if that's present but the
solve still fails, noting the likely cause is a missing vendor license/installation,
not something this class can diagnose further.
Bootstrap. draw_bootstrap_indices() resamples within blocks by
default (bootstrap_type = "within_blocks" or NULL, via
stratified_bootstrap_indices_cpp()) or resamples whole blocks otherwise (via
resample_group_rows_cpp()), mirroring the block structure in the resampling
scheme, as in DesignFixedBlocking.
Super classes
Design -> DesignFixed -> DesignFixedOptimalBlocks
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedOptimalBlocks$supports_batch_w_pregeneration()
Returns TRUE so the calling framework pre-generates
all replicate w vectors for a simulation cell in one batch,
paying the one-time block-formation cost (K-way anticlustering,
greedy matching, or the exact ompr/GLPK solve) once per cell and
reusing the resulting block assignment across replicates, rather than
recomputing it per replicate.
Usage
DesignFixedOptimalBlocks$supports_batch_w_pregeneration()
Returns
Always TRUE for this class.
DesignFixedOptimalBlocks$new()
Initialize a fixed optimal-blocks design. Block formation
itself is deferred until the first draw (see class documentation);
this constructor only validates and records configuration, including
checking that B (or its floor(sqrt(n)) default) admits a
feasible block-size partition of n when n is already
known.
Usage
DesignFixedOptimalBlocks$new( B = NULL, method = "K-way", dist = "mahal", roi_solver = "glpk", response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
BNumber of blocks to form. If omitted and
nis supplied, defaults tofloor(sqrt(n)), with a minimum of 1.methodAlgorithm used to partition subjects into blocks.
"K-way"(default)Balanced k-means anticlustering via
anticlust::balanced_clustering. Requires the anticlust package. Produces well-spread blocks and is significantly faster than"greedy"(e.g., ~10x faster forn=200, p=10, B=10) while achieving a better within-block distance objective (e.g., ~4% lower)."greedy"Greedy nearest-neighbour matching via
blockTools::block. Requires the blockTools package. Fast even for largen."ompr"Exact mixed-integer programme solved with GLPK via ompr. Globally optimal but scales as
O(n^2 B)in variables and is only practical for smalln.
distDistance specification used only when
method = "ompr". Either a function or one of"euclidean","sum_abs_diff", or"mahal". Default is"mahal".roi_solverMILP backend used only when
method = "ompr". A closed setc("glpk", "gurobi", "cplex"), default"glpk". See the class documentation's "Solver backend" and wiring-guide sections for the Gurobi/CPLEX setup steps.response_typeThe response type for the design.
prob_TTreatment assignment probability within each block.
include_is_missing_as_a_new_featureWhether to include missingness indicators.
nPlanned sample size.
verboseWhether to print progress messages.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new DesignFixedOptimalBlocks object.
DesignFixedOptimalBlocks$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedOptimalBlocks$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Higgins, M. J., Sävje, F., and Sekhon, J. S. (2016). "Improving massive
experiments with threshold blocking." Proceedings of the National Academy of
Sciences, 113(27), 7369-7376, doi:10.1073/pnas.1510504113, for optimal/near-optimal
covariate-based blocking prior to randomization. See also
randomized block
design for orientation and
Mahalanobis distance for
the default "ompr" distance.
Examples
des = DesignFixedOptimalBlocks$new(n = 9, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x = rnorm(9)))
des$assign_w_to_all_subjects()
A Fixed Rerandomization Design (Rejection-Sampled on Covariate Balance)
Description
A fixed-sample-size DesignFixed implementing
rerandomization (Morgan and Rubin, 2012): candidate allocations w are drawn
from the design's base randomization law (balanced complete randomization when
prob_T = 0.5, i.i.d. \mathrm{Bernoulli}(prob\_T) draws otherwise — see
generate_one_rerandomized_w()) and only accepted if a covariate
imbalance criterion M(w) between the treated and control groups falls below a
threshold a (obj_val_cutoff), i.e. the accepted allocations are drawn
from the base randomization distribution truncated to \{w : M(w) \le a\}.
Unlike DesignFixedGreedy/
DesignFixedGreedyDOptimal, which search for a
single well-balanced allocation, rerandomization instead filters the design's
own randomization distribution, which is what makes it directly compatible with
Fisherian/randomization-based inference (the accepted allocations remain a
well-defined — if truncated — randomization distribution, valid for randomization
tests/CIs restricted to that truncated support).
Two mutually exclusive acceptance modes are supported (specifying both errors):
obj_val_cutoff accepts/rejects each draw against a fixed threshold a;
prop_acceptable instead draws r / prop\_acceptable candidates and keeps
the r with the lowest M(w) (an empirical top-quantile acceptance
region, equivalent in the large-draw limit to an implied cutoff at the
prop_acceptable quantile of M(w)'s distribution under the base
randomization law).
Objective M(w). objective = "mahal_dist" (default) uses the
squared Mahalanobis distance between treated and control covariate means,
M(w) = (\bar X_T - \bar X_C)^\top S^{-1} (\bar X_T - \bar X_C), where S is
the sample covariance of all covariates (ridge-regularized by 10^{-6}I if
|\det S| < 10^{-10}), computed once and cached in private$S_inv.
objective = "abs_sum_diff" instead uses the sum of absolute mean differences,
M(w) = \sum_j |\bar X_{T,j} - \bar X_{C,j}|, with no correlation adjustment.
Only these two objectives are supported; any other value errors at draw time.
Fast path (native C++). When prob_T = 0.5 and n is even,
candidate generation and filtering run via a parallel C++ rejection sampler
(rerandomization_search_cpp()), which internally works with a rescaled
objective f_{\mathrm{cpp}}: f_{\mathrm{cpp}} = M(w)/4 for
"mahal_dist" and f_{\mathrm{cpp}} = M(w)/2 (on a GED-standardized scale)
for "abs_sum_diff"; the user-facing obj_val_cutoff is converted to this
internal scale before being passed to C++, so the accepted-allocation semantics are
unaffected, but this rescaling is a backend implementation detail worth knowing when
comparing C++-path and pure-R-path acceptance rates for the same nominal cutoff. The
sampler draws up to max(r * 1000, 100000) candidates internally; if fewer than
r allocations are accepted within that budget (an overly tight cutoff), this
errors naming how many were actually found – loosen obj_val_cutoff or use
prop_acceptable instead. (Earlier versions silently recycled the accepted set
to pad out to r, duplicating some draws; fixed, since that meant some
"independent" replicates were literal duplicates of an accepted allocation.)
Seed reproducibility and multi-core parallelism. The C++ fast path's
rejection sampler is a genuine work-stealing search: with more than one core
(set_num_cores/a fork cluster/mirai daemons; the package default is a
single core), threads race via atomic operations for both which candidate draws to
try next and which output column an accepted draw claims, so which per-thread-seeded
RNG stream ends up producing a given replicate – and in what order – depends on
real-time OS scheduling, not just seed. With the default single core, draws
are exactly seed-reproducible; this is not guaranteed once more than one core
is in use. Contrast with DesignFixedGreedy/
DesignFixedBinaryMatch, whose C++ kernels
use static (not work-stealing) thread scheduling and remain seed-reproducible
regardless of core count.
prop_acceptable path. Uses
complete_randomization_forced_balanced_cpp() (balanced case) or
complete_randomization_imbalanced_cpp() (prob_T != 0.5) to draw
n_{\mathrm{draw}} = \mathrm{round}(r / prop\_acceptable) candidate allocations
in one batched call, computes M(w) for all of them via
compute_objective_vals_cpp(), and keeps the r with smallest M(w).
Pure-R fallback (unbalanced or odd n, obj_val_cutoff mode
only). Draws one candidate at a time via generate_one_rerandomized_w() in an
unbounded repeat loop that accepts the first candidate with
M(w) \le a. Unlike the C++ fast path, this fallback has no draw-count
safety limit: if obj_val_cutoff is set tight enough that acceptance
probability under the base randomization law is extremely small for this n/
covariate structure, this loop can run for a very long time (in principle
indefinitely) before finding an acceptable draw.
Super classes
Design -> DesignFixed -> DesignFixedRerandomization
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixedRerandomization$new()
Initialize a rerandomization fixed experimental design.
Exactly one of obj_val_cutoff/prop_acceptable may be
specified (or neither, which accepts every candidate, i.e. no filtering);
supplying both raises an error. See class documentation for the exact
acceptance semantics of each mode and the covariate-imbalance objective.
Usage
DesignFixedRerandomization$new( response_type, prob_T = 0.5, obj_val_cutoff = NULL, prop_acceptable = NULL, objective = "mahal_dist", include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values.
prob_TThe probability of the treatment assignment.
obj_val_cutoffThe maximum allowable objective value
a; a candidate allocation is accepted iffM(w) \le a. Cannot be specified together with prop_acceptable.prop_acceptableThe proportion of randomizations to accept (draws r/prop_acceptable total, returns r lowest). Cannot be specified together with obj_val_cutoff.
objectiveThe covariate-imbalance objective
M(w)to filter on: either"mahal_dist"(default, squared Mahalanobis distance) or"abs_sum_diff"(sum of absolute mean differences); see class documentation for the exact formulas.include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixedRerandomization' object
DesignFixedRerandomization$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixedRerandomization$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Morgan, K. L., and Rubin, D. B. (2012). "Rerandomization to improve covariate balance in experiments." The Annals of Statistics, 40(2), 1263-1282, doi:10.1214/12-AOS1008, for the rerandomization framework and its randomization-inference validity. See also Mahalanobis distance for the default objective.
Examples
des = DesignFixedRerandomization$new(n = 10, response_type = 'continuous')
A Fixed, Individually Balanced Completely Randomized Design (iBCRD)
Description
A fixed-sample-size DesignFixed implementing the
individually balanced complete randomized design (iBCRD): the number of treated
subjects is fixed at exactly n_T = \mathrm{round}(n \cdot prob\_T), and each
allocation w with exactly n_T ones is drawn uniformly at random from the
\binom{n}{n_T} possible such allocations (via a Fisher-Yates shuffle of a base
vector with n_T ones and n - n_T zeros). This is the classical "complete
randomization" reference design of randomization inference: unlike
DesignFixedBernoulli (independent per-subject
coin flips, random n_T), n_T is fixed here, which is what makes exact
permutation/randomization tests over the \binom{n}{n_T} allocations well-defined;
unlike DesignFixedGreedyDOptimal/
DesignFixedGreedy, no covariate information is
used to select among those allocations — every one of the \binom{n}{n_T}
allocations is equally likely.
Draw mechanism. draw_ws_raw(r) delegates to
generate_permutations_ibcrd_cpp(), which builds one base allocation vector
(n_T ones followed by n - n_T zeros) and independently
shuffles (Fisher-Yates via std::shuffle) a fresh copy of it
per replicate column, seeded from R's own RNG stream (so seed does govern
reproducibility here, unlike the A-/D-optimal exchange searches).
assign_w_to_all_subjects() draws one such allocation (r = 1) and applies
it to all subjects.
Single implicit block. The constructor sets private$m to a constant
vector of 1s (a single block containing every subject) once n is known, so that
shared blocking/matching machinery that expects a block-membership vector treats the
whole sample as one block by default.
Super classes
Design -> DesignFixed -> DesignFixediBCRD
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignFixediBCRD$new()
Initialize a fixed individually balanced completely randomized experimental design (see class documentation for the exact randomization law).
Usage
DesignFixediBCRD$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TTarget probability of treatment assignment; the realized number of treated subjects is fixed at
round(n * prob_T)for every draw (unlikeDesignFixedBernoulli, where it is random).include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignFixediBCRD' object
DesignFixediBCRD$clone()
The objects of this class are cloneable with this method.
Usage
DesignFixediBCRD$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Fisher, R. A. (1935). The Design of Experiments. Oliver and Boyd, for complete randomization as the canonical reference design of randomization inference. See also randomized experiment for orientation on complete vs. Bernoulli randomization.
Examples
des = DesignFixediBCRD$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects()
Sequential One-by-One Experimental Design
Description
Abstract R6 class encapsulating data and functionality for a sequential one-by- one experimental design.
Sample size and stopping
Subjects are assigned one at a time, but this class does not implement interim
outcome monitoring or an outcome-dependent stopping rule. For the usual
fixed-sample analysis, specify the target sample size n before enrollment
and stop after exactly n subjects; the caller is responsible for ending
enrollment at that point. With n = NULL, the class leaves the final sample
size unspecified and does not determine when enrollment ends. Inference methods
that assume a fixed sample size require the final size to be chosen independently
of accumulating outcomes.
Super class
Design -> DesignSeqOneByOne
Methods
Public methods
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOne$new()
Initialize a sequential one-by-one design.
Usage
DesignSeqOneByOne$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL, ... )
Arguments
response_typeThe data type of response values.
prob_TThe probability of the treatment assignment.
include_is_missing_as_a_new_featureIf missing data is present, include a dummy variable for it.
nThe prespecified target sample size for fixed-sample analysis. If
NULL, the final sample size is left to the caller; the class does not provide a stopping rule.verboseWhether to print progress messages.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
...Extra arguments passed to the
Designsuperclass.
DesignSeqOneByOne$add_one_subject()
Add subject-specific measurements for the next subject entrant.
Usage
DesignSeqOneByOne$add_one_subject(x_new, allow_new_cols = TRUE)
Arguments
x_newA data frame with one row representing the new subject's covariates.
allow_new_colsAllow new features in the new subject's covariates.
DesignSeqOneByOne$add_one_subject_to_experiment_and_assign()
Adds a subject and assigns treatment.
Usage
DesignSeqOneByOne$add_one_subject_to_experiment_and_assign(x_new)
Arguments
x_newA data frame with one row representing the new subject's covariates.
Returns
The treatment assignment as {0,1} (1 = treated, 0 = control).
DesignSeqOneByOne$assign_wt()
Assigns treatment to the current subject.
Usage
DesignSeqOneByOne$assign_wt()
Returns
The treatment assignment (0 or 1).
DesignSeqOneByOne$print_current_subject_assignment()
Prints the current subject's assignment.
Usage
DesignSeqOneByOne$print_current_subject_assignment()
DesignSeqOneByOne$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOne$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
# DesignSeqOneByOne is abstract and cannot be instantiated directly;
# construct a concrete subclass instead, e.g.:
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
Atkinson's (1982) Covariate-Adjusted Biased Coin Sequential Design
Description
A DesignSeqOneByOne that assigns each newly
arriving subject's treatment via Atkinson's (1982) D_A-optimum biased coin: a
coin whose treatment probability is skewed away from prob_T toward whichever
assignment would most improve the current design's efficiency for estimating the
treatment effect D_A-optimally, given the covariates observed so far. Compared
to a fixed-probability coin (e.g.
DesignSeqOneByOneBernoulli), this
improves covariate balance/estimation efficiency online without ever fully
determinizing the assignment (the coin is always strictly between 0 and 1, so
randomization-based inference remains valid), at the cost of requiring a numerically
well-conditioned design matrix to compute the bias.
Assignment rule. For subject t, let Z_{t-1} = [w_{1:t-1}, 1,
X_{1:t-1}] be the (treatment, intercept, covariates) design matrix accumulated from
the first t-1 subjects, and let M = (t-1)(Z_{t-1}^\top Z_{t-1})^{-1}.
Writing x_t for the new subject's covariate vector (with a leading 1 for the
intercept) and A = M_{[1, 2:]} \cdot x_t (the treatment row of M,
projected onto x_t), the treatment probability is
\pi_t = \frac{\big(M_{11}/A + 1\big)^2}{\big(M_{11}/A + 1\big)^2 + 1},
clamped to [0, 1], and subject t is assigned to treatment with
probability \pi_t. This is Atkinson's biased-coin formula for
D_A-optimal sequential design: the coin biases toward the assignment that
would most reduce the variance of the treatment-effect estimate under the linear
model implied by Z_t, converging toward more extreme (but never
fully deterministic) probabilities as the current covariate imbalance grows in
directions that matter for that estimate.
Fallback to a fair(-ish) coin. For the first
ncol(private$Xraw) + 3 subjects (too few observations for
Z_{t-1}^\top Z_{t-1} to be reliably invertible), and whenever the C++
computation encounters a non-invertible design matrix, a non-finite bias term, or any
other numerical failure (caught via tryCatch()), assignment falls back to an
unbiased \mathrm{Bernoulli}(prob\_T) draw instead of Atkinson's rule.
Reproducibility. The per-subject C++ draw (atkinson_assign_weight_cpp())
seeds its own generator from R's RNG stream per call, so seed governs
reproducibility of the resulting assignment sequence in the usual way.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneAtkinson
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneAtkinson$new()
Initialize an Atkinson (1982) biased-coin sequential experimental design (see class documentation for the assignment rule).
Usage
DesignSeqOneByOneAtkinson$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values.
prob_TThe nominal probability of treatment assignment; used as the fallback coin probability early in the trial and whenever Atkinson's rule cannot be computed (see class documentation), and as the reference probability the biased coin is skewed away from otherwise.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneAtkinson' object
DesignSeqOneByOneAtkinson$assign_wt()
Draw the next subject's treatment assignment via Atkinson's
(1982) D_A-optimum biased coin (see class documentation for the
exact probability formula), falling back to an unbiased
\mathrm{Bernoulli}(prob\_T) draw early in the trial or on numerical
failure of the biased-coin computation.
Usage
DesignSeqOneByOneAtkinson$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneAtkinson$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneAtkinson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Atkinson, A. C. (1982). "Optimum biased coin designs for sequential clinical trials with prognostic factors." Biometrika, 69(1), 61-67, doi:10.1093/biomet/69.1.61. See also randomized experiment for orientation on biased-coin sequential designs.
Examples
seq_des = DesignSeqOneByOneAtkinson$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
A Sequential Bernoulli (Independent-Coin-Flip) Randomized Design
Description
A DesignSeqOneByOne in which each arriving
subject's treatment assignment is drawn independently as
w_t \stackrel{iid}{\sim} \mathrm{Bernoulli}(prob\_T), with no dependence on
covariates or on prior assignments — the direct sequential-enrollment analog of
DesignFixedBernoulli. As in the fixed-sample
version, the realized number of treated subjects after t arrivals is random
(\mathrm{Binomial}(t, prob\_T)), in contrast to sequential designs that
actively balance assignment counts or covariates (e.g.
DesignSeqOneByOneAtkinson).
Nonparametric bootstrap
The ordinary row bootstrap for this design is supported under a prespecified
fixed sample size: choose n before enrollment and stop after exactly
n subjects, without using interim outcomes to decide when to stop.
EDI does not implement sequential monitoring or check that this stopping
condition was followed. In particular, n = NULL does not satisfy the
documented fixed-sample justification, even though the bootstrap method is
not blocked at runtime. The usual assumptions of iid subjects and potential
outcomes, and a regular estimator, also apply.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneBernoulli
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneBernoulli$is_a_bernoulli_capable()
Characterization: this design draws each subject's treatment
assignment as an independent \mathrm{Bernoulli}(prob\_T) coin flip
(see class documentation), so it is Bernoulli-capable by construction.
Usage
DesignSeqOneByOneBernoulli$is_a_bernoulli_capable()
Returns
Always TRUE for this class.
DesignSeqOneByOneBernoulli$new()
Initialize a Bernoulli (independent-coin-flip) sequential experimental design.
Usage
DesignSeqOneByOneBernoulli$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_typeThe data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival", "ordinal".
prob_TThe probability of the treatment assignment. This defaults to
0.5.include_is_missing_as_a_new_featureIf missing data is present in a variable, should we include another dummy variable for its missingness? The default is
TRUE.nThe prespecified sample size for fixed-sample inference. Default is
NULL; the nonparametric bootstrap justification above requires a fixedn.verboseA flag indicating whether messages should be displayed.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneBernoulli' object
DesignSeqOneByOneBernoulli$assign_wt()
Draw the next subject's treatment assignment as a single
independent \mathrm{Bernoulli}(prob\_T) coin flip (see class
documentation); does not consult covariates or prior assignments.
Usage
DesignSeqOneByOneBernoulli$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneBernoulli$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneBernoulli$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
Efron's (1971) Biased Coin Sequential Design
Description
A DesignSeqOneByOne implementing Efron's (1971)
biased coin: no covariates are used, only the running counts of treated
(n_T) and control (n_C) subjects assigned so far. If the counts are
currently equal, the next subject is assigned by a fair \mathrm{Bernoulli}(0.5)
coin; otherwise, the next subject is assigned to the currently
under-represented group with probability weighted_coin_prob
(> 0.5, e.g. the classical 2/3) and to the over-represented group with
probability 1 - weighted_coin_prob. This keeps the running treatment/control
counts close to balanced throughout enrollment (unlike
DesignSeqOneByOneBernoulli, whose
running counts can drift arbitrarily far from balanced) while remaining strictly
randomized at every step (the coin is always strictly between
1 - weighted_coin_prob and weighted_coin_prob, never fully
deterministic), unlike a purely deterministic alternating allocation. This is a
count-balancing design only — it does not use covariates at all, in contrast to
DesignSeqOneByOneAtkinson/
DesignSeqOneByOneKK21, which bias the coin
toward covariate balance rather than (or in addition to) count balance.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneEfron
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneEfron$new()
Initialize an Efron (1971) biased coin sequential experimental design (see class documentation for the exact assignment rule).
Usage
DesignSeqOneByOneEfron$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, weighted_coin_prob = 2/3, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TNominal probability of treatment assignment; used only as the fair-coin probability when the running treated/control counts are exactly equal (see
assign_wt()).include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
weighted_coin_probThe probability (
> 0.5) of assigning the next subject to whichever of treatment/control currently has fewer subjects, when the running counts are unequal. Default2/3, the value from Efron (1971).missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneEfron' object
DesignSeqOneByOneEfron$assign_wt()
Draw the next subject's treatment assignment via Efron's
(1971) biased coin (see class documentation): a fair coin if the running
treated/control counts are equal, otherwise a coin biased toward the
currently under-represented group at probability weighted_coin_prob.
Usage
DesignSeqOneByOneEfron$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneEfron$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneEfron$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Efron, B. (1971). "Forcing a sequential experiment to be balanced." Biometrika, 58(3), 403-417, doi:10.1093/biomet/58.3.403. See also randomized experiment for orientation on biased-coin sequential designs.
Examples
seq_des = DesignSeqOneByOneEfron$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
Kapelner and Krieger's (2014) Sequential "Matching-on-the-Fly" Design
Description
A DesignSeqOneByOne that matches subjects to
each other as they arrive, rather than requiring all subjects up front (as
DesignFixedBinaryMatch does): each new
subject is either matched to its nearest (Mahalanobis-distance) unmatched prior
subject in the "reservoir" — if that distance is small enough to pass a
statistical closeness test — and assigned the opposite treatment from its
match, or, if no sufficiently close match exists (or matching hasn't started yet),
randomized and added to the reservoir for future subjects to potentially match
against.
Burn-in. For the first t_0_pct * n subjects (default 35%), or
whenever no covariates are yet available, subjects are simply randomized
(\mathrm{Bernoulli}(prob\_T)) and placed in the reservoir (private$m
set to 0 for that subject) — matching does not begin until enough subjects have
accumulated to estimate a stable covariate covariance structure.
Matching test. Once past burn-in, for new subject t with covariate
vector x_t, the squared Mahalanobis distance (via
compute_proportional_mahal_distances_cpp(), using the sample covariance of
all prior subjects' covariates, ridge-regularized by .Machine$double.eps) to
every subject currently in the reservoir is computed, and the closest one is a
candidate match. The match is accepted only if that squared distance falls
below a threshold T^2_{\mathrm{cutoff}} derived from an F critical value,
T^2_{\mathrm{cutoff}} = \frac{p(n-1)}{n-p} \, F_{p,\, t-p}(\lambda),
where p is the rank of the covariate matrix so far, n is the design's
target (planned) sample size, t is the number of subjects enrolled so far, and
\lambda (lambda, default 0.1) is the F-distribution quantile level —
i.e. this is a
Hotelling's T^2-type test of whether the candidate pair's covariate
difference is small enough to plausibly be exchangeable "noise" rather than a
meaningful covariate mismatch; lambda controls how strict that test is
(smaller lambda accepts fewer, closer matches). If accepted, both subjects
are recorded as a new match (private$m), and the new subject receives the
opposite treatment of its match, guaranteeing exactly one treated and one
control per matched pair — the same guarantee
DesignFixedBinaryMatch provides, but
formed incrementally rather than all at once. If rejected (or the reservoir is
empty), the subject is randomized and added to the reservoir instead.
Lifecycle note. The morrison and p constructor arguments are
currently recorded on the object but not consulted anywhere in the matching or
assignment logic in this version of the class; treat them as reserved for future use
rather than as active configuration.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneKK14
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneKK14$is_a_kk_matching_capable()
Characterization: this design matches subjects to each other incrementally as they arrive (see class documentation), so it is KK matching-on-the-fly-capable by construction.
Usage
DesignSeqOneByOneKK14$is_a_kk_matching_capable()
Returns
Always TRUE for this class.
DesignSeqOneByOneKK14$new()
Initialize a KK14 sequential matching-on-the-fly experimental design (see class documentation for the burn-in and matching-test rules).
Usage
DesignSeqOneByOneKK14$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, lambda = NULL, t_0_pct = NULL, morrison = FALSE, p = NULL, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TProbability of treatment assignment used for burn-in/ unmatched (reservoir) subjects; matched subjects instead always receive the opposite assignment of their match (see class documentation).
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
lambdaThe F-distribution quantile level controlling how strict the matching-acceptance test is (default 0.1; smaller values accept fewer, closer matches). See class documentation for the exact threshold formula.
t_0_pctThe fraction of
nsubjects to randomize into the reservoir before matching begins (default 0.35).morrisonCurrently unused by this class's matching/assignment logic; reserved for future use.
pCurrently unused by this class's matching/assignment logic; reserved for future use.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneKK14' object
DesignSeqOneByOneKK14$assign_wt()
Draw the next subject's treatment assignment via KK14 matching-on-the-fly (see class documentation): during burn-in, or if no sufficiently close reservoir match exists, randomize and add the subject to the reservoir; otherwise match to the nearest reservoir subject and assign the opposite treatment.
Usage
DesignSeqOneByOneKK14$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneKK14$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneKK14$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148. See also Hotelling's T-squared distribution for the matching-test statistic, and Mahalanobis distance.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
Kapelner and Krieger's (2021) Outcome-Weighted Sequential Matching-on-the-Fly Design
Description
A DesignSeqOneByOneKK14 extension that
replaces KK14's unweighted Mahalanobis matching distance with a
response-weighted squared distance: at each assignment, a per-covariate
weight vector is re-estimated from the responses observed so far (weighting each
covariate by the estimated strength of its association with the response, e.g. an
absolute standardized regression coefficient), and the new subject is matched to its
nearest reservoir subject under that weighted distance rather than the raw
(unweighted) Mahalanobis distance KK14 uses. Weighting the match by outcome
association means matching effort is spent preferentially on prognostic
covariates (those that actually explain outcome variance) rather than equally on
every covariate, which is intended to improve estimator efficiency beyond what
outcome-agnostic matching achieves. morrison = TRUE additionally switches to
Morrison and Owen's (2025) alternative calibration of the matching-acceptance
threshold (differing in the fixed- vs. variable-n settings) and removes KK14's
burn-in wait before matching begins.
Weight estimation. compute_weights() dispatches on
response_type to a corresponding kk21_*_weights_cpp() backend that fits
a per-covariate simple regression of the response on that covariate (given all
responses observed so far) and returns the absolute t-statistic (coefficient over its
standard error) as that covariate's weight: OLS for "continuous", logistic for
"incidence", negative-binomial (or, if count_use_speedup = TRUE, OLS on
log(y + 1)) for "count", beta regression (or OLS on the logit scale if
proportion_use_speedup = TRUE) for "proportion", Weibull/lognormal/
log-logistic AFT (or OLS on log(y) if
survival_use_speedup_for_no_censoring = TRUE and there is no censoring yet)
for "survival", and proportional-odds (or OLS on the numeric-coerced level if
ordinal_use_speedup = TRUE) for "ordinal". Weights are normalized to
sum to 1 and cached per-iteration in private$iteration_weights (retrievable
via get_iteration_weights()); the *_use_speedup flags trade weight
accuracy for speed by substituting a fast continuous-regression proxy for the
response-type-appropriate GLM/AFT/ordinal fit on every single assignment call.
Weighted matching test. The weighted squared distance from the new subject
to every reservoir subject is computed
(compute_weighted_sqd_distances_cpp()), and the match is accepted only if the
minimum weighted distance falls below the private$compute_lambda() quantile of
a bootstrapped reference distribution of weighted pairwise distances among
subjects enrolled so far (compute_bootstrapped_weighted_sqd_distances_cpp(),
num_boot resamples) — a nonparametric, simulation-based acceptance threshold,
in contrast to KK14's closed-form F-distribution threshold. As in KK14, an accepted
match receives the opposite treatment of its match; a rejected (or empty-reservoir)
draw is randomized and added to the reservoir.
Fallback to KK14. Before enough responses have accumulated to fit the
weight-estimation regressions reliably (fewer than 2 * (ncol(X) + 2)
non-missing responses — two observations per regression parameter, at minimum
ncol(X) + 2), assign_wt() falls back to the inherited (unweighted) KK14
assignment rule entirely, via super$assign_wt().
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneKK14 -> DesignSeqOneByOneKK21
Methods
Public methods
+ inherited public methods from DesignSeqOneByOneKK14
DesignSeqOneByOneKK14$add_all_subject_matched_pair_ids()DesignSeqOneByOneKK14$assert_blocking_design()DesignSeqOneByOneKK14$assert_equal_block_sizes()DesignSeqOneByOneKK14$assert_matching_design()DesignSeqOneByOneKK14$get_block_ids()DesignSeqOneByOneKK14$get_cmh_se_w_mat()DesignSeqOneByOneKK14$get_matching_cluster_ids()DesignSeqOneByOneKK14$inject_cmh_se_w_mat()DesignSeqOneByOneKK14$is_a_kk_matching_capable()DesignSeqOneByOneKK14$is_blocking_design()DesignSeqOneByOneKK14$is_complete_blocking_design()DesignSeqOneByOneKK14$is_matching_design()DesignSeqOneByOneKK14$set_m()DesignSeqOneByOneKK14$summarize_blocks()
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_fixed_sample_size()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneKK21$new()
Initialize a matching-on-the-fly sequential experimental design which matches based on Kapelner and Krieger (2021) with option to use matching parameters of Morrison and Owen (2025)
Usage
DesignSeqOneByOneKK21$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, lambda = NULL, t_0_pct = NULL, morrison = FALSE, p = NULL, num_boot = NULL, count_use_speedup = TRUE, proportion_use_speedup = TRUE, survival_use_speedup_for_no_censoring = TRUE, ordinal_use_speedup = TRUE, missingness_method = "impute", design_formula = ~., seed = NULL, ... )
Arguments
response_typeThe data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival". This package will enforce that all added responses via the
add_one_subject_responsemethod will be of the appropriate type.prob_TThe probability of the treatment assignment. This defaults to
0.5.include_is_missing_as_a_new_featureIf missing data is present in a variable, should we include another dummy variable for its missingness in addition to imputing its value? If the feature is type factor, instead of creating a new column, we allow missingness to be its own level. The default is
TRUE.nThe sample size (if fixed). Default is
NULLfor not fixed.verboseA flag indicating whether messages should be displayed to the user. Default is
FALSE.lambdaThe quantile cutoff of the subject distance distribution for determining matches. If unspecified and
morrison = FALSE, default is 10%.t_0_pctThe percentage of total sample size n where matching begins. If unspecified and
morrison = FALSE, default is 35%.morrisonDefault is
FALSEwhich implies matching via the KK14 algorithm usinglambdaandt_0_pctmatching. IfTRUE, we use Morrison and Owen (2025)'s formula forlambdawhich differs in the fixed n versus variable n settings and matching begins immediately with no wait for a certain reservoir size like in KK14.pThe number of covariate features. Must be specified when
morrison = TRUEotherwise do not specify this argument.num_bootthe number of bootstrap samples taken to approximate the subject-distance distribution. Default is 500.
count_use_speedupShould we speed up the estimation of the weights in the response = count case via a continuous regression on log(y + 1). instead of a negative binomial regression each time? This is at the expense of the weights being less accurate. Default is
TRUE.proportion_use_speedupShould we speed up the estimation of the weights in the response = proportion case via a continuous regression on log(y / (1 - y)) instead of a beta regression each time? This is at the expense of the weights being less accurate. Default is
TRUE.survival_use_speedup_for_no_censoringShould we speed up the estimation of the weights in the response = survival case via a continuous regression on log(y) instead of a Weibull AFT regression each time, but only when there is no censoring in the data collected so far? This is at the expense of the weights being less accurate when censoring is present. Default is
TRUE.ordinal_use_speedupShould we speed up the estimation of the weights in the response = ordinal case via a continuous regression on the ordinal levels coerced to numeric. instead of a proportional odds model each time? This is at the expense of the weights being less accurate. Default is
TRUE.missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
...Extra arguments passed to the
DesignSeqOneByOneKK14superclass.
Returns
A new 'DesignSeqOneByOneKK21' object
Examples
seq_des = DesignSeqOneByOneKK21$new(n = 6, response_type = "continuous") seq_des$add_one_subject_to_experiment_and_assign(data.frame(x = rnorm(1)))
DesignSeqOneByOneKK21$get_iteration_weights()
Retrieve the full history of normalized covariate weight
vectors computed by compute_weights() across every assignment call
so far (see class documentation), keyed by subject index t, for
inspecting how the outcome-informed weighting evolved as data accrued.
Usage
DesignSeqOneByOneKK21$get_iteration_weights()
Returns
A list of numeric weight vectors (one per assignment call at which weights were computed), each summing to 1 and named by covariate.
DesignSeqOneByOneKK21$get_covariate_weights()
Retrieve the normalized covariate weight vector from the most recent assignment call (see class documentation for how weights are estimated).
Usage
DesignSeqOneByOneKK21$get_covariate_weights()
Returns
A numeric vector of weights, one per covariate, summing to 1 and
named by covariate; NULL if weights have not yet been computed
(e.g. still in the KK14 fallback regime).
DesignSeqOneByOneKK21$assign_wt()
Draw the next subject's treatment assignment via the KK21 outcome-weighted matching-on-the-fly rule (see class documentation): falls back to unweighted KK14 matching if too few responses have accumulated to estimate covariate weights, otherwise re-estimates weights, matches to the nearest reservoir subject under the weighted distance if it clears the bootstrapped acceptance threshold (assigning the opposite treatment), or randomizes into the reservoir otherwise.
Usage
DesignSeqOneByOneKK21$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneKK21$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneKK21$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential
allocation with higher power and efficiency." Biometrics, 70(2), 378-388,
doi:10.1111/biom.12148, for the base matching-on-the-fly algorithm this class
extends with outcome-weighted distances (Kapelner and Krieger, 2021); see also
Morrison, T., and Owen, A. B. (2025) for the alternative morrison = TRUE
threshold calibration referenced by the morrison argument.
Examples
seq_des = DesignSeqOneByOneKK21$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
## ------------------------------------------------
## Method `DesignSeqOneByOneKK21$new()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneKK21$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x = rnorm(1)))
Stepwise Variant of the KK21 Outcome-Weighted Sequential Matching Design
Description
A DesignSeqOneByOneKK21 variant that
computes its per-covariate matching weights via forward stepwise selection
(compute_weights_KK21stepwise()) instead of KK21's independent
marginal-association regressions: covariates are added to a growing "selected" set
one at a time, at each step choosing whichever remaining covariate has the largest
absolute association statistic conditional on (i.e. in a model that also
includes) the covariates already selected and the treatment-assignment column,
rather than each covariate's association with the response considered in isolation.
This targets the case where covariates are mutually correlated: KK21's marginal
weights can assign similar high weight to several collinear prognostic covariates
(effectively double-counting the same information), whereas the stepwise conditional
weights down-weight a covariate once its explanatory content is already captured by
previously selected covariates.
Weight computation. For each response type, a family-appropriate model
(OLS/logistic/negative-binomial/beta/AFT survival/proportional-odds, matching the
same response-type dispatch and *_use_speedup fast-path conventions as
DesignSeqOneByOneKK21) is repeatedly refit,
each time regressing the response on one candidate remaining covariate plus all
previously selected covariates plus the treatment column ws; the candidate
with the largest absolute association statistic is selected next and assigned that
statistic as its weight, then removed from the candidate pool, and the process
repeats until every covariate has been assigned a weight (an O(p^2) number of
model fits per assignment call, for p covariates). If a candidate's model fit
fails to converge (e.g. perfect separation or rank deficiency) partway through, the
remaining not-yet-selected covariates' weights are left NA internally and then
replaced with 0 (excluding them from the weighted matching distance) rather than
propagating the failure.
Everything else is inherited from KK21. Burn-in fallback to KK14, the
bootstrapped acceptance-threshold test, matched-pair assignment, and the
morrison/lambda/t_0_pct matching-schedule options are all
unchanged from DesignSeqOneByOneKK21; only
how the covariate weight vector is computed differs.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneKK14 -> DesignSeqOneByOneKK21 -> DesignSeqOneByOneKK21stepwise
Methods
Public methods
+ inherited public methods from DesignSeqOneByOneKK21
+ inherited public methods from DesignSeqOneByOneKK14
DesignSeqOneByOneKK14$add_all_subject_matched_pair_ids()DesignSeqOneByOneKK14$assert_blocking_design()DesignSeqOneByOneKK14$assert_equal_block_sizes()DesignSeqOneByOneKK14$assert_matching_design()DesignSeqOneByOneKK14$get_block_ids()DesignSeqOneByOneKK14$get_cmh_se_w_mat()DesignSeqOneByOneKK14$get_matching_cluster_ids()DesignSeqOneByOneKK14$inject_cmh_se_w_mat()DesignSeqOneByOneKK14$is_a_kk_matching_capable()DesignSeqOneByOneKK14$is_blocking_design()DesignSeqOneByOneKK14$is_complete_blocking_design()DesignSeqOneByOneKK14$is_matching_design()DesignSeqOneByOneKK14$set_m()DesignSeqOneByOneKK14$summarize_blocks()
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_fixed_sample_size()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneKK21stepwise$new()
Initialize a matching-on-the-fly sequential experimental design whose covariate matching weights are computed via forward stepwise selection (see class documentation), based on the stepwise version of Kapelner and Krieger (2021) with option to use matching parameters of Morrison and Owen (2025)
Usage
DesignSeqOneByOneKK21stepwise$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, lambda = NULL, t_0_pct = NULL, morrison = FALSE, p = NULL, num_boot = NULL, count_use_speedup = TRUE, proportion_use_speedup = TRUE, survival_use_speedup_for_no_censoring = TRUE, ordinal_use_speedup = TRUE, missingness_method = "impute", design_formula = ~., ... )
Arguments
response_typeThe data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival". This package will enforce that all added responses via the
add_one_subject_responsemethod will be of the appropriate type.prob_TThe probability of the treatment assignment. This defaults to
0.5.include_is_missing_as_a_new_featureIf missing data is present in a variable, should we include another dummy variable for its missingness in addition to imputing its value? If the feature is type factor, instead of creating a new column, we allow missingness to be its own level. The default is
TRUE.nThe sample size (if fixed). Default is
NULLfor not fixed.verboseA flag indicating whether messages should be displayed to the user. Default is
FALSE.lambdaThe quantile cutoff of the subject distance distribution for determining matches. If unspecified and
morrison = FALSE, default is 10%.t_0_pctThe percentage of total sample size n where matching begins. If unspecified and
morrison = FALSE, default is 35%.morrisonDefault is
FALSEwhich implies matching via the KK14 algorithm usinglambdaandt_0_pctmatching. IfTRUE, we use Morrison and Owen (2025)'s formula forlambdawhich differs in the fixed n versus variable n settings and matching begins immediately with no wait for a certain reservoir size like in KK14.pThe number of covariate features. Must be specified when
morrison = TRUEotherwise do not specify this argument.num_bootthe number of bootstrap samples taken to approximate the subject-distance distribution. Default is 500.
count_use_speedupShould we speed up the estimation of the weights in the response = count case via a continuous regression on log(y + 1). instead of a negative binomial regression each time? This is at the expense of the weights being less accurate. Default is
TRUE.proportion_use_speedupShould we speed up the estimation of the weights in the response = proportion case via a continuous regression on log(y / (1 - y)) instead of a beta regression each time? This is at the expense of the weights being less accurate. Default is
TRUE.survival_use_speedup_for_no_censoringShould we speed up the estimation of the weights in the response = survival case via a continuous regression on log(y) instead of a Weibull AFT regression each time, but only when there is no censoring in the data collected so far? This is at the expense of the weights being less accurate when censoring is present. Default is
TRUE.ordinal_use_speedupShould we speed up the estimation of the weights in the response = ordinal case via a continuous regression on the ordinal levels coerced to numeric. instead of a proportional odds model each time? This is at the expense of the weights being less accurate. Default is
TRUE.missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
...Extra arguments passed to the
DesignSeqOneByOneKK21superclass.
Returns
A new 'DesignSeqOneByOneKK21stepwise' object
Examples
seq_des = DesignSeqOneByOneKK21stepwise$new(n = 6, response_type = "continuous")
DesignSeqOneByOneKK21stepwise$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneKK21stepwise$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential
allocation with higher power and efficiency." Biometrics, 70(2), 378-388,
doi:10.1111/biom.12148, for the base matching-on-the-fly algorithm; the
outcome-weighted extension follows Kapelner and Krieger (2021), with this class
using a forward-stepwise (rather than marginal) weight-estimation scheme. See also
Morrison, T., and Owen, A. B. (2025) for the alternative morrison = TRUE
threshold calibration referenced by the morrison argument.
Examples
seq_des = DesignSeqOneByOneKK21stepwise$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
## ------------------------------------------------
## Method `DesignSeqOneByOneKK21stepwise$new()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneKK21stepwise$new(n = 6, response_type = "continuous")
Pocock and Simon's (1975) Minimization Sequential Design
Description
A DesignSeqOneByOne implementing Pocock and
Simon's minimization method: for each categorical covariate in strata_cols,
the design tracks a running treated/control count per covariate level
(private$counts), and for a new subject computes, for each candidate
treatment arm k \in \{0, 1\}, a weighted total imbalance
G_k = \sum_{j} weights_j \cdot \mathrm{Var}\big(\text{counts at subject's level of covariate } j,
\text{ after hypothetically assigning arm } k\big),
where the variance is taken across the two treatment arms' hypothetical counts at
that covariate level (so G_k is large when arm k would leave the
subject's covariate-level counts unbalanced, summed with weights across
covariates). The subject is then assigned to whichever arm minimizes G_k with
probability p_best (and to the other arm with probability
1 - p_best), or — if the two arms are exactly tied — via a plain
\mathrm{Bernoulli}(prob\_T) draw. Unlike
DesignSeqOneByOneAtkinson/
DesignSeqOneByOneKK14, which use continuous
covariate distances, minimization operates on categorical/discretized strata and
balances marginal covariate-level counts directly rather than a multivariate
distance or matched-pair structure.
Level bookkeeping. private$ensure_factor_metadata() maintains a
mapping from each observed level of each strata_cols column to a row index
in private$counts (an (total levels across all covariates) x 2 matrix of
running treated/control counts), growing both the level map and counts as
new levels are encountered; missing values are treated as their own level
("NA").
Non-resampling bootstrap. draw_bootstrap_indices() always performs a
plain i.i.d. nonparametric bootstrap over subjects (sample_int_replace_cpp()),
since minimization's adaptive assignment process has no simple exchangeable
resampling unit to preserve (each subject's assignment probability depends on the
full sequence of covariate levels and assignments that preceded it).
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOnePocockSimon
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOnePocockSimon$new()
Initialize a Pocock and Simon (1975) minimization sequential experimental design (see class documentation for the exact imbalance criterion and assignment rule).
Usage
DesignSeqOneByOnePocockSimon$new( strata_cols, weights = NULL, p_best = 0.8, response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
strata_colsThe names of the covariates to be used for minimization. These must be factor or categorical variables.
weightsA numeric vector of per-covariate weights
weights_jin the imbalance criterionG_k(see class documentation), one per entry ofstrata_cols, in the same order. Defaults to 1 for all (equal-weighted covariates).p_bestThe probability of assigning the treatment arm that minimizes
G_k(see class documentation); the complementary arm is assigned with probability1 - p_best. Defaults to 0.8 (an 80/20 biased coin favoring the balancing arm, rather than a fully deterministic minimization rule).response_typeThe data type of response values.
prob_TThe probability of the treatment assignment used only when the two arms' imbalance is exactly tied (see class documentation).
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseFlag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOnePocockSimon' object
DesignSeqOneByOnePocockSimon$assign_wt()
Draw the next subject's treatment assignment via Pocock and
Simon minimization (see class documentation for the exact imbalance
criterion G_k and the p_best/prob_T assignment rule),
and update the running per-covariate-level treated/control counts
in-place to reflect this assignment.
Usage
DesignSeqOneByOnePocockSimon$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOnePocockSimon$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOnePocockSimon$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Pocock, S. J., and Simon, R. (1975). "Sequential treatment assignment with balancing for prognostic factors in the controlled clinical trial." Biometrics, 31(1), 103-115, doi:10.2307/2529712. See also minimisation (clinical trials) for orientation.
Examples
seq_des = DesignSeqOneByOnePocockSimon$new(strata_cols = 'x1', n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = factor(1, levels=1:2)))
A Sequential Permuted-Block Design with Randomly Varying Block Sizes
Description
A DesignSeqOneByOne implementing permuted-block
randomization with randomly varying block size: subjects are assigned from a queue
of pre-shuffled treatment labels (a "block"), refilled with a fresh
sampled block whenever it empties. Each new block's size is
itself drawn uniformly at random from block_sizes (rather than being fixed),
and each block internally contains exactly round(block_size * prob_T) treated
and block_size - round(block_size * prob_T) control labels in random order.
Randomizing the block size (rather than using a single fixed block length, as in
classical permuted-block designs) is a standard clinical-trials safeguard against
selection bias: with a fixed, known block size, unblinded staff could
predict the last assignment(s) in a block from the ones already observed, whereas an
unpredictable block size makes this much harder while still guaranteeing
near-perfect treatment/control balance throughout enrollment (balance is exact at
every block boundary and never worse than one full block's imbalance in between). If
strata_cols is supplied, a separate independent sequence of blocks is
maintained per stratum (one queue per distinct combination of strata_cols
values), so balance holds within each stratum, not just overall.
Block-size / prob_T compatibility. Every entry of block_sizes
must yield an integer number of treated subjects when multiplied by prob_T
(checked at construction: abs(bs * prob_T - round(bs * prob_T)) <= 1e-10 for
every bs); a block size that would require a fractional number of treated
subjects is rejected.
Per-stratum queues. private$strata_states is a hashed environment
mapping each stratum key (or the literal key "overall" when
strata_cols is NULL) to the vector of not-yet-used assignments
remaining in that stratum's current block; assign_wt() pops the next
assignment from the relevant queue, refilling it with a freshly drawn block (random
size, randomly ordered) whenever it is empty.
Bootstrap. draw_bootstrap_indices() resamples within strata (via
stratified_bootstrap_indices_cpp()) when strata_cols is supplied, or
performs a plain i.i.d. nonparametric bootstrap over subjects otherwise.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneRandomBlockSize
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneRandomBlockSize$new()
Initialize a sequential permuted-block experimental design with randomly varying block size (see class documentation for the exact block-refill rule and its selection-bias rationale).
Usage
DesignSeqOneByOneRandomBlockSize$new( strata_cols = NULL, block_sizes = c(4, 6, 8), response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
strata_colsA character vector of column names to use for stratification. If NULL, simple blocking is used.
block_sizesA vector of positive integers representing the possible block sizes to choose from. Each must be a multiple of the inverse of
prob_Tto ensure integer treatment/control counts.response_typeThe data type of response values which must be one of the following: "continuous", "incidence", "proportion", "count", "survival", "ordinal".
prob_TThe probability of the treatment assignment. This defaults to
0.5.include_is_missing_as_a_new_featureIf missing data is present in a variable, should we include another dummy variable for its missingness? Default is
TRUE.nThe sample size (if fixed). Default is
NULLfor not fixed.verboseA flag indicating whether messages should be displayed. Default is
FALSE.missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneRandomBlockSize' object
DesignSeqOneByOneRandomBlockSize$assign_wt()
Pop the next treatment assignment from the current subject's stratum block queue (see class documentation), refilling that queue with a freshly drawn random-size, randomly-ordered block first if it is empty.
Usage
DesignSeqOneByOneRandomBlockSize$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneRandomBlockSize$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneRandomBlockSize$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Efron, B. (1971). "Forcing a sequential experiment to be balanced." Biometrika, 58(3), 403-417, doi:10.1093/biomet/58.3.403, for sequential balanced-block randomization background. See also block randomisation for orientation on permuted-block designs and the selection-bias rationale for varying block size.
Examples
seq_des = DesignSeqOneByOneRandomBlockSize$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
A Stratified Permuted-Block Sequential Design (SPBR) with Fixed Block Size
Description
A DesignSeqOneByOne implementing classical
stratified permuted-block randomization: subjects are assigned from a per-stratum
queue of pre-shuffled treatment labels (a "block" of block_size labels,
containing exactly round(block_size * prob_T) treated and the remainder
control, in random order), refilled with a fresh randomly-ordered block of the
same fixed size whenever a stratum's queue empties. This is the
fixed-block-size, mandatory-stratification counterpart of
DesignSeqOneByOneRandomBlockSize
(which varies block size across draws and makes stratification optional): here,
strata_cols is required, and a single fixed block_size is used for
every stratum and every block, guaranteeing exact treatment/control balance within
each stratum at every block boundary.
Block-size / prob_T compatibility. block_size must yield an
integer number of treated subjects: the constructor errors unless
abs(block_size * prob_T - round(block_size * prob_T)) <= 1e-10.
Per-stratum queues. As in
DesignSeqOneByOneRandomBlockSize,
private$strata_states is a hashed environment mapping each stratum key
(concatenated strata_cols values, "NA" for missing) to the vector of
not-yet-used assignments remaining in that stratum's current block;
assign_wt() pops the next assignment, refilling with a fresh block when
empty. draw_bootstrap_indices() resamples within strata by default
(bootstrap_type = "within_blocks" or NULL) or resamples whole strata
otherwise.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneSPBR
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneSPBR$new()
Initialize a stratified permuted-block sequential experimental design with fixed block size (see class documentation for the exact block-refill rule).
Usage
DesignSeqOneByOneSPBR$new( strata_cols, block_size = 4, response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
strata_colsA character vector of column names to use for stratification.
block_sizeThe size of the permuted blocks (fixed; see class documentation for its compatibility requirement with
prob_T).response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TProbability of treatment assignment.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneSPBR' object
DesignSeqOneByOneSPBR$assign_wt()
Pop the next treatment assignment from the current subject's stratum block queue (see class documentation), refilling that queue with a freshly drawn fixed-size, randomly-ordered block first if it is empty.
Usage
DesignSeqOneByOneSPBR$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneSPBR$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneSPBR$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Zelen, M. (1974). "The randomization and stratification of patients to
clinical trials." Journal of Chronic Diseases, 27(7-8), 365-375,
doi:10.1016/0021-9681(74)90015-0, for stratified permuted-block randomization.
See also block
randomisation for orientation, and
DesignSeqOneByOneRandomBlockSize
for the randomly-varying-block-size variant.
Examples
seq_des = DesignSeqOneByOneSPBR$new(strata_cols = 'x1', n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = factor(1, levels=1:2)))
Wei's (1977, 1978) Adaptive Urn Sequential Design, UD(\alpha, \beta)
Description
A DesignSeqOneByOne implementing Wei's
adaptive biased-coin urn design UD(\alpha, \beta): conceptually, an urn starts
with \alpha balls of each type (treatment and control), and each assignment
is drawn proportionally to the current ball counts, then \beta balls of the
opposite type to whatever was drawn are added back to the urn (so drawing
treatment adds \beta control balls, and vice versa), pushing subsequent draws
toward the under-represented arm. No covariates are used; only the running
treated/control counts n_T, n_C matter, via the closed-form assignment
probability
\Pr(w_t = 1) = \frac{\alpha + \beta \, n_C}{2\alpha + \beta (n_T + n_C)}.
Like DesignSeqOneByOneEfron, this design
balances running assignment counts online while remaining strictly randomized (the
probability is always strictly between 0 and 1 for finite \alpha, \beta > 0);
unlike Efron's design (which only distinguishes "balanced" vs. "imbalanced" and
applies a single fixed weighted_coin_prob in the imbalanced case), the urn
design's bias toward the under-represented arm scales continuously and smoothly with
the current degree of imbalance, tuned by the ratio \beta/\alpha: larger
\beta/\alpha yields stronger balancing pressure, and \beta = 0 recovers a
fixed \mathrm{Bernoulli}(0.5) coin (no adaptation at all).
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneUrn
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneUrn$new()
Initialize Wei's UD(\alpha, \beta) adaptive urn
sequential experimental design (see class documentation for the exact
assignment-probability formula).
Usage
DesignSeqOneByOneUrn$new( alpha = 1, beta = 1, response_type, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
alphaThe initial number of balls of each type (Treatment/Control) in the conceptual urn; larger
alpharelative tobetaweakens the balancing effect (assignment probabilities stay closer to 0.5 for longer).betaThe number of balls of the opposite type added to the urn after each assignment;
beta = 0recovers an unbiased\mathrm{Bernoulli}(0.5)coin (no balancing).response_typeThe data type of response values.
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneUrn' object
DesignSeqOneByOneUrn$assign_wt()
Draw the next subject's treatment assignment from Wei's
UD(\alpha, \beta) urn probability (see class documentation),
computed from the running treated/control counts.
Usage
DesignSeqOneByOneUrn$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneUrn$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneUrn$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Wei, L. J. (1977). "A class of designs for sequential clinical trials." Journal of the American Statistical Association, 72(358), 382-386, doi:10.1080/01621459.1977.10481006; Wei, L. J. (1978). "The adaptive biased coin design for sequential experiments." The Annals of Statistics, 6(1), 92-100, doi:10.1214/aos/1176344068. See also randomized experiment for orientation on adaptive biased-coin sequential designs.
Examples
seq_des = DesignSeqOneByOneUrn$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
A Sequential Design Guaranteeing Exact Terminal Balance (Random Allocation Rule)
Description
A DesignSeqOneByOne implementing the "random
allocation rule" (a sequential realization of complete randomization): treatment is
assigned to each arriving subject with probability equal to the fraction of
remaining treatment slots among all remaining slots,
\Pr(w_t = 1) = n_{T,\mathrm{rem}} / (n_{T,\mathrm{rem}} + n_{C,\mathrm{rem}}),
where n_{T,\mathrm{rem}} = \mathrm{round}(n \cdot prob\_T) - n_T and
n_{C,\mathrm{rem}} = (n - \mathrm{round}(n \cdot prob\_T)) - n_C are the
treatment/control slots not yet used, given the running counts n_T, n_C.
This guarantees the realized sequence, once all n subjects have arrived, has
exactly \mathrm{round}(n \cdot prob\_T) treated subjects — the same
terminal allocation-count guarantee as
DesignFixediBCRD's complete randomization, but
realized online as subjects arrive one at a time rather than all at once, and with
every prefix of the sequence itself drawn from the correct conditional (hypergeometric)
distribution given the slots used so far. If a slot type is exhausted
(n_{T,\mathrm{rem}} \le 0 or n_{C,\mathrm{rem}} \le 0), the remaining
subjects are deterministically assigned to whichever type still has open slots.
No target n: falls back to Bernoulli. If n was not supplied at
construction (private$n is NULL), there is no terminal target to
balance toward, so assign_wt() falls back to an unbiased
\mathrm{Bernoulli}(prob\_T) draw for every subject instead (equivalent to
DesignSeqOneByOneBernoulli).
Single implicit block. add_one_subject_to_experiment_and_assign()
overrides the inherited method only to additionally set private$m to a
constant vector of 1s (a single block containing every subject enrolled so far)
after each assignment, mirroring the fixed-sample
DesignFixediBCRD's single-block convention for
shared blocking/matching machinery.
Super classes
Design -> DesignSeqOneByOne -> DesignSeqOneByOneiBCRD
Methods
Public methods
+ inherited public methods from DesignSeqOneByOne
+ inherited public methods from Design
Design$add_all_subject_responses()Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_even_allocation()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$overwrite_all_subject_assignments()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_randomization_draw()Design$supports_resampling()Design$supports_resampling_replay()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
DesignSeqOneByOneiBCRD$new()
Initialize a sequential design targeting exact terminal
treatment/control balance (see class documentation for the assignment
rule and the no-fixed-n fallback).
Usage
DesignSeqOneByOneiBCRD$new( response_type, prob_T = 0.5, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
prob_TTarget probability of treatment assignment; the terminal number of treated subjects is fixed at
round(n * prob_T)whennis known (see class documentation).include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe planned (target) sample size; if
NULL, there is no terminal balance target and assignment falls back to an unbiased Bernoulli coin (see class documentation).verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'DesignSeqOneByOneiBCRD' object
DesignSeqOneByOneiBCRD$add_one_subject_to_experiment_and_assign()
Add one subject to the experiment and assign treatment via
assign_wt() (delegating to the inherited
DesignSeqOneByOne$add_one_subject_to_experiment_and_assign()),
then set private$m to a single-block vector of 1s covering every
subject enrolled so far (see class documentation), overwritten on every
call rather than only once all subjects have arrived.
Usage
DesignSeqOneByOneiBCRD$add_one_subject_to_experiment_and_assign(x_new)
Arguments
x_newA data frame with one row representing the new subject's covariates.
Returns
The treatment assignment (0 or 1) for the newly added subject.
DesignSeqOneByOneiBCRD$assign_wt()
Draw the next subject's treatment assignment via the random
allocation rule (see class documentation): with probability equal to the
fraction of remaining treatment slots among all remaining slots, or a
deterministic assignment if one slot type is exhausted; falls back to an
unbiased Bernoulli coin if no fixed n was supplied.
Usage
DesignSeqOneByOneiBCRD$assign_wt()
Returns
The treatment assignment (0 or 1) for the next subject.
DesignSeqOneByOneiBCRD$clone()
The objects of this class are cloneable with this method.
Usage
DesignSeqOneByOneiBCRD$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Rosenberger, W. F., and Lachin, J. M. (2016). Randomization in
Clinical Trials: Theory and Practice (2nd ed.), Wiley, for the random allocation
rule as a sequential implementation of complete randomization. See also
DesignFixediBCRD for the fixed-sample (all-at-once)
version of the same terminal randomization law.
Examples
seq_des = DesignSeqOneByOneiBCRD$new(n = 6, response_type = 'continuous')
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
Comprehensive-test slow-path registry
Description
Performance-based exclusions used by EDI's comprehensive test harness. These rules describe paths that are implemented but intentionally omitted from routine exhaustive execution because their observed runtime is too high. They do not change an inference class's public capabilities and must not be interpreted as "not implemented" declarations.
Usage
EDI_COMPREHENSIVE_SLOW_PATHS
Format
A named list. 'exact_operations' contains keys of the form 'response_type||InferenceClass||operation', optionally suffixed with '||model_formula' (e.g. '||~1') to restrict the entry to that one formula – an operation with no formula suffix is skipped for every formula, matching the pre-2026-09-08 behavior. 'model_formula' is matched as the deparsed formula text ('"~1"', '"~."', ...). Every other element contains formula-, dataset-, and design-independent concrete inference-class names for the named slow-path family.
Details
Rule-name prefixes use the harness vocabulary: 'boot' is ordinary nonparametric bootstrap, 'bbt' is Bayesian bootstrap, 'brt' is bootstrap randomization, 'pboot'/'param_bootstrap' are parametric bootstrap, and 'rand' is randomization inference. Suffixes identify the affected CI, p-value, or typed variant. A class can appear in more than one category.
'InferenceSuite$run_all_inference()' also omits matching class/method/type combinations when its 'methods' argument is left at the default 'NULL'. Supplying 'methods' explicitly opts into the requested paths even when they appear in this registry. Because of that coupling, an 'exact_operations' entry must never name the base 'compute_estimate' operation itself (only a CI/p-value sub-method, e.g. 'compute_rand_two_sided_pval') – doing so doesn't just skip one slow sub-computation, it silently drops the whole class from ‘run_all_inference()'’s default output (found 2026-08-26: a pre-existing '"survival||InferenceSurvivalCoxPHRegr||compute_estimate"' entry, harmless while this registry was internal-only, started doing exactly that the moment 'run_all_inference()' began consulting it – 'compute_estimate()' for that class takes ~0.09s, nowhere near "too slow"; removed).
See Also
[InferenceSuite]
Split from a single "RR" tag into "RR" (raw risk ratio) vs. "log_risk_ratio" – per user decision, 2026-08-23: 'InferenceIncidGCompRiskRatio'/ 'InferenceIncidKKGCompRiskRatio' compute the raw ratio 'risk1/risk0' directly (a nonlinear function of an underlying logistic model's coefficients) and only touch log space as a delta-method device to get a positive-respecting CI/SE, exponentiating back before returning – "RR" is their natural, directly-computed scale. 'InferenceIncidModifiedPoisson'/ 'InferenceIncidLogBinomial'/'InferenceIncidKKModifiedPoisson' instead fit a genuine log-link regression model (Zou's modified-Poisson working likelihood, or a log-link binomial GLM) whose own coefficient *is* log(RR) by construction, with a directly-computed (non-delta-method) coefficient SE – log(RR) is their natural scale, and "RR" is the derived quantity ('exp(coefficient)'). Both groups previously shared one "RR" tag, which put a log10 x-axis and null-reference line at 1 on what were actually already-log-scale estimates/CIs for the second group – a genuine scale mismatch, not just a display nicety.
Description
Split from a single "RR" tag into "RR" (raw risk ratio) vs. "log_risk_ratio" – per user decision, 2026-08-23: 'InferenceIncidGCompRiskRatio'/ 'InferenceIncidKKGCompRiskRatio' compute the raw ratio 'risk1/risk0' directly (a nonlinear function of an underlying logistic model's coefficients) and only touch log space as a delta-method device to get a positive-respecting CI/SE, exponentiating back before returning – "RR" is their natural, directly-computed scale. 'InferenceIncidModifiedPoisson'/ 'InferenceIncidLogBinomial'/'InferenceIncidKKModifiedPoisson' instead fit a genuine log-link regression model (Zou's modified-Poisson working likelihood, or a log-link binomial GLM) whose own coefficient *is* log(RR) by construction, with a directly-computed (non-delta-method) coefficient SE – log(RR) is their natural scale, and "RR" is the derived quantity ('exp(coefficient)'). Both groups previously shared one "RR" tag, which put a log10 x-axis and null-reference line at 1 on what were actually already-log-scale estimates/CIs for the second group – a genuine scale mismatch, not just a display nicety.
Usage
EDI_INFERENCE_ESTIMAND_TAGS
Exact binomial incidence component source
Description
Source list for the exact-binomial incidence component.
Usage
ExactBinomialIncidenceSource
Exact Fisher incidence component source
Description
Source list for the exact Fisher incidence component.
Usage
ExactFisherIncidenceSource
Exact Zhang incidence component source
Description
Source list for the exact Zhang incidence component.
Usage
ExactZhangIncidenceSource
KK conditional-logit IVWC component source
Description
Initialize conditional-logistic IVWC inference for KK binary
responses and prepare separate matched-pair and reservoir likelihood
components used by InferenceIncidKKCondLogitIVWC.
Computes the class-specific treatment-effect estimate; see
Inference.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
IncidKKCondLogitIVWCSource
Details
Source list for the KK conditional-logit inverse-variance-weighted-combination (IVWC) incidence component.
Conditional Logistic Combined-Likelihood Inference for KK Designs with Binary Responses
Description
Initialize conditional-logistic combined-likelihood inference
for KK binary responses and prepare matched-pair conditional-logit plus
reservoir Bernoulli likelihood components. See
InferenceAsympLik for shared
likelihood-test methods.
Computes the class-specific treatment-effect estimate; see
Inference.
Recomputes the combined conditional-logistic estimate under Bayesian-bootstrap weights.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
IncidKKCondLogitOneLikLikelihoodSource
Details
Fits a single joint likelihood over all KK design data for incidence responses. The matched-pair component uses the conditional logistic likelihood, and the reservoir component uses the standard Bernoulli log-likelihood.
Identity-link binomial regression component source
Description
Initialize inference for the identity-link binomial risk-
difference model P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top
\gamma; see
InferenceIncidBinomialIdentityRiskDiff for the model form. Does not
fit the model; the fit is deferred to the first call to
compute_estimate() or a method that requires it.
Refits the identity-link binomial model with subject/block-level
weights applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_identity_binomial_regression_weighted_cpp, and returns
the reweighted risk-difference estimate \hat\beta_T^{(w)}. Uses the
same QR column-dropping hardening and fit-reasonableness check as
compute_estimate(); a hardened-but-still-unreasonable fit is
cached as nonestimable and returns NA.
Likelihood-ratio confidence interval for \beta_T by test
inversion (find the set of delta not rejected at level
alpha by the likelihood-ratio test); see
InferenceAsympLik for the shared
inversion contract. Falls back to a nonestimable result (NA
bounds) if the underlying root-finding fails.
Usage
IncidenceBinomialIdentityLikelihoodSource
Details
Source list for the IncidenceBinomialIdentityLikelihood component composed by InferenceIncidBinomialIdentityRiskDiff.
KK incidence g-computation component source
Description
Initialize KK marginal g-computation inference for a completed incidence design; prepares the KK match structure used by the cluster-robust sandwich covariance.
Usage
IncidenceKKGComputationSource
Details
Source list for the IncidenceKKGComputation component shared by the KK g-computation incidence classes.
Log-binomial likelihood component source
Description
Initialize inference for the log-link binomial risk-ratio
model \log P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top \gamma;
see InferenceIncidLogBinomial
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Refits the log-binomial model with subject/block-level weights
applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_log_binomial_regression_weighted_cpp, and returns the
reweighted log-risk-ratio estimate \hat\beta_T^{(w)}. Uses the same
QR column-dropping hardening and fit-reasonableness check as
compute_estimate(); a hardened-but-still-unreasonable fit is
cached as nonestimable and returns NA.
Score confidence interval for \beta_T by test inversion
of the score test (find the set of delta not rejected at level
alpha); see InferenceAsympLik
for the shared inversion contract. Falls back to a nonestimable result
(NA bounds) if the underlying root-finding fails or degenerates.
Gradient confidence interval for \beta_T by test
inversion of the gradient test; see
InferenceAsympLik for the shared
inversion contract. Falls back to a nonestimable result (NA
bounds) if the underlying root-finding fails or degenerates.
Usage
IncidenceLogBinomialLikelihoodSource
Details
Source list for the IncidenceLogBinomialLikelihood component composed by InferenceIncidLogBinomial.
Logistic likelihood component source
Description
Initialize inference for the logistic regression model
\mathrm{logit}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top
\gamma; see
InferenceIncidLogRegr for the
model form. Does not fit the model; the fit is deferred to the first
call to compute_estimate() or a method that requires it.
Refits the logistic model with subject/block-level weights
applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_logistic_regression_weighted_cpp, and returns the
reweighted log-odds-ratio estimate \hat\beta_T^{(w)}. Uses the
same QR column-dropping hardening and fit-reasonableness check as
compute_estimate(); a hardened-but-still-unreasonable fit
(e.g. near-perfect separation under the resampled weights) is cached
as nonestimable and returns NA.
Usage
IncidenceLogisticLikelihoodSource
Details
Source list for the IncidenceLogisticLikelihood component composed by InferenceIncidLogRegr.
Modified-Poisson likelihood component source
Description
Initialize inference for the modified Poisson model
\log E[Y_i \mid w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top
\gamma; see
InferenceIncidModifiedPoisson
for the model form and the non-robust-SE caveat. Does not fit the
model; the fit is deferred to the first call to
compute_estimate() or a method that requires it.
Fits the modified Poisson model by maximizing the Poisson
working log-likelihood on the binary response and returns the
log-risk-ratio estimate \hat\beta_T.
Wald confidence interval for \beta_T using the
model-based (non-robust) Poisson-working-likelihood standard error; see
InferenceIncidModifiedPoisson's
non-robust-SE caveat and
InferenceAsymp for the shared Wald
contract.
Two-sided Wald test of H_0: \beta_T = \code{delta}
using the model-based (non-robust) Poisson-working-likelihood standard
error; see
InferenceIncidModifiedPoisson's
non-robust-SE caveat.
Refits the modified Poisson model with subject/block-level
weights applied to the working log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights) via
fast_poisson_regression_weighted_cpp, and returns the
reweighted log-risk-ratio estimate \hat\beta_T^{(w)}. Uses the
same QR column-dropping hardening and fit-reasonableness check as
compute_estimate(); a hardened-but-still-unreasonable fit is
cached as nonestimable and returns NA.
Usage
IncidenceModifiedPoissonLikelihoodSource
Details
Source list for the IncidenceModifiedPoissonLikelihood component composed by InferenceIncidModifiedPoisson.
Probit likelihood component source
Description
Initialize inference for the probit regression model
\Phi^{-1}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma;
see InferenceIncidProbitRegr
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Refits the probit model with subject/block-level weights
applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_probit_regression_weighted_cpp, and returns the
reweighted estimate \hat\beta_T^{(w)} on the latent
standard-normal-index scale. Uses the same QR column-dropping hardening
and fit-reasonableness check as compute_estimate(); a
hardened-but-still-unreasonable fit is cached as nonestimable and
returns NA.
Usage
IncidenceProbitLikelihoodSource
Details
Source list for the IncidenceProbitLikelihood component composed by InferenceIncidProbitRegr.
Inference for A Sequential Design
Description
An abstract R6 Class that estimates, tests and provides intervals for a
treatment effect in a completed design.
This class takes a completed Design object as an input where this object
contains data for a fully completed experiment (i.e. all treatment
assignments were allocated and all responses were collected).
Active bindings
num_coresCurrent number of cores for this inference object. Defaults to the global budget unless overridden on the object.
Methods
Public methods
Inference$new()
Initialize an estimation and test object after the design is completed.
Usage
Inference$new( des_obj, verbose = FALSE, harden = TRUE, model_formula = NULL, smart_cold_start_default = NULL, seed = NULL )
Arguments
des_objA completed
Designobject whose entire n subjects are assigned and response y is recorded within.verboseWhether to print progress messages.
hardenWhether to apply robustness measures (default
TRUE). WhenTRUE, the inference methods employ defensive strategies including QR-based rank reduction of the design matrix, progressive correlation-threshold dropping, and fallback fits (e.g.\ robust survival regression, treatment-only models) to avoid crashes on ill-conditioned data. WhenFALSE, the vanilla algorithm runs on the full design matrix as supplied; any rank deficiency or convergence failure will surface as an error rather than being silently worked around. Set toFALSEwhen you want to verify that the raw model converges without intervention.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.smart_cold_start_defaultWhether to use smart cold start values by default for likelihood-based models. Explicit starts always override this object-level policy.
NULL(default) consults the global cold-start dispatch policy.seedInteger seed for reproducibility.
Inference$capabilities()
Returns the effective metadata-backed capabilities for this inference object.
Usage
Inference$capabilities()
Returns
A character vector of capability names.
Inference$supports()
Returns whether this inference object supports a metadata-backed capability.
Usage
Inference$supports(capability)
Arguments
capabilityCapability name or names.
Returns
A logical vector aligned with capability.
Inference$compute_exact_two_sided_pval_for_treatment_effect()
Computes an exact two-sided p-value. Subclasses that support exact inference override this; inference objects that do not support exact methods throw an error.
Usage
Inference$compute_exact_two_sided_pval_for_treatment_effect(...)
Arguments
...Other arguments passed to the method.
Inference$compute_exact_confidence_interval()
Computes an exact confidence interval. Subclasses that support exact inference override this; inference objects that do not support exact methods throw an error.
Usage
Inference$compute_exact_confidence_interval(...)
Arguments
...Other arguments passed to the method.
Inference$compute_asymp_two_sided_pval()
Computes an asymptotic two-sided p-value. Subclasses that support asymptotic inference override this; inference objects that do not support asymptotic methods throw an error.
Usage
Inference$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect.
Inference$compute_asymp_confidence_interval()
Computes an asymptotic confidence interval. Subclasses that support asymptotic inference override this; inference objects that do not support asymptotic methods throw an error.
Usage
Inference$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level.
Inference$compute_estimate()
Computes the treatment-effect estimate. Concrete subclasses implement the model-specific estimator, such as a fitted regression coefficient, maximum-likelihood parameter, estimating-equation solution, mean or risk contrast, survival contrast, or rank statistic. Interval, p-value, bootstrap, jackknife, and randomization methods use this method as the canonical point-estimate contract.
Usage
Inference$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
A numeric treatment estimate.
Inference$is_nonestimable()
Returns whether the most recent inference attempt explicitly marked the result as non-estimable.
Usage
Inference$is_nonestimable(type = c("any", "estimate", "se"))
Arguments
typeWhich stage to query:
"any","estimate", or"se".
Returns
A logical scalar.
Inference$get_nonestimable_reason()
Returns the reason recorded for the most recent explicit non-estimability.
Usage
Inference$get_nonestimable_reason()
Returns
A character scalar or NULL.
Inference$get_nonestimable_stage()
Returns the stage recorded for the most recent explicit non-estimability.
Usage
Inference$get_nonestimable_stage()
Returns
A character scalar or NULL.
Inference$duplicate()
Duplicate this inference object
Usage
Inference$duplicate(verbose = FALSE, make_fork_cluster = FALSE)
Arguments
verboseA flag indicating whether messages should be displayed.
make_fork_clusterWhether the duplicate should be allowed to create a fork cluster. Default FALSE.
Returns
A new Inference object with the same data
Inference$get_response()
Return the response vector used by inference extension classes.
This accessor is part of the supported extension contract for user-defined R6 inference classes. Prefer this method over direct access to private fields.
Usage
Inference$get_response()
Returns
A numeric response vector.
Inference$get_treatment()
Return the treatment-assignment vector used by inference extension classes.
This accessor is part of the supported extension contract for user-defined R6 inference classes. Treatment is encoded as 0/1.
Usage
Inference$get_treatment()
Returns
A numeric or integer 0/1 treatment vector.
Inference$get_covariates()
Return the processed covariate matrix used by inference extension classes.
This accessor returns the design object's model-matrix covariates, after
the package's missingness handling and encoding. It may be NULL if
no covariates are available.
Usage
Inference$get_covariates()
Returns
A numeric matrix of covariates or NULL.
Inference$get_analysis_data()
Return a data frame with response, treatment, censoring status, and covariates.
This accessor is the preferred data interface for user-defined R6
inference classes. It avoids reliance on private implementation fields.
The returned data frame always contains y, w, and
dead; covariate columns are appended when available.
Usage
Inference$get_analysis_data()
Returns
A data frame suitable for user-defined model fitting.
Inference$get_design_object()
Return the completed design object backing this inference object.
This accessor is part of the supported extension contract. Extension
classes should use this method instead of private$des_obj.
Usage
Inference$get_design_object()
Returns
The completed Design object.
Inference$get_response_type()
Return the response type for the backing design.
Usage
Inference$get_response_type()
Returns
A character scalar such as "continuous",
"incidence", "proportion", "count",
"survival", or "ordinal".
Inference$get_model_formula()
Return the model formula used for covariate adjustment.
Usage
Inference$get_model_formula()
Returns
A formula object or NULL.
Inference$set_optimization_alg()
Set the optimizer used by likelihood-based inference implementations.
Usage
Inference$set_optimization_alg( optimization_alg = NULL, allow_irls = private$optimization_alg_allow_irls, default = private$optimization_alg_default )
Arguments
optimization_algThe optimizer name. Valid values are configured by the concrete inference class.
allow_irlsWhether to allow IRLS (Iteratively Reweighted Least Squares) as a fallback or primary optimization algorithm.
defaultThe default optimizer to use if none is specified.
Returns
Invisibly returns self.
Inference$get_optimization_alg()
Return the optimizer used by likelihood-based inference implementations.
Usage
Inference$get_optimization_alg()
Inference$set_seed()
Set the seed for reproducibility.
Usage
Inference$set_seed(seed)
Arguments
seedInteger seed for reproducibility.
Inference$clone()
The objects of this class are cloneable with this method.
Usage
Inference$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Abstract Conditional Logistic GLMM Inference
Description
Fits one likelihood with a conditional-logistic contribution from discordant matched pairs and a random-intercept logistic GLMM contribution from concordant matched pairs and reservoir subjects.
Super class
Inference -> InferenceAbstractKKCondLogitGLMM
Methods
Public methods
-
InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights() -
InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval() -
InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAbstractKKCondLogitGLMM$new()
Initialize KK conditional-logit GLMM incidence
inference, validate the binary response, and prepare the matched-pair
conditional likelihood and reservoir mixed-model components. See
InferenceAbstractKKCondLogitGLMM.
Usage
InferenceAbstractKKCondLogitGLMM$new( des_obj, model_formula = NULL, max_abs_reasonable_coef = 50, max_abs_reasonable_se = 10, max_abs_log_sigma = 8, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with an incidence or proportion response.model_formulaOptional formula for covariate adjustment.
max_abs_reasonable_coefCap for reasonable coefficient estimates.
max_abs_reasonable_seCap for reasonable treatment standard errors.
max_abs_log_sigmaCap for reasonable log random effect variance.
verboseLogical. Whether to print progress messages.
smart_cold_start_defaultLogical. Whether to use smart starting values for the optimizer.
optimization_algCharacter. Optimization algorithm (default "lbfgs").
InferenceAbstractKKCondLogitGLMM$compute_estimate()
Computes the class-specific treatment-effect estimate; see
Inference.
Usage
InferenceAbstractKKCondLogitGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyLogical. If TRUE, skip variance component calculations.
InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights()
Recomputes the class-specific treatment estimate for a bootstrap sample; see
InferenceNonParamBootstrap.
Usage
InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric vector. Row weights for bootstrap.
estimate_onlyLogical. If TRUE, skip variance component calculations.
InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval()
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Usage
InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaNumeric. Significance level (default 0.05).
InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval()
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNumeric. Null treatment effect value (default 0).
InferenceAbstractKKCondLogitGLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAbstractKKCondLogitGLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Abstract class for all-subject marginal incidence inference in KK designs
Description
Abstract class for all-subject marginal incidence inference in KK designs
Super class
Inference -> InferenceAbstractKKMarginalIncid
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAbstractKKMarginalIncid$new()
Initialize the shared KK marginal-incidence inference base,
validate the binary matched/reservoir design, and prepare caches used by
InferenceAbstractKKMarginalIncid.
Usage
InferenceAbstractKKMarginalIncid$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseA flag indicating whether messages should be displayed.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceAbstractKKMarginalIncid$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAbstractKKMarginalIncid$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Abstract class for all-subject modified-Poisson inference in KK designs
Description
Abstract class for all-subject modified-Poisson inference in KK designs
Super classes
Inference -> InferenceAbstractKKMarginalIncid -> InferenceAbstractKKModifiedPoisson
Methods
Public methods
-
InferenceAbstractKKModifiedPoisson$compute_estimate_with_bootstrap_weights() -
InferenceAbstractKKModifiedPoisson$compute_asymp_confidence_interval() -
InferenceAbstractKKModifiedPoisson$compute_asymp_two_sided_pval()
+ inherited public methods from InferenceAbstractKKMarginalIncid
InferenceAbstractKKMarginalIncid$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_gradient_confidence_interval()InferenceAbstractKKMarginalIncid$compute_gradient_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_jackknife_bias_estimate()InferenceAbstractKKMarginalIncid$compute_jackknife_estimate()InferenceAbstractKKMarginalIncid$compute_jackknife_std_error()InferenceAbstractKKMarginalIncid$compute_jackknife_wald_confidence_interval()InferenceAbstractKKMarginalIncid$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_param_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_param_bootstrap_estimate()InferenceAbstractKKMarginalIncid$compute_param_bootstrap_pval()InferenceAbstractKKMarginalIncid$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_rand_confidence_interval()InferenceAbstractKKMarginalIncid$compute_rand_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_score_confidence_interval()InferenceAbstractKKMarginalIncid$compute_score_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_subsampling_confidence_interval()InferenceAbstractKKMarginalIncid$compute_subsampling_sensitivity()InferenceAbstractKKMarginalIncid$compute_subsampling_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_wald_confidence_interval()InferenceAbstractKKMarginalIncid$compute_wald_two_sided_pval()InferenceAbstractKKMarginalIncid$get_information_preference()InferenceAbstractKKMarginalIncid$get_information_source_used()InferenceAbstractKKMarginalIncid$get_last_param_bootstrap_diagnostics()InferenceAbstractKKMarginalIncid$get_last_param_bootstrap_estimate_diagnostics()InferenceAbstractKKMarginalIncid$get_mod()InferenceAbstractKKMarginalIncid$get_summary()InferenceAbstractKKMarginalIncid$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKMarginalIncid$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKMarginalIncid$get_supported_bootstrap_ci_types()InferenceAbstractKKMarginalIncid$get_supported_bootstrap_pval_types()InferenceAbstractKKMarginalIncid$get_supported_information_preferences()InferenceAbstractKKMarginalIncid$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKMarginalIncid$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKMarginalIncid$get_supported_testing_types()InferenceAbstractKKMarginalIncid$get_testing_type()InferenceAbstractKKMarginalIncid$initialize()InferenceAbstractKKMarginalIncid$select_optimal_b_subsampling()InferenceAbstractKKMarginalIncid$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKMarginalIncid$set_information_preference()InferenceAbstractKKMarginalIncid$set_testing_type()InferenceAbstractKKMarginalIncid$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAbstractKKModifiedPoisson$compute_estimate()
Compute the KK marginal incidence treatment-effect estimate
using the class-specific marginal risk-difference or risk-ratio
estimator and cache it for related
InferenceAbstractKKMarginalIncid
methods.
Usage
InferenceAbstractKKModifiedPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyLogical. If TRUE, skip variance component calculations.
InferenceAbstractKKModifiedPoisson$compute_estimate_with_bootstrap_weights()
Recomputes the KK marginal incidence estimate under Bayesian-bootstrap weights.
Usage
InferenceAbstractKKModifiedPoisson$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric vector. Row weights for bootstrap.
estimate_onlyLogical. If TRUE, skip variance component calculations.
InferenceAbstractKKModifiedPoisson$compute_asymp_confidence_interval()
Compute the KK marginal incidence asymptotic confidence
interval using the cached marginal effect and design-aware standard
error. See InferenceAsymp.
Usage
InferenceAbstractKKModifiedPoisson$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaNumeric. Significance level (default 0.05).
InferenceAbstractKKModifiedPoisson$compute_asymp_two_sided_pval()
Compute the KK marginal incidence asymptotic two-sided
p-value using the cached marginal effect and design-aware standard
error. See InferenceAsymp.
Usage
InferenceAbstractKKModifiedPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNumeric. Null treatment effect value (default 0).
InferenceAbstractKKModifiedPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAbstractKKModifiedPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Abstract class for ordinal CLMM-based Inference in KK designs
Description
Abstract class for ordinal CLMM-based Inference in KK designs
Super class
Inference -> InferenceAbstractKKOrdinalCLMM
Methods
Public methods
-
InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights() -
InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval() -
InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAbstractKKOrdinalCLMM$new()
Initialize KK cumulative-link mixed-model inference for
ordinal responses, validate the matched design, and prepare the ordinal
likelihood used by InferenceAbstractKKOrdinalCLMM.
Usage
InferenceAbstractKKOrdinalCLMM$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, harden = TRUE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment.
use_rcppLogical. If
TRUE(default), use the internal Rcpp implementation (no external packages required). SetFALSEto fall back to ordinal::clmm.verboseA flag indicating whether messages should be displayed.
hardenWhether to apply robustness measures.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceAbstractKKOrdinalCLMM$compute_estimate()
Compute the ordinal CLMM treatment-effect estimate by fitting
the cumulative-link mixed model and caching the treatment coefficient for
related InferenceAsymp methods.
Usage
InferenceAbstractKKOrdinalCLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyLogical. If TRUE, skip variance component calculations.
InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights()
Recomputes the KK ordinal CLMM treatment estimate under Bayesian-bootstrap weights.
Usage
InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric vector. Row weights for bootstrap.
estimate_onlyLogical. If TRUE, skip variance component calculations.
InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval()
Compute the ordinal CLMM asymptotic confidence interval for
the treatment coefficient using the fitted-model standard error. See
InferenceAsymp.
Usage
InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaNumeric. Significance level (default 0.05).
InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()
Compute the ordinal CLMM asymptotic two-sided p-value for the
treatment coefficient using the fitted-model standard error. See
InferenceAsymp.
Usage
InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNumeric. Null treatment effect value (default 0).
InferenceAbstractKKOrdinalCLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAbstractKKOrdinalCLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Abstract mixin: Zhang combined randomisation CI for quantile regression
Description
Provides compute_rand_confidence_interval()
via Zhang's combined test-inversion method for both Bernoulli (m = 0,
all subjects in the reservoir) and KK matching-on-the-fly designs
(m > 0).
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable -> InferenceAsympLik -> InferenceKKPassThroughCompoundNoParamBootstrap -> InferenceAbstractQuantileRandCI
Methods
Public methods
+ inherited public methods from InferenceKKPassThroughCompoundNoParamBootstrap
InferenceKKPassThroughCompoundNoParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceKKPassThroughCompoundNoParamBootstrap$compute_estimate_with_bootstrap_weights()
+ inherited public methods from InferenceAsympLik
InferenceAsympLik$compute_asymp_confidence_interval()InferenceAsympLik$compute_asymp_two_sided_pval()InferenceAsympLik$compute_gradient_confidence_interval()InferenceAsympLik$compute_gradient_two_sided_pval()InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval()InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval()InferenceAsympLik$compute_lik_ratio_confidence_interval()InferenceAsympLik$compute_lik_ratio_two_sided_pval()InferenceAsympLik$compute_score_confidence_interval()InferenceAsympLik$compute_score_two_sided_pval()InferenceAsympLik$get_information_preference()InferenceAsympLik$get_information_source_used()InferenceAsympLik$get_supported_information_preferences()InferenceAsympLik$get_supported_testing_types()InferenceAsympLik$get_testing_type()InferenceAsympLik$set_information_preference()InferenceAsympLik$set_testing_type()
+ inherited public methods from InferenceMLEorKMSummaryTable
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()InferenceJackknife$compute_jackknife_bias_estimate()InferenceJackknife$compute_jackknife_estimate()InferenceJackknife$compute_jackknife_std_error()InferenceJackknife$compute_jackknife_wald_confidence_interval()InferenceJackknife$compute_jackknife_wald_two_sided_pval()
+ inherited public methods from InferenceBayesianBootstrap
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAbstractQuantileRandCI$new()
Initialize the inference object.
Usage
InferenceAbstractQuantileRandCI$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA DesignSeqOneByOne object.
model_formulaOptional formula for covariate adjustment.
verboseWhether to print messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceAbstractQuantileRandCI$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAbstractQuantileRandCI$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Mean-Difference IVWC Inference for KK Matching-on-the-Fly Designs
Description
Fits a compound (inverse-variance-weighted combination, "IVWC") mean-difference
estimator of the treatment effect for continuous responses under a
DesignSeqOneByOne-family KK matching-on-the-fly
design (see DesignSeqOneByOneKK14 and
DesignSeqOneByOneKK21). Such a design
produces two structurally different kinds of subjects: subjects successfully
matched into pairs during the sequential design, and unmatched "reservoir"
subjects randomized independently. This estimator combines both:
\hat\beta_T = w^* \bar d + (1 - w^*)\, \bar r, \qquad
w^* = \frac{\widehat{\mathrm{Var}}(\bar r)}{\widehat{\mathrm{Var}}(\bar r) +
\widehat{\mathrm{Var}}(\bar d)},
where \bar d is the mean within-pair (treated minus control) difference
among matched subjects and \bar r is the treated-minus-control difference
in means among reservoir subjects, weighted inversely by their estimated
variances (see $compute_asymp_confidence_interval() for the full
variance formula and the fallback behavior when only one of the two
sub-estimates is usable). Inference is Wald-only: this class has no
likelihood tier (likelihood_tier = "none") and provides asymptotic
Wald, randomization, and bootstrap (including Bayesian bootstrap) confidence
intervals and p-values, but no score/likelihood-ratio/gradient tests.
Initialize KK IVWC mean-difference inference.
Computes the compound IVWC (inverse-variance-weighted
compound) mean-difference point estimate \hat\beta_T: the
inverse-variance-weighted combination w^* \bar d + (1-w^*)\, \bar
r of the matched-pair mean within-pair difference \bar d and the
reservoir treated-minus-control mean difference \bar r, falling
back to whichever of the two is usable if the other is not (see
$compute_asymp_confidence_interval() for the full weighting
formula and usability conditions).
Computes a 1-\alpha level frequentist confidence interval
for the compound IVWC (inverse-variance-weighted compound) mean-difference
estimator \hat\beta_T.
Computes a two-sided Wald p-value for the compound
IVWC mean-difference estimator \hat\beta_T testing
H_0: \beta_T = \code{delta}, using the same asymptotically-normal
point estimate and standard error (z = (\hat\beta_T -
\code{delta})/\widehat{\mathrm{SE}}(\hat\beta_T)) that
$compute_asymp_confidence_interval() inverts to form its interval
— see that method's documentation for the full inverse-variance-weighted
combination formula. This class has no likelihood tier
(likelihood_tier = "none"), so no score, likelihood-ratio, or
gradient test is available here; this is a plain Wald test, not a
likelihood-backed one.
Details
The point estimate combines two sub-estimates depending on which are
usable: the mean within-pair difference among matched subjects,
\bar d, with estimated variance \widehat{\mathrm{Var}}(\bar
d), and the treated-minus-control difference in means among reservoir
(unmatched) subjects, \bar r, with estimated variance
\widehat{\mathrm{Var}}(\bar r). When both are usable (at least 2
matched pairs and at least 2 treated/2 control reservoir subjects, with
finite positive variance estimates), they are combined by classical
inverse-variance weighting,
\hat\beta_T = w^* \bar d + (1 - w^*)\, \bar r, \qquad w^* =
\frac{\widehat{\mathrm{Var}}(\bar r)}{\widehat{\mathrm{Var}}(\bar r) +
\widehat{\mathrm{Var}}(\bar d)},
with combined variance the standard inverse-variance-pooled form
\widehat{\mathrm{Var}}(\hat\beta_T) = \left(\widehat{\mathrm{Var}}(\bar
r)^{-1} + \widehat{\mathrm{Var}}(\bar d)^{-1}\right)^{-1} =
\widehat{\mathrm{Var}}(\bar r)\,\widehat{\mathrm{Var}}(\bar d) \big/
\left(\widehat{\mathrm{Var}}(\bar r) + \widehat{\mathrm{Var}}(\bar
d)\right). If only one of the two sub-estimates is usable (e.g. the
reservoir is empty or degenerate, or no pairs matched), \hat\beta_T
and its variance fall back to that sub-estimate alone. The compound
estimator is treated as asymptotically normal, so the interval is
\hat\beta_T \pm z_{1-\alpha/2}\sqrt{\widehat{\mathrm{Var}}(\hat\beta_T)}
(or a t-based critical value, depending on
private$compute_z_or_t_ci_from_s_and_df's degrees-of-freedom
resolution).
Value
The setting-appropriate (see description) numeric estimate of the treatment effect
A (1 - alpha)-sized frequentist confidence interval for the treatment effect
The approximate frequentist p-value
Legacy status
Legacy class. Not fully tested in
comprehensive_tests.R; prefer a more actively maintained KK
continuous-response inference class (e.g.
InferenceContinKKOLSIVWC) for new
analyses unless this specific unadjusted mean-difference estimator is
required.
Super class
Inference -> InferenceAllKKMeanDiffIVWC
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAllKKMeanDiffIVWC$new()
Usage
InferenceAllKKMeanDiffIVWC$new( des_obj, verbose = FALSE, harden = TRUE, model_formula = NULL, smart_cold_start_default = NULL )
Arguments
des_objA KK matching-on-the-fly design object.
verboseWhether to print progress messages.
hardenWhether to use hardened model-matrix fitting.
model_formulaOptional formula for covariate adjustment.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceAllKKMeanDiffIVWC$compute_estimate()
Usage
InferenceAllKKMeanDiffIVWC$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf
TRUE, compute only the point estimate\hat\beta_Tand skip the variance-component computations needed for confidence intervals or p-values (faster when only the point estimate is needed).
InferenceAllKKMeanDiffIVWC$compute_asymp_confidence_interval()
Usage
InferenceAllKKMeanDiffIVWC$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
InferenceAllKKMeanDiffIVWC$compute_asymp_two_sided_pval()
Usage
InferenceAllKKMeanDiffIVWC$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. For any treatment effect at all this is set to zero (the default).
InferenceAllKKMeanDiffIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAllKKMeanDiffIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this estimator targets, and for the inverse-variance combination of matched-pair and reservoir estimates.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))
seq_des_inf = InferenceAllKKMeanDiffIVWC$
new(seq_des)
seq_des_inf$compute_estimate()
seq_des_inf$compute_asymp_confidence_interval()
seq_des_inf$compute_asymp_two_sided_pval()
Non-parametric Wilcoxon-based Compound Inference for KK Matching-on-the-Fly Designs
Description
Fits a non-parametric, rank-based compound (inverse-variance-weighted, IVWC)
estimator of the treatment effect under a
DesignSeqOneByOne-family KK
matching-on-the-fly design (see
DesignSeqOneByOneKK14). For matched
pairs, the sub-estimate \hat\beta_m is the Hodges-Lehmann
estimate from a Wilcoxon signed-rank test on the within-pair differences (the
median of the Walsh averages (d_i+d_j)/2); for reservoir (unmatched)
subjects, \hat\beta_r is the Hodges-Lehmann estimate from a Wilcoxon
rank-sum (Mann-Whitney U) test on treated-vs-control reservoir responses
(the median of all pairwise differences). The two are combined by classical
inverse-variance weighting,
\hat\beta_T = w^* \hat\beta_m + (1-w^*)\, \hat\beta_r, \qquad
w^* = \frac{\widehat{\mathrm{Var}}(\hat\beta_r)}{\widehat{\mathrm{Var}}(\hat\beta_r)
+ \widehat{\mathrm{Var}}(\hat\beta_m)},
with variance the standard inverse-variance-pooled form (see
$compute_estimate()'s method-level documentation for the full formula
and fallback behavior when only one sub-estimate is usable). Because it is
built from Hodges-Lehmann/Wilcoxon estimators rather than sample means, this
method is robust to outliers and does not assume a specific parametric
distribution for the response — but it does not currently support censored
survival data or incidence (binary) responses (see $initialize()), and
its jackknife methods all report explicit non-estimability rather than
computing a (statistically unreliable) delete-1 jackknife of the
Hodges-Lehmann functional.
Override to avoid O(n^2) per-resample HL computation during the bootstrap warm-start inside compute_rand_confidence_interval. The asymptotic MLE CI is a perfectly adequate starting bound for the bisection and is computed in O(1).
Initialize KK inverse-variance combined Wilcoxon inference
and prepare matched/reservoir rank-based components used by
InferenceAllKKWilcoxIVWC.
Requires des_obj to be a KK matching-on-the-fly-capable design
(des_obj$is_a_kk_matching_capable()); errors otherwise. Also
rejects response_type = "incidence" (a message-and-error
recommends a compound mean-difference or conditional-logistic estimator
instead — rank-based methods are not well suited to binary outcomes)
and rejects censored survival data (recommends a restricted-mean or
Cox-based method instead, since this estimator has no censoring
handling). Legal response_type values are "continuous",
"count", "proportion", "survival" (uncensored
only), and "ordinal".
Returns the estimated treatment effect: an inverse-variance-weighted compound (IVWC) of two Hodges-Lehmann median-shift estimates.
Computes a 1-\alpha level confidence interval for the
compound Hodges-Lehmann treatment effect estimator. Although each
sub-estimate is itself derived from a non-parametric rank test, the
inverse-variance-weighted combination \hat\beta_T (see
$compute_estimate() for the full formula) is treated as
asymptotically normal, so the interval is
\hat\beta_T \pm z_{1-\alpha/2}\sqrt{\widehat{\mathrm{Var}}(\hat\beta_T)}
(or a t-based critical value, depending on
private$compute_z_or_t_ci_from_s_and_df's degrees-of-freedom
resolution).
Compute the KK Wilcoxon compound two-sided p-value testing
H_0: \beta_T = \code{delta}, from the same asymptotically-normal
compound estimate/variance (z = (\hat\beta_T -
\code{delta})/\widehat{\mathrm{SE}}(\hat\beta_T)) that
$compute_asymp_confidence_interval() inverts to form its
interval — see that method's documentation, and
$compute_estimate(), for the compound Hodges-Lehmann estimator's
full formula. Only delta = 0 is currently supported: a non-zero
null shift raises an error (when assertions are enabled) rather than
testing it, because the underlying Wilcoxon tests' null-shift handling
has not been extended to the compound combined estimator. See related
simple Wilcoxon behavior in
InferenceAllSimpleWilcox.
Reports the jackknife point-estimate as explicitly
non-estimable for this compound Hodges-Lehmann estimator, rather than
computing a leave-one-out jackknife. Deletion-based (jackknife)
resampling of a Hodges-Lehmann/Wilcoxon-derived statistic is known to
behave poorly — the median-of-Walsh-averages functional is not smooth
enough for the delete-1 jackknife's linear-approximation machinery to
be reliable at the small matched-pair/reservoir sample sizes typical of
KK designs, and combining two already-jackknife-unstable sub-estimates
compounds the problem. This method exists purely to record that
unavailability (via private$cache_nonestimable_estimate()) rather
than silently returning a misleading number; see
InferenceJackknife for the shared
jackknife contract this method participates in.
Reports the jackknife bias-correction estimate as
non-estimable for this Wilcoxon compound estimator, for the same
reason as $compute_jackknife_estimate() (the Hodges-Lehmann
functional is not smooth enough for the delete-1 jackknife); see
InferenceJackknife for the shared
jackknife contract.
Reports the jackknife standard error as non-estimable for
this Wilcoxon compound estimator, for the same reason as
$compute_jackknife_estimate(); see
InferenceJackknife for the shared
jackknife contract.
Reports the jackknife-Wald p-value as non-estimable here,
for the same reason as $compute_jackknife_estimate(); see
InferenceJackknife.
Reports the jackknife-Wald confidence interval as
non-estimable here, for the same reason as
$compute_jackknife_estimate(); see
InferenceJackknife.
Details
For matched pairs, \hat\beta_m is the Hodges-Lehmann
estimate from a Wilcoxon signed-rank test on the within-pair
differences (stats::wilcox.test(diffs, conf.int = TRUE)'s
estimate, the median of the Walsh averages
(d_i + d_j)/2), with variance estimated as the sample variance of
those Walsh averages divided by the number of pairs m. For
reservoir (unmatched) subjects, \hat\beta_r is the
Hodges-Lehmann estimate from a Wilcoxon rank-sum test between
treated and control reservoir responses (median of all pairwise
differences y_{T,i} - y_{C,j}), with an analogous
pairwise-difference-variance-based estimate. When both sub-estimates
are usable, the compound estimate is the inverse-variance-weighted
combination
\hat\beta_T = w^* \hat\beta_m + (1-w^*)\, \hat\beta_r, \qquad
w^* = \frac{\widehat{\mathrm{Var}}(\hat\beta_r)}{\widehat{\mathrm{Var}}(\hat\beta_r)
+ \widehat{\mathrm{Var}}(\hat\beta_m)},
with combined variance \widehat{\mathrm{Var}}(\hat\beta_r)\,
\widehat{\mathrm{Var}}(\hat\beta_m) / (\widehat{\mathrm{Var}}(\hat\beta_r) +
\widehat{\mathrm{Var}}(\hat\beta_m)) — the same combination scheme as
InferenceAllKKMeanDiffIVWC,
but applied to rank-based rather than mean-based sub-estimates. If
only one sub-estimate is usable (e.g. no matched pairs, or a degenerate
reservoir), \hat\beta_T falls back to that sub-estimate alone.
Super class
Inference -> InferenceAllKKWilcoxIVWC
Methods
Public methods
-
InferenceAllKKWilcoxIVWC$compute_bootstrap_confidence_interval() -
InferenceAllKKWilcoxIVWC$compute_asymp_confidence_interval() -
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_two_sided_pval() -
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAllKKWilcoxIVWC$compute_bootstrap_confidence_interval()
Usage
InferenceAllKKWilcoxIVWC$compute_bootstrap_confidence_interval( alpha = 0.05, ... )
Arguments
alphaThe confidence level. Default is 0.05.
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.alphaSignificance level. Default 0.05.
...Additional arguments passed to super.
InferenceAllKKWilcoxIVWC$new()
Usage
InferenceAllKKWilcoxIVWC$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA DesignSeqOneByOne object (must be a KK design).
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceAllKKWilcoxIVWC$compute_estimate()
Usage
InferenceAllKKWilcoxIVWC$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceAllKKWilcoxIVWC$compute_asymp_confidence_interval()
Usage
InferenceAllKKWilcoxIVWC$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level. Default is 0.05.
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.alphaSignificance level. Default 0.05.
InferenceAllKKWilcoxIVWC$compute_asymp_two_sided_pval()
Usage
InferenceAllKKWilcoxIVWC$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. For any treatment effect at all this is set to zero (the default).
deltaNull treatment-effect value. Default 0.
InferenceAllKKWilcoxIVWC$compute_jackknife_estimate()
Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllKKWilcoxIVWC$compute_jackknife_bias_estimate()
Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllKKWilcoxIVWC$compute_jackknife_std_error()
Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_std_error(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_two_sided_pval()
Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_two_sided_pval( delta = 0, unit = "auto" )
Arguments
deltaThe null difference to test against. For any treatment effect at all this is set to zero (the default).
deltaNull treatment-effect value. Default 0.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_confidence_interval()
Usage
InferenceAllKKWilcoxIVWC$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaThe confidence level. Default is 0.05.
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.alphaSignificance level. Default 0.05.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllKKWilcoxIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAllKKWilcoxIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Hodges, J. L., and Lehmann, E. L. (1963). "Estimates of Location Based on Rank Tests." The Annals of Mathematical Statistics, 34(2), 598-611, doi:10.1214/aoms/1177704172, for the Hodges-Lehmann estimator underlying both sub-estimates; Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design and the inverse-variance combination of matched-pair and reservoir estimates.
Legacy class. Not fully tested in comprehensive_tests.R.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceAllKKWilcoxIVWC$new(seq_des)
inf$compute_estimate()
Simple Mean-Difference Inference for Continuous Responses
Description
Fits the simplest possible treatment-effect estimator for a continuous
response: the unadjusted difference in sample means between the treated and
control arms, \hat\beta_T = \bar y_T - \bar y_C, with no covariate
adjustment. Inference is by Welch's unequal-variance t-test: standard
error \sqrt{s_T^2/n_T + s_C^2/n_C} (per-arm sample variances, not
pooled) with Satterthwaite-Welch degrees of freedom — see
$compute_asymp_confidence_interval() for the exact formula. This class
has no likelihood tier (likelihood_tier = "none") and provides
asymptotic Wald, randomization, and bootstrap (including Bayesian bootstrap)
confidence intervals and p-values. Warm starts are disabled for this class,
since the simple mean difference is a closed-form estimator (no iterative fit
to warm-start).
Initialize a simple mean-difference inference object.
Computes a 1-\alpha level confidence interval for the
simple (unadjusted) mean-difference treatment effect
\hat\beta_T = \bar y_T - \bar y_C, using Welch's
unequal-variance formula: standard error
\widehat{\mathrm{SE}}(\hat\beta_T) = \sqrt{s_T^2/n_T + s_C^2/n_C}
(sample variances s_T^2, s_C^2 computed separately per arm,
not pooled) with Satterthwaite-Welch degrees of freedom
\mathrm{df} = (s_T^2/n_T + s_C^2/n_C)^2 \big/ \left(\frac{(s_T^2/n_T)^2}{n_T-1}
+ \frac{(s_C^2/n_C)^2}{n_C-1}\right); the interval is
\hat\beta_T \pm t_{\mathrm{df}, 1-\alpha/2}\,\widehat{\mathrm{SE}}(\hat\beta_T).
Requires at least 2 observations per arm; otherwise the standard error
and interval are NA. See
InferenceAsymp for the shared
asymptotic confidence-interval contract this delegates to.
Computes a two-sided Welch's t-test p-value testing
H_0: \beta_T = \code{delta}, from the same Welch
unequal-variance standard error and Satterthwaite-Welch degrees of
freedom used by $compute_asymp_confidence_interval() — see that
method's documentation for the full formula. See
InferenceAsymp for the shared
asymptotic two-sided p-value contract this delegates to.
Computes the simple (unadjusted) mean-difference point
estimate \hat\beta_T = \bar y_T - \bar y_C, the difference in
sample means between the treated and control arms. NA if
either arm has zero observations. See
InferenceMLEorKMSummaryTable
for the shared estimate-contract this participates in.
Recomputes the simple mean-difference estimate under
subject/block bootstrap weights (used by the Bayesian bootstrap and
related weighted-resampling machinery — see
InferenceNonParamBootstrap).
The weighted point estimate is \hat\beta_T = \bar y_T^w - \bar
y_C^w, weighted arm means \bar y_T^w = \sum_i r_i y_i \mathbb{1}[w_i=1]
/ \sum_i r_i \mathbb{1}[w_i=1] (and analogously for control), where
r_i are the expanded row weights. Unless estimate_only =
TRUE, the standard error uses a weighted, effective-sample-size
Welch formula: n_{\mathrm{eff}} = (\sum r_i)^2 / \sum r_i^2
(the usual Kish effective-sample-size correction for unequal weights)
in place of the raw n in both the per-arm weighted variance
denominator and the Satterthwaite-Welch degrees-of-freedom formula (see
$compute_asymp_confidence_interval() for the unweighted version
of the same formula). Rows with non-finite or non-positive weight, or a
non-finite response, are dropped before computing; if no rows survive,
returns NA with all cached variance components set to NA.
Value
A two-sided p-value.
The setting-appropriate (see description) numeric estimate of the treatment effect
Super class
Inference -> InferenceAllSimpleAverageDiff
Methods
Public methods
-
InferenceAllSimpleAverageDiff$compute_asymp_confidence_interval() -
InferenceAllSimpleAverageDiff$compute_asymp_two_sided_pval() -
InferenceAllSimpleAverageDiff$compute_estimate_with_bootstrap_weights()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAllSimpleAverageDiff$compute_rand_two_sided_pval()
Uses the randomization-CI layer's two-sided p-value contract
(InferenceRandCI's version, not InferenceRand's): for
incidence responses this dispatches to the Zhang exact randomization
test where applicable rather than refusing outright, matching this
class's pre-migration old-ladder behavior (see
InferenceIncidRiskDiff's identical rationale). Previously bound
to InferenceRand's version instead, which silently regressed
Zhang dispatch after migration – see
inference_all_abstract_rand_ci.R's
compute_rand_two_sided_pval for why it's now safe to splice
this in outside the old inheritance chain.
Usage
InferenceAllSimpleAverageDiff$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, type = NULL, args_for_type = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors.
deltaNull treatment effect value.
deltaNull treatment effect value.
transform_responsesResponse transformation to apply during the test. For survival responses the default
"log"multiplies the recorded times of the units treated under each reference allocation bye^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); seecompute_rand_confidence_interval()for the assumptions.na.rmWhether to remove non-finite simulated statistics.
show_progressWhether to show progress.
permutationsOptional pre-generated assignment draws.
typeOptional incidence-specific exact randomization type.
args_for_typeOptional arguments keyed by
type.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceAllSimpleAverageDiff$new()
Usage
InferenceAllSimpleAverageDiff$new( des_obj, model_formula = NULL, verbose = FALSE, max_resample_attempts = 50L, smart_cold_start_default = NULL )
Arguments
des_objA DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages. Default
FALSE.max_resample_attemptsMaximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. If all attempts fail the replicate is recorded as
NA, silently reducing the effectiveB. Must be a positive integer. Default50L.smart_cold_start_defaultWhether to use smart cold start values.
InferenceAllSimpleAverageDiff$compute_asymp_confidence_interval()
Usage
InferenceAllSimpleAverageDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaConfidence level.
InferenceAllSimpleAverageDiff$compute_asymp_two_sided_pval()
Usage
InferenceAllSimpleAverageDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect value.
deltaNull treatment effect value.
InferenceAllSimpleAverageDiff$compute_estimate()
Usage
InferenceAllSimpleAverageDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
estimate_onlyIf TRUE, skip variance calculations.
InferenceAllSimpleAverageDiff$compute_estimate_with_bootstrap_weights()
Usage
InferenceAllSimpleAverageDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsRow weights for the bootstrap sample.
estimate_onlyIf TRUE, skip variance component calculations.
estimate_onlyIf TRUE, skip variance calculations.
InferenceAllSimpleAverageDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAllSimpleAverageDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Welch, B. L. (1947). "The Generalization of 'Student's' Problem when Several Different Population Variances are Involved." Biometrika, 34(1-2), 28-35, doi:10.1093/biomet/34.1-2.28, for the unequal-variance t-test and its Satterthwaite-Welch degrees-of-freedom approximation used here.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))
seq_des_inf = InferenceAllSimpleAverageDiff$new(seq_des)
seq_des_inf$compute_estimate()
seq_des_inf$compute_asymp_confidence_interval()
seq_des_inf$compute_asymp_two_sided_pval()
Simple Mean-Difference Inference with Pooled Variance
Description
Fits the same unadjusted mean-difference point estimate as
InferenceAllSimpleAverageDiff,
\hat\beta_T = \bar y_T - \bar y_C, but performs inference via the
classical pooled equal-variance Student's t-test instead of
Welch's unequal-variance version: pooled variance s_p^2 =
\left((n_T-1)s_T^2 + (n_C-1)s_C^2\right)/(n_T+n_C-2), standard error
s_p\sqrt{1/n_T + 1/n_C}, and exact degrees of freedom n_T+n_C-2
— see $compute_asymp_confidence_interval() for the full formula.
This assumes the two arms have equal population variance; prefer
InferenceAllSimpleAverageDiff
when that assumption is doubtful, since the pooled estimator's nominal
coverage degrades under heteroskedasticity with unequal arm sizes. This
class does not support censored survival data (enforced at construction).
This class has no likelihood tier (likelihood_tier = "none") and
provides asymptotic Wald, randomization, and bootstrap (including Bayesian
bootstrap) confidence intervals and p-values. Warm starts are disabled for
this class, since the simple mean difference is a closed-form estimator
(no iterative fit to warm-start).
Initialize simple pooled-variance mean-difference inference
for continuous responses and prepare the pooled standard-error
calculation used by
InferenceAllSimpleMeanDiffPooledVar.
Disables warm starts (closed-form estimator) and asserts
des_obj has no censored observations (unsupported by this
class).
Computes a 1-\alpha level confidence interval for the
simple (unadjusted) mean-difference treatment effect
\hat\beta_T = \bar y_T - \bar y_C, using the classical
pooled equal-variance Student's t-test formula (unlike
InferenceAllSimpleAverageDiff's
Welch unequal-variance version): the pooled variance estimate
s_p^2 = \left((n_T-1)s_T^2 + (n_C-1)s_C^2\right) / (n_T+n_C-2)
gives standard error \widehat{\mathrm{SE}}(\hat\beta_T) =
s_p\sqrt{1/n_T + 1/n_C} with exact degrees of freedom n_T + n_C -
2; the interval is \hat\beta_T \pm t_{\mathrm{df}, 1-\alpha/2}\,
\widehat{\mathrm{SE}}(\hat\beta_T). Assumes equal population
variances in the two arms — use
InferenceAllSimpleAverageDiff
instead when that assumption is doubtful. Requires at least 2
observations per arm; otherwise returns c(NA, NA). See
InferenceAsymp for the shared
asymptotic confidence-interval contract this participates in.
Computes a two-sided pooled-variance Student's t-test
p-value testing H_0: \beta_T = \code{delta}, from the same
pooled standard error and exact n_T+n_C-2 degrees of freedom
used by $compute_asymp_confidence_interval() — see that
method's documentation for the full formula. See
InferenceAsymp for the shared
asymptotic two-sided p-value contract this participates in.
Value
A new InferenceAllSimpleMeanDiffPooledVar object.
Super class
Inference -> InferenceAllSimpleMeanDiffPooledVar
Methods
Public methods
-
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_confidence_interval() -
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAllSimpleMeanDiffPooledVar$new()
Usage
InferenceAllSimpleMeanDiffPooledVar$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_confidence_interval()
Usage
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaConfidence level.
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_two_sided_pval()
Usage
InferenceAllSimpleMeanDiffPooledVar$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect value.
InferenceAllSimpleMeanDiffPooledVar$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAllSimpleMeanDiffPooledVar$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Student [Gosset, W. S.] (1908). "The Probable Error of a Mean." Biometrika, 6(1), 1-25, doi:10.1093/biomet/6.1.1, for the pooled-variance two-sample t-test used here.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceAllSimpleMeanDiffPooledVar$new(seq_des)
inf$compute_estimate()
Simple Wilcoxon Rank-Sum (Hodges-Lehmann) Inference
Description
Fits a non-parametric treatment-effect estimator based on the two-sample
Wilcoxon rank-sum test: the point estimate is the Hodges-Lehmann
location-shift estimate (the median of all pairwise treatment-minus-control
differences y_{T,i} - y_{C,j}), and both the confidence interval and
two-sided p-value are the standard rank-based Wilcoxon quantities from
stats::wilcox.test() (normal approximation with continuity correction),
not Wald intervals/tests built around the point estimate and a separately
estimated standard error. Robust to outliers and does not assume normality or
equal arm variances. Not supported for incidence (binary) responses (the
Hodges-Lehmann estimator degenerates on 0/1 data — use
InferenceAllSimpleAverageDiff or a
conditional-logistic estimator instead) or censored survival data (use
InferenceSurvivalGehanWilcox
instead). This class has no likelihood tier (likelihood_tier = "none")
and does not support the Bayesian bootstrap; its jackknife methods all report
explicit non-estimability rather than computing a (statistically unreliable)
delete-1 jackknife of the Hodges-Lehmann functional.
Initialize simple Wilcoxon inference and prepare the
rank-based treatment statistic used by
InferenceAllSimpleWilcox.
Rejects response_type = "incidence" (Hodges-Lehmann degenerates
on binary data) and rejects censored survival data at construction; see
the class-level documentation for recommended alternatives in both
cases. Legal response_type values are "continuous",
"count", "proportion", "survival" (uncensored
only), and "ordinal".
Returns the Hodges-Lehmann estimate of location
shift: the median of all pairwise treatment-minus-control differences
y_{T,i} - y_{C,j} (via wilcox_hl_point_estimate_cpp()), the
standard point estimate associated with the Wilcoxon rank-sum test.
Robust to outliers and does not assume normality or equal variances.
Wilcoxon rank-sum test two-sided p-value testing
H_0: \beta_T = \code{delta} (via stats::wilcox.test(yT, yC -
delta, exact = FALSE)$p.value, the normal approximation with
continuity correction) — a genuine rank-based test, not a
Wald test built from the Hodges-Lehmann estimate and its standard
error, despite living alongside $compute_asymp_confidence_interval()
in this class's "asymptotic" method family. For delta != 0, the
control arm's values are shifted by delta before testing, so the
test checks whether y_T and y_C + \code{delta} come from
the same distribution.
Returns the Hodges-Lehmann confidence interval directly from
stats::wilcox.test(yT, yC, conf.int = TRUE, exact = FALSE,
conf.level = 1 - alpha) — the standard nonparametric interval
associated with the Wilcoxon rank-sum test, based on inverting the
rank-sum test statistic rather than a Wald normal-approximation
interval around $compute_estimate()'s point estimate (though the
two coincide asymptotically).
Delegates to the genuine rank-based
$compute_asymp_two_sided_pval() rather than the generic
Wald-component z/t formula.
Fixed 2026-09-06: this class did not override
compute_wald_two_sided_pval, so it fell through to the
composed Wald component's generic
(estimate - delta) / se formula built from
compute_estimate() (the Hodges-Lehmann median-of-pairwise-
differences) and get_standard_error(). On heavily tied,
small-integer count/ordinal data the Hodges-Lehmann estimate lands
on exactly 0 far more often than a continuous estimator would, so
the Wald statistic came out exactly 0/se = 0 regardless of
se, forcing p = 1 deterministically (observed:
pinned at 1 in ~75-98
compute_asymp_two_sided_pval() does not have this failure
mode.
Delegates to the genuine rank-based
$compute_asymp_confidence_interval() rather than the
generic Wald normal-approximation interval, for the same reason
as compute_wald_two_sided_pval above.
Fixed 2026-09-06: the generic Wald component's normal-approximation
interval is built from get_standard_error(), which this
class derives by back-solving se = (ci[2]-ci[1]) /
(2*1.96) from stats::wilcox.test()'s own asymptotic CI
width. Under heavy ties, that root search can converge to a
numerically near-zero-width interval as a search artifact, not a
real sampling-uncertainty statement; that spurious near-zero SE
then produced a near-[0,0] Wald interval (observed in
over 1,200 rows of comprehensive-results data). The rank-based
compute_asymp_confidence_interval() inverts the rank-sum
test directly and does not go through this derived SE at all.
Reports the jackknife point-estimate as explicitly
non-estimable for this Hodges-Lehmann estimator, rather than computing
a leave-one-out jackknife: the median-of-pairwise-differences
functional is not smooth enough for the delete-1 jackknife's
linear-approximation machinery to be reliable. This method exists
purely to record that unavailability (via
private$cache_nonestimable_estimate()) rather than silently
returning a misleading number; see
InferenceJackknife for the shared
jackknife contract this method participates in.
Reports the jackknife bias-correction estimate as
non-estimable for this simple Wilcoxon estimator, for the same reason
as $compute_jackknife_estimate() (the Hodges-Lehmann functional
is not smooth enough for the delete-1 jackknife); see
InferenceJackknife for the shared
jackknife contract.
Reports the jackknife standard error as non-estimable for
this simple Wilcoxon estimator, for the same reason as
$compute_jackknife_estimate(); see
InferenceJackknife for the shared
jackknife contract.
Reports the jackknife-Wald p-value as non-estimable here,
for the same reason as $compute_jackknife_estimate(); see
InferenceJackknife.
Reports the jackknife-Wald confidence interval as
non-estimable here, for the same reason as
$compute_jackknife_estimate(); see
InferenceJackknife.
Super class
Inference -> InferenceAllSimpleWilcox
Methods
Public methods
-
InferenceAllSimpleWilcox$compute_asymp_confidence_interval() -
InferenceAllSimpleWilcox$compute_jackknife_wald_two_sided_pval() -
InferenceAllSimpleWilcox$compute_jackknife_wald_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAllSimpleWilcox$new()
Usage
InferenceAllSimpleWilcox$new( des_obj, model_formula = NULL, verbose = FALSE, max_resample_attempts = 50L, smart_cold_start_default = NULL )
Arguments
des_objA completed
DesignSeqOneByOneobject.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages. Default
FALSE.max_resample_attemptsMaximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. If all attempts fail the replicate is recorded as
NA, silently reducing the effectiveB. Must be a positive integer. Default50L.smart_cold_start_defaultFlag for consistent API.
InferenceAllSimpleWilcox$compute_estimate()
Usage
InferenceAllSimpleWilcox$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceAllSimpleWilcox$compute_asymp_two_sided_pval()
Usage
InferenceAllSimpleWilcox$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
deltaNull treatment effect. Default 0.
deltaNull treatment-effect value. Default 0.
InferenceAllSimpleWilcox$compute_asymp_confidence_interval()
Usage
InferenceAllSimpleWilcox$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
InferenceAllSimpleWilcox$compute_wald_two_sided_pval()
Usage
InferenceAllSimpleWilcox$compute_wald_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
deltaNull treatment effect. Default 0.
deltaNull treatment-effect value. Default 0.
InferenceAllSimpleWilcox$compute_wald_confidence_interval()
Usage
InferenceAllSimpleWilcox$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
InferenceAllSimpleWilcox$compute_jackknife_estimate()
Usage
InferenceAllSimpleWilcox$compute_jackknife_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllSimpleWilcox$compute_jackknife_bias_estimate()
Usage
InferenceAllSimpleWilcox$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllSimpleWilcox$compute_jackknife_std_error()
Usage
InferenceAllSimpleWilcox$compute_jackknife_std_error(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllSimpleWilcox$compute_jackknife_wald_two_sided_pval()
Usage
InferenceAllSimpleWilcox$compute_jackknife_wald_two_sided_pval( delta = 0, unit = "auto" )
Arguments
deltaNull treatment effect. Default 0.
deltaNull treatment effect. Default 0.
deltaNull treatment-effect value. Default 0.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllSimpleWilcox$compute_jackknife_wald_confidence_interval()
Usage
InferenceAllSimpleWilcox$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceAllSimpleWilcox$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAllSimpleWilcox$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Hodges, J. L., and Lehmann, E. L. (1963). "Estimates of Location Based on Rank Tests." The Annals of Mathematical Statistics, 34(2), 598-611, doi:10.1214/aoms/1177704172, for the Hodges-Lehmann estimator; Wilcoxon, F. (1945). "Individual Comparisons by Ranking Methods." Biometrics Bulletin, 1(6), 80-83, doi:10.2307/3001968, for the underlying rank-sum test.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))
seq_des_inf = InferenceAllSimpleWilcox$new(seq_des)
seq_des_inf$compute_estimate()
Asymptotic Inference
Description
Abstract class for asymptotic inference.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp
Methods
Public methods
+ inherited public methods from InferenceJackknife
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()InferenceJackknife$compute_jackknife_bias_estimate()InferenceJackknife$compute_jackknife_estimate()InferenceJackknife$compute_jackknife_std_error()InferenceJackknife$compute_jackknife_wald_confidence_interval()InferenceJackknife$compute_jackknife_wald_two_sided_pval()
+ inherited public methods from InferenceBayesianBootstrap
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAsymp$compute_asymp_confidence_interval()
Computes an asymptotic confidence interval for the treatment
effect using the configured large-sample test. For the default Wald path,
the method first calls compute_estimate(), retrieves the
class-specific standard error, and forms a normal or t interval around
the estimate. Likelihood-backed subclasses may override the dispatch; see
InferenceAsympLik.
Usage
InferenceAsymp$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level 1 -
alpha. Default 0.05.
Returns
A confidence interval.
InferenceAsymp$compute_asymp_two_sided_pval()
Computes an asymptotic two-sided p-value for the treatment
effect using the configured large-sample test. For the default Wald path,
the method compares compute_estimate() to the null value
delta using the class-specific standard error and a normal or t
reference distribution. Likelihood-backed subclasses may override the
dispatch; see InferenceAsympLik.
Usage
InferenceAsymp$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect to test against. Default 0.
Returns
The asymptotic p-value.
InferenceAsymp$get_supported_testing_types()
Gets the asymptotic testing methods supported by this inference object.
Usage
InferenceAsymp$get_supported_testing_types()
InferenceAsymp$set_testing_type()
Sets the asymptotic testing method used by p-values and CIs.
This base (Wald-only) implementation accepts only "wald" and
rejects everything else with a clear message; likelihood-tier classes
override this with a richer version supporting score/gradient/lik_ratio
testing types (see InferenceAsympLik).
Without this base method, a Wald-only class (one composing only the
Wald component, e.g. a robust-sandwich or Bai-adjusted-t
estimator) has no set_testing_type() at all, so calling it
fails with an opaque "attempt to apply non-function" instead of a
clear rejection.
Usage
InferenceAsymp$set_testing_type(testing_type = "wald")
Arguments
testing_typeOne of
"wald"for this base implementation (likelihood-tier subclasses accept more values).
Returns
The inference object, invisibly.
InferenceAsymp$compute_wald_two_sided_pval()
Computes the Wald two-sided p-value regardless of configured
testing type. This directly uses the treatment estimate, its standard
error, and the available degrees of freedom; compare with
compute_asymp_two_sided_pval() for configured-test dispatch.
Usage
InferenceAsymp$compute_wald_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect.
InferenceAsymp$compute_wald_confidence_interval()
Computes the Wald confidence interval regardless of configured
testing type. This directly uses the treatment estimate, its standard
error, and the available degrees of freedom; compare with
compute_asymp_confidence_interval() for configured-test dispatch.
Usage
InferenceAsymp$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceAsymp$compute_estimate()
Abstract method to compute the treatment-effect estimate. Concrete subclasses implement the model-specific calculation, such as an MLE coefficient, estimating-equation coefficient, standardized contrast, or rank/statistic-based treatment effect. Related p-value and interval methods call this method before using class-specific uncertainty estimates.
Usage
InferenceAsymp$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
A scalar treatment estimate.
InferenceAsymp$get_mod()
Returns the model object from the last call that produced the treatment
estimate and SE. Calls compute_estimate() first if needed.
Usage
InferenceAsymp$get_mod()
Returns
The cached model object (type depends on the concrete class).
InferenceAsymp$get_summary()
Prints a summary of the model from the last call that produced the treatment estimate and SE.
Usage
InferenceAsymp$get_summary()
InferenceAsymp$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAsymp$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Likelihood-Backed Asymptotic Inference
Description
Intermediate base class for asymptotic inference families that expose
likelihood / partial-likelihood / working-likelihood test paths in addition
to Wald inference. The term "likelihood" is used broadly: subclasses may be
backed by a true full likelihood, a partial likelihood (e.g. Cox PH), a
quasi-likelihood (e.g. GEE, quasi-Poisson), or a composite/combined
likelihood. Classes requiring a full generative likelihood — i.e. those
supporting parametric-bootstrap LR calibration — inherit instead from
InferenceParamBootstrap.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable -> InferenceAsympLik
Methods
Public methods
-
InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval() -
InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval() -
InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval() -
InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval() -
InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval() -
InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval()
+ inherited public methods from InferenceMLEorKMSummaryTable
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()InferenceJackknife$compute_jackknife_bias_estimate()InferenceJackknife$compute_jackknife_estimate()InferenceJackknife$compute_jackknife_std_error()InferenceJackknife$compute_jackknife_wald_confidence_interval()InferenceJackknife$compute_jackknife_wald_two_sided_pval()
+ inherited public methods from InferenceBayesianBootstrap
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceAsympLik$compute_asymp_confidence_interval()
Computes an asymptotic confidence interval for the treatment
effect using the configured likelihood-backed test. Wald intervals use
the fitted estimate and standard error; score, likelihood-ratio,
gradient, and Bartlett-corrected likelihood-ratio intervals are obtained
by inverting the corresponding test. For purely Wald asymptotic dispatch,
see InferenceAsymp.
Usage
InferenceAsympLik$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level 1 -
alpha. Default 0.05.
Returns
A confidence interval.
InferenceAsympLik$compute_asymp_two_sided_pval()
Computes an asymptotic two-sided p-value for the treatment
effect using the configured likelihood-backed test. Depending on
testing_type, this evaluates a Wald, score, likelihood-ratio,
gradient, or Bartlett-corrected likelihood-ratio statistic under the
null value delta. For count-specific likelihood families, see
InferenceCountLikelihood.
Usage
InferenceAsympLik$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect to test against. Default 0.
Returns
The asymptotic p-value.
InferenceAsympLik$set_testing_type()
Sets the asymptotic testing method used by p-values and CIs.
Usage
InferenceAsympLik$set_testing_type(
testing_type = c("wald", "score", "gradient", "lik_ratio", "lik_ratio_bartlett_approx",
"lik_ratio_bartlett_exact")
)
Arguments
testing_typeOne of
"wald","score","gradient","lik_ratio","lik_ratio_bartlett_approx", or"lik_ratio_bartlett_exact".
Returns
The inference object, invisibly.
InferenceAsympLik$set_information_preference()
Sets the information matrix preference used by score-test dispatch.
Usage
InferenceAsympLik$set_information_preference(
information_preference = c("auto", "fisher", "observed")
)
Arguments
information_preferenceOne of
"auto","fisher", or"observed".
Returns
The inference object, invisibly.
InferenceAsympLik$get_testing_type()
Gets the asymptotic testing method used by p-values and CIs.
Usage
InferenceAsympLik$get_testing_type()
InferenceAsympLik$get_information_preference()
Gets the score-test information matrix preference.
Usage
InferenceAsympLik$get_information_preference()
InferenceAsympLik$get_information_source_used()
Gets the actual information source used by the most recent information-backed computation.
Usage
InferenceAsympLik$get_information_source_used()
InferenceAsympLik$get_supported_testing_types()
Gets the asymptotic testing methods supported by this inference object.
Usage
InferenceAsympLik$get_supported_testing_types()
InferenceAsympLik$get_supported_information_preferences()
Gets the score-test information matrix preferences supported by this inference object.
Usage
InferenceAsympLik$get_supported_information_preferences()
InferenceAsympLik$compute_score_two_sided_pval()
Computes the score two-sided p-value regardless of configured
testing type. The score test evaluates the null-restricted fit and uses
the configured information matrix preference; compare with
compute_asymp_two_sided_pval() for configured-test dispatch.
Usage
InferenceAsympLik$compute_score_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect.
InferenceAsympLik$compute_score_confidence_interval()
Computes the score confidence interval regardless of configured
testing type by inverting score-test p-values over candidate treatment
effects. Compare with compute_asymp_confidence_interval() for
configured-test dispatch.
Usage
InferenceAsympLik$compute_score_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceAsympLik$compute_lik_ratio_two_sided_pval()
Computes the likelihood-ratio two-sided p-value regardless of
configured testing type. This compares unrestricted and null-restricted
fits at delta; subclasses may use a full, partial, quasi-, or
composite likelihood as described in
InferenceAsympLik.
Usage
InferenceAsympLik$compute_lik_ratio_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect.
InferenceAsympLik$compute_lik_ratio_confidence_interval()
Computes the likelihood-ratio confidence interval regardless of configured testing type by inverting likelihood-ratio p-values over candidate treatment effects. Compare with score and gradient intervals in this class.
Usage
InferenceAsympLik$compute_lik_ratio_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval()
Computes the approximate (Monte-Carlo) Bartlett-corrected likelihood-ratio
two-sided p-value regardless of configured testing type. Returns NA_real_
for subclasses that do not implement an approximate Bartlett correction factor.
The approximate Bartlett factor is estimated by Monte Carlo (e.g. the generic
InferenceParamBootstrap factor): B datasets are simulated under
the null-restricted fit at delta and refit to approximate
E[LR | H0], the quantity a classical analytic Bartlett correction
targets exactly. The Monte-Carlo draws are seeded from this object's own
seed (see set_seed()), so repeated calls at the same
delta with the same B are reproducible; there is no separate
seed argument here.
See compute_lik_ratio_bartlett_exact_two_sided_pval() for the
closed-form analytic counterpart (no simulation, no B).
Usage
InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval( delta = 0, B = 99 )
Arguments
deltaNull treatment effect. Default 0.
BNumber of Monte-Carlo replicates used to estimate the Bartlett factor. Default 99.
InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval()
Computes the approximate (Monte-Carlo) Bartlett-corrected likelihood-ratio
confidence interval regardless of configured testing type. Returns
c(NA_real_, NA_real_) for subclasses that do not implement an
approximate Bartlett correction factor.
See compute_lik_ratio_bartlett_approx_two_sided_pval() for what
B controls and how the Monte-Carlo seed is inherited from this
object's own seed. Each p-value evaluation during the
confidence-interval search re-simulates B replicates, so this can
be substantially more expensive than the p-value alone.
Usage
InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval( alpha = 0.05, B = 99 )
Arguments
alphaSignificance level. Default 0.05.
BNumber of Monte-Carlo replicates used to estimate the Bartlett factor. Default 99.
InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval()
Computes the exact (closed-form analytic) Bartlett-corrected
likelihood-ratio two-sided p-value regardless of configured testing type.
Returns NA_real_ for subclasses that do not implement an exact,
bespoke analytic Bartlett correction factor. Concrete families opt in
through get_bartlett_factor_exact().
Unlike compute_lik_ratio_bartlett_approx_two_sided_pval(), this path
involves no simulation and no Monte-Carlo replicate count.
Usage
InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval()
Computes the exact (closed-form analytic) Bartlett-corrected
likelihood-ratio confidence interval regardless of configured testing type.
Returns c(NA_real_, NA_real_) for subclasses that do not implement an
exact, bespoke analytic Bartlett correction factor.
Usage
InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval( alpha = 0.05 )
Arguments
alphaSignificance level. Default 0.05.
InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval()
Computes "the best available" Bartlett-corrected likelihood-ratio
two-sided p-value regardless of configured testing type: uses the exact
(closed-form analytic) factor if this class implements one, otherwise
falls back to the approximate (Monte-Carlo) factor. Errors if the class
supports neither (see supports_bartlett_likelihood_ratio_exact()/
supports_bartlett_likelihood_ratio_approx()).
This is a convenience entry point for callers who want a Bartlett-corrected
p-value without caring which mechanism produced it. Because exact and
approximate factors are computed differently (deterministic closed form vs.
seeded Monte-Carlo simulation), the same call can silently start
returning different numeric results on a future package version once a
family gains an exact implementation where previously only the
approximate path existed. Callers who need results stable across package
versions (e.g. for reproducibility or regression tests) should call
compute_lik_ratio_bartlett_approx_two_sided_pval() or
compute_lik_ratio_bartlett_exact_two_sided_pval() directly instead.
Usage
InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval(delta = 0, B = 99)
Arguments
deltaNull treatment effect. Default 0.
BNumber of Monte-Carlo replicates, used only when the exact factor is unavailable and the approximate factor is used instead. If explicitly supplied but the exact factor is used (so
Bhas no effect), a warning is issued;Bleft at its default is silently ignored in that case. Default 99.
InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval()
Computes "the best available" Bartlett-corrected likelihood-ratio confidence interval regardless of configured testing type: uses the exact (closed-form analytic) factor if this class implements one, otherwise falls back to the approximate (Monte-Carlo) factor. Errors if the class supports neither.
See compute_lik_ratio_bartlett_two_sided_pval() for the exact-over-approx
selection rule, the B-ignored warning behavior, and why callers who
need version-to-version reproducibility should prefer the explicit
_approx/_exact methods instead.
Usage
InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval( alpha = 0.05, B = 99 )
Arguments
alphaSignificance level. Default 0.05.
BNumber of Monte-Carlo replicates, used only when the exact factor is unavailable and the approximate factor is used instead. Default 99.
InferenceAsympLik$compute_gradient_two_sided_pval()
Computes the gradient two-sided p-value regardless of configured testing type.
Usage
InferenceAsympLik$compute_gradient_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect.
InferenceAsympLik$compute_gradient_confidence_interval()
Computes the gradient confidence interval regardless of configured testing type.
Usage
InferenceAsympLik$compute_gradient_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceAsympLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceAsympLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Bai Adjusted-t Mean-Difference Inference for KK14 Designs
Description
Continuous-response mean-difference inference for designs assigned by
DesignSeqOneByOneKK14 (the
Kapelner-Krieger 2014 sequential matching-on-the-fly design). The point
estimate and its variance are the closed-form Bai-adjusted-t combination of
the matched-pairs mean difference and the unmatched-reservoir mean
difference, inverse-variance-weighted when both are usable; the full
formula, pair-distance definition, and confidence-interval/p-value
construction are shared with
InferenceBaiAdjustedTKK21.
The two leaves differ only in how pair distance is defined during matching:
this class (KK14) uses the plain squared Euclidean distance
\sum_j (x_{1j} - x_{2j})^2 between candidate subjects' covariate
vectors, unlike KK21's covariate-weighted distance. Because the estimator
is closed-form, initialization does not use warm starts (there is no
iterative fit to warm-start).
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceBaiAdjustedTKK14
Methods
Public methods
-
InferenceBaiAdjustedTKK14$approximate_randomization_distribution_beta_hat_T() -
InferenceBaiAdjustedTKK14$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceBaiAdjustedTKK14$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceBaiAdjustedTKK14$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceBaiAdjustedTKK14$supports_rand_pval_for_incidence()
Usage
InferenceBaiAdjustedTKK14$supports_rand_pval_for_incidence()
InferenceBaiAdjustedTKK14$compute_rand_two_sided_pval()
Usage
InferenceBaiAdjustedTKK14$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceBaiAdjustedTKK14$clone()
The objects of this class are cloneable with this method.
Usage
InferenceBaiAdjustedTKK14$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceBaiAdjustedTKK14$new(seq_des)
inf$compute_estimate()
Bai Adjusted-t Mean-Difference Inference for KK21 Designs
Description
Continuous-response mean-difference inference for designs assigned by
DesignSeqOneByOneKK21 (the
Kapelner-Krieger 2021 sequential matching-on-the-fly design with
covariate-weighted matching). The point estimate and its variance are the
closed-form Bai-adjusted-t combination of the matched-pairs mean
difference and the unmatched-reservoir mean difference,
inverse-variance-weighted when both are usable; see
InferenceBaiAdjustedTKK14 for
the full formula, the pair-distance definition, and the
confidence-interval/p-value construction shared with this class.
The two leaves differ only in how pair distance is defined during
matching: this class (KK21) uses the design's covariate-weighted squared
distance \sum_j w_j (x_{1j} - x_{2j})^2, where w_j are the
design's covariate_weights (see
DesignSeqOneByOneKK21), unlike
KK14's unweighted distance. Because the estimator is closed-form,
initialization does not use warm starts (there is no iterative fit to
warm-start).
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceBaiAdjustedTKK21
Methods
Public methods
-
InferenceBaiAdjustedTKK21$approximate_randomization_distribution_beta_hat_T() -
InferenceBaiAdjustedTKK21$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceBaiAdjustedTKK21$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceBaiAdjustedTKK21$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceBaiAdjustedTKK21$supports_rand_pval_for_incidence()
Usage
InferenceBaiAdjustedTKK21$supports_rand_pval_for_incidence()
InferenceBaiAdjustedTKK21$compute_rand_two_sided_pval()
Usage
InferenceBaiAdjustedTKK21$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceBaiAdjustedTKK21$clone()
The objects of this class are cloneable with this method.
Usage
InferenceBaiAdjustedTKK21$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceBaiAdjustedTKK21$new(seq_des)
inf$compute_estimate()
Bayesian Bootstrap-capable Inference
Description
Abstract class for Dirichlet-weight Bayesian bootstrap inference layered on top of the existing nonparametric bootstrap infrastructure.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap
Methods
Public methods
-
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types() -
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types() -
InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights() -
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T() -
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval() -
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
Returns the type values
compute_bayesian_bootstrap_two_sided_pval() accepts.
Usage
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()
Returns the type values
compute_bayesian_bootstrap_confidence_interval() accepts.
Usage
InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()
InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()
Recomputes the treatment estimate under Bayesian-bootstrap subject-, block-, cluster-, or matched-set weights.
This is an abstract hook implemented by concrete inference families that support weighted re-estimation.
Usage
InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric Bayesian-bootstrap weights at the design's exchangeable resampling unit. For ordinary designs these are subject-level weights. For blocking, clustering, or matching designs these may instead be block-, cluster-, pair-, or matched-set-level weights, depending on
weighting_unit_type.estimate_onlyIf
TRUE, compute only the point estimate for the weighted replicate.
Returns
A numeric treatment-effect estimate for the weighted replicate.
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()
Creates the Bayesian-bootstrap distribution of the treatment estimate using Dirichlet weights.
Usage
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T( B = 501, show_progress = TRUE, debug = FALSE, weighting_unit_type = NULL )
Arguments
BNumber of Bayesian-bootstrap replicates. The default is 501.
show_progressA flag indicating whether a progress bar should be displayed.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.weighting_unit_typeOptional Bayesian-bootstrap weighting-unit scheme. Legal public values are:
NULLUse the design's default weighting-unit logic. For ordinary non-blocking designs this is the usual subject-level Bayesian bootstrap. For certain blocking designs,
NULLmaps to the same behavior as"within_blocks"."within_blocks"Only legal for blocking-style designs that support block-aware weighting:
DesignFixedBlocking,DesignFixedOptimalBlocks,DesignSeqOneByOneSPBR, andDesignFixedBlockedCluster. Draws Dirichlet weights on observational units within each observed block/stratum. For blocked cluster designs this means cluster-within-stratum weights."resample_blocks"Only legal for the same blocking-style designs as
"within_blocks". Draws Dirichlet weights on whole observed blocks/strata rather than on units within each block.
Any non-
NULLvalue is rejected for designs outside that blocking family.
Returns
When debug = FALSE (default), a numeric vector of length
B containing the Bayesian-bootstrap estimates. When
debug = TRUE, a list with: values, errors (list of
character vectors, one per iteration), warnings (list of
character vectors, one per iteration), num_errors,
num_warnings, prop_iterations_with_errors,
prop_iterations_with_warnings, and
prop_illegal_values.
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()
Computes a Bayesian-bootstrap-based two-sided p-value for the treatment effect.
Usage
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval( delta = 0, B = 501, type = NULL, na.rm = FALSE, show_progress = TRUE, min_number_usable_samples = 5L, weighting_unit_type = NULL )
Arguments
deltaNull hypothesis value. Default 0.
BNumber of Bayesian-bootstrap replicates. Default 501.
typeType of Bayesian-bootstrap p-value. Supported values are
"percentile"(default),"symmetric","wald","studentized"/"bootstrap-t"(pivots by replicate SE fromcompute_estimate_with_bootstrap_weights(..., estimate_only = FALSE)), and"bca"(bias-corrected and accelerated via leave-one-unit-out Bayesian jackknife).na.rmIf
TRUE, discard non-finite bootstrap replicates before computing the p-value. Otherwise, any non-finite replicate returnsNA.show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite Bayesian-bootstrap replicates required after filtering. Default 5.
weighting_unit_typeOptional Bayesian-bootstrap weighting-unit scheme. See
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T().
Returns
A numeric two-sided p-value, or NA_real_ if too few usable
replicates remain or the estimate is non-finite.
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()
Computes a Bayesian-bootstrap confidence interval for the treatment effect.
Usage
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, min_number_usable_samples = 5L, weighting_unit_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
BNumber of Bayesian-bootstrap replicates. Default 501.
typeType of Bayesian-bootstrap interval. Supported values are
"percentile"(default),"basic","wald","studentized"/"bootstrap-t"(pivots by replicate SE fromcompute_estimate_with_bootstrap_weights(..., estimate_only = FALSE)), and"bca"(bias-corrected and accelerated via leave-one-unit-out Bayesian jackknife).na.rmIf
TRUE, discard non-finite bootstrap replicates before constructing the interval.show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite Bayesian-bootstrap replicates required after filtering. Default 5.
weighting_unit_typeOptional Bayesian-bootstrap weighting-unit scheme. See
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T().
Returns
A length-2 numeric confidence interval. Returns
c(NA_real_, NA_real_) when the estimate is non-finite or too few
usable replicates remain.
InferenceBayesianBootstrap$clone()
The objects of this class are cloneable with this method.
Usage
InferenceBayesianBootstrap$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Linear Mixed Model Inference for KK Designs with Continuous Response
Description
Fits a linear mixed model for continuous responses under a KK
matching-on-the-fly design. The matched-pair strata enter as a subject-level
random intercept (1 | group_id), accounting for within-pair correlation.
When use_rcpp = TRUE (default) the likelihood is maximised by an
internal Rcpp/L-BFGS routine that requires no external packages. Set
use_rcpp = FALSE to fall back to glmmTMB.
The treatment coefficient \beta_T is on the response's natural
(untransformed) scale — a mean difference, not a ratio or log-scale
effect. likelihood_tier = "full": likelihood-ratio, score, and
Wald tests are all available when the model converges (see
$get_likelihood_test_spec() inherited from the shared count/GLMM
likelihood plumbing). Validity requires the random-intercept-per-pair
structure to correctly capture the design's matching dependence and the
usual linear mixed model assumptions (conditional normality of responses
and pair effects, correctly specified fixed-effects formula).
Super class
Inference -> InferenceContinKKGLMM
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKGLMM$new()
Initialize inference for the linear mixed model
Y_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + b_{g(i)} +
\epsilon_i, b_g \sim N(0, \sigma_b^2), \epsilon_i \sim
N(0, \sigma_e^2), where g(i) is subject i's matched-pair
group id, W_i is the treatment indicator, X_i are
covariates, and \beta_T is the treatment effect (mean
difference on the response's natural scale). The random intercept
b_g absorbs the within-pair correlation induced by matching, so
\beta_T's standard error correctly reflects the design.
Usage
InferenceContinKKGLMM$new( des_obj, model_formula = NULL, use_rcpp = TRUE, use_gls_fast_path = TRUE, use_gls_fast_path_bootstrap = FALSE, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a continuous response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.use_rcppLogical. If
TRUE(default), use the optimised Rcpp Gaussian LMM implementation (no external package required). IfFALSE, use glmmTMB.use_gls_fast_pathLogical. If
TRUE(default), use a fast GLS estimator (no optimisation) forestimate_onlycalls during randomisation inference once variance components are cached from a prior full fit. Statistically exact: fixing VC at the null-fit MLE and permuting only the treatment assignment gives a valid permutation test by exchangeability. SetFALSEto always run full L-BFGS.use_gls_fast_path_bootstrapLogical. If
TRUE, also use the fast GLS estimator for non-studentised bootstrap draws (estimate_only = TRUEweighted calls). Asymptotically valid by the plug-in principle (VC orthogonal to beta_T in the Fisher information), but not exact in finite samples. DefaultFALSE.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
optimization_algThe optimization algorithm to use. Default is dispatched via policy.
InferenceContinKKGLMM$compute_estimate()
Fits the linear mixed model by maximum likelihood and
returns \hat\beta_T. When use_rcpp = TRUE (the
default), the log-likelihood
\ell(\beta, \sigma_b^2, \sigma_e^2) is maximized directly by an
internal Rcpp/L-BFGS routine over \beta and the log-variance
components; otherwise glmmTMB performs the fit. If
use_gls_fast_path = TRUE and variance components are already
cached from a prior full fit (used during randomization inference,
where only the treatment column changes across permutations),
estimate_only = TRUE calls instead solve the generalized
least-squares problem at the cached variance components — exact
under exchangeability of the permuted treatment assignment, since
fixing the variance components at their null-fit MLE and permuting
only W preserves the permutation test's validity.
Usage
InferenceContinKKGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations (standard error, degrees of freedom) needed for confidence intervals or p-values; only
\hat\beta_Tis returned.
InferenceContinKKGLMM$compute_estimate_with_bootstrap_weights()
Refits the linear mixed model with subject/block-level
weights applied to each row's contribution to the likelihood
(Bayesian-bootstrap or nonparametric-bootstrap draw weights,
expanded from subject/block level to individual rows via
private$expand_subject_or_block_weights_to_row_weights()),
and returns the reweighted estimate \hat\beta_T^{(w)}. When
weights are effectively constant, this collapses to the unweighted
compute_estimate() call (returns df = Inf to signal a
degenerate/skipped bootstrap replicate rather than refitting). When
use_rcpp = TRUE, a weighted Rcpp fast path is tried first via
private$weighted_rcpp_estimate(); otherwise
private$compute_weighted_glmm_bootstrap_estimate() refits via
glmmTMB-based machinery.
Usage
InferenceContinKKGLMM$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric vector of nonnegative weights, one per matched-pair group or reservoir subject (bootstrap draw weights), expanded to per-row weights before fitting.
estimate_onlyLogical. If TRUE, skip standard-error computation.
Returns
The reweighted treatment estimate \hat\beta_T^{(w)}.
InferenceContinKKGLMM$compute_asymp_confidence_interval()
Computes a 1-\alpha Wald confidence interval for
\beta_T from the fitted-model standard error, using a normal
or t critical value depending on the resolved degrees of
freedom (see private$compute_z_or_t_ci_from_s_and_df).
Usage
InferenceContinKKGLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level of the interval is
1 - \code{alpha}. Default0.05.
InferenceContinKKGLMM$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value for
H_0: \beta_T = \code{delta} using the fitted-model estimate
and standard error (same statistic that
$compute_asymp_confidence_interval() inverts).
Usage
InferenceContinKKGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment-effect value. Default
0.
InferenceContinKKGLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKGLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
See Also
Comparable Python API: statsmodels MixedLM. See also: Mixed model (Wikipedia).
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKGLMM$new(seq_des)
inf$compute_estimate()
OLS IVWC Compound Inference for KK Designs
Description
Fits a variance-weighted compound estimator for KK matching-on-the-fly designs with continuous responses using OLS regression for matched-pair differences and reservoir outcomes, with the treatment indicator and, optionally, all recorded covariates as predictors. Note that warm starts are disabled for this class as OLS is a closed-form estimator and does not benefit from initialization.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
The point estimate \hat\beta_T is the inverse-variance-weighted
combination of an OLS fit on matched-pair within-pair differences and an
OLS fit on reservoir (unmatched) subjects' outcomes, falling back to
whichever sub-fit is usable if the other is not — the same compound
combination rule used by
InferenceAllKKMeanDiffIVWC,
generalized here to allow covariate adjustment via model_formula.
likelihood_tier = "none": this is an estimating-equation (least
squares) estimator, not a fitted likelihood, so only Wald-type asymptotic
inference is available (no likelihood-ratio or score test).
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceContinKKOLSIVWC
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKOLSIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceContinKKOLSIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKOLSIVWC$supports_rand_pval_for_incidence()
Usage
InferenceContinKKOLSIVWC$supports_rand_pval_for_incidence()
InferenceContinKKOLSIVWC$compute_rand_two_sided_pval()
Usage
InferenceContinKKOLSIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKOLSIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKOLSIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKOLSIVWC$new(seq_des)
inf$compute_estimate()
OLS Combined-Likelihood Inference for KK Designs
Description
Fits a single stacked OLS regression over matched-pair differences and reservoir observations for KK matching-on-the-fly designs with continuous responses, using the treatment indicator and, optionally, all recorded covariates as predictors. Note that warm starts are disabled for this class as OLS is a closed-form estimator and does not benefit from initialization.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Model. Let m be the number of matched pairs and
n_R = n_{RT} + n_{RC} the number of unmatched reservoir subjects. The
design matrix stacks two blocks: m matched-pair difference rows
(each row's response is the within-pair outcome difference, coded with an
implicit unit treatment column and covariate differences X_{d}), and
n_R reservoir rows (raw covariates plus a treatment/matching-status
indicator column). The stacked regression is fit by ordinary least squares
(lm.fit), and \hat\beta_T is the coefficient on
the treatment/matched-difference column, i.e. an additive mean-difference
estimand on the outcome's natural scale. If only matched pairs exist,
j_treat = 1; if only reservoir data exist, j_treat = 2; if
both exist, the combined design uses j_treat = 2. If neither
matched pairs nor a treatment-and-control-populated reservoir exist, the
estimate is marked nonestimable ("no_usable_matched_or_reservoir_data").
Variance. Standard errors use the HC2 heteroskedasticity-consistent
sandwich estimator (ols_hc2_post_fit_cpp), not the classical OLS
variance, so the Wald confidence interval/p-value are robust to
heteroskedasticity across the matched/reservoir blocks.
Likelihood tier. likelihood_tier = "full": this is a genuine
Gaussian likelihood (not a quasi-likelihood or partial likelihood), so
score, gradient, and likelihood-ratio testing types are available in
addition to Wald, and an exact (not higher-order-accurate) Bartlett
correction reproduces base R's lm() classical partial F-test exactly
under the classical homoskedastic-Gaussian-errors assumption (a stronger
assumption than the HC2-robust Wald path uses, so the two paths need not
agree numerically).
Assumptions. Continuous response; independent matched pairs and/or
independent reservoir subjects; no censoring
(assertNoCensoring() is enforced); a
KK matching-on-the-fly design
(DesignSeqOneByOneKK14 or subclass).
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceContinKKOLSOneLik
Methods
Public methods
-
InferenceContinKKOLSOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceContinKKOLSOneLik$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKOLSOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceContinKKOLSOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKOLSOneLik$supports_rand_pval_for_incidence()
Usage
InferenceContinKKOLSOneLik$supports_rand_pval_for_incidence()
InferenceContinKKOLSOneLik$compute_rand_two_sided_pval()
Usage
InferenceContinKKOLSOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKOLSOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKOLSOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential
allocation with higher power and efficiency. Biometrics, 70(2),
378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)
See Also
InferenceContinKKOLSIVWC
for the inverse-variance-weighted-combination alternative to this
one-likelihood combined-fit approach; analogous Python API:
statsmodels GLM/OLS.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKOLSOneLik$new(seq_des)
inf$compute_estimate()
Quantile Regression Compound Estimator for KK Matching-on-the-Fly Designs
Description
A variance-weighted compound quantile regression estimator for KK matching-on-the-fly designs with continuous responses. The estimator combines:
Quantile regression on within-pair differences (matched pairs)
Quantile regression on reservoir subjects (treatment vs control)
using the same variance-weighted combination logic as the OLS compound estimator.
Default quantile: tau = 0.5 (median regression).
At tau = 0.5 this estimates the median treatment effect, which is the canonical
nonparametric location estimator and is more robust to outliers and heavy-tailed
response distributions than the OLS mean-based estimator. To target a different
quantile of the treatment effect distribution — for example the 25th or 75th
percentile — pass tau = 0.25 or tau = 0.75 to the constructor:
inf = InferenceContinKKQuantileRegrIVWC$ new(seq_des, tau = 0.75)
Any value strictly between 0 and 1 is accepted.
Standard errors use Powell's "nid" sandwich estimator (non-iid), which is more robust than the "iid" (constant-density) assumption; the implementation falls back to "iid" on failure. Asymptotic z-based inference is used throughout.
The randomization-based confidence interval is inherited from the base class and is valid for location-shift models at all quantiles: shifting y by delta maps the tau-th quantile treatment effect to delta under the null.
This class requires the quantreg package, which is listed in Suggests and is not installed automatically with EDI. Install quantreg before using this class.
Legacy class. Not fully tested in comprehensive_tests.R.
Super class
Inference -> InferenceContinKKQuantileRegrIVWC
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKQuantileRegrIVWC$new()
Initialize continuous-response KK IVWC quantile-regression
inference; see
InferenceContinKKQuantileRegrIVWC.
Usage
InferenceContinKKQuantileRegrIVWC$new( des_obj, model_formula = NULL, tau = 0.5, verbose = FALSE )
Arguments
des_objA DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.tauThe quantile level for regression, strictly between 0 and 1. The default
tau = 0.5estimates the median treatment effect. Pass a different value (e.g.tau = 0.25ortau = 0.75) to target the corresponding percentile of the treatment effect distribution.verboseA flag indicating whether messages should be displayed to the user. Default is
FALSE.
Examples
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrIVWC$new(seq_des, verbose = FALSE)
infer
InferenceContinKKQuantileRegrIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKQuantileRegrIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKQuantileRegrIVWC$new(seq_des)
inf$compute_estimate()
## ------------------------------------------------
## Method `InferenceContinKKQuantileRegrIVWC$new()`
## ------------------------------------------------
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrIVWC$new(seq_des, verbose = FALSE)
infer
Quantile Regression Combined-Likelihood Compound Estimator for KK Designs (Continuous)
Description
Fits the combined stacked quantile regression (matched-pair differences + reservoir) using the treatment indicator and all recorded covariates for continuous responses. Minimises the joint check-function loss over both data sources simultaneously. Inference is based on the stacked combined-likelihood quantile-regression fit.
Model. Analogous to
InferenceContinKKOLSOneLik's single
stacked design (matched-pair difference rows plus reservoir rows fit
jointly), but the objective is the quantreg check-function loss
\rho_\tau(u) = u(\tau - \mathbf{1}_{u<0}) rather than squared error,
so the estimand \beta_T is the treatment effect on the \tau-th
quantile of the response, not the mean. tau = 0.5 (default) targets
the median treatment effect, which is more robust to outliers and
heavy-tailed responses than the OLS mean-based estimator; any value
strictly between 0 and 1 is accepted. likelihood_tier = "none": no
likelihood-based (score/gradient/lik_ratio) testing types, only Wald.
Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring; a KK matching-on-the-fly design. Requires the quantreg package (listed in Suggests, not installed automatically).
Super class
Inference -> InferenceContinKKQuantileRegrOneLik
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKQuantileRegrOneLik$new()
Initialize continuous-response KK combined-likelihood quantile-regression
inference. The shared stacked matched-pair/reservoir quantile-regression fit and its
tau semantics are documented on
InferenceContinKKQuantileRegrOneLik
and on the shared compute_estimate() method it inherits from the
KKQuantileRegrOneLik component.
Usage
InferenceContinKKQuantileRegrOneLik$new( des_obj, model_formula = NULL, tau = 0.5, verbose = FALSE )
Arguments
des_objA DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.tauThe quantile level for regression, strictly between 0 and 1. Default is 0.5.
verboseWhether to print progress messages.
Examples
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrOneLik$new(seq_des, verbose
= FALSE)
infer
InferenceContinKKQuantileRegrOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKQuantileRegrOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential
allocation with higher power and efficiency. Biometrics, 70(2),
378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)
Koenker, R. (2005). Quantile Regression. Cambridge University Press.
See Also
InferenceContinKKQuantileRegrIVWC
for the inverse-variance-weighted-combination alternative to this
one-likelihood combined-fit approach.
Quantile
regression (orientation).
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKQuantileRegrOneLik$new(seq_des)
inf$compute_estimate()
## ------------------------------------------------
## Method `InferenceContinKKQuantileRegrOneLik$new()`
## ------------------------------------------------
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK14$new(n = nrow(x_dat), response_type = "continuous", verbose =
FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(c(1.2, 0.9, 1.5, 1.8, 2.1, 1.7, 2.6, 2.2))
infer <- InferenceContinKKQuantileRegrOneLik$new(seq_des, verbose
= FALSE)
infer
Robust-Regression IVWC Compound Inference for KK Designs
Description
Fits a variance-weighted compound estimator for KK matching-on-the-fly designs with continuous responses using robust regression for matched-pair differences and reservoir outcomes, with treatment and, optionally, all recorded covariates as predictors.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceContinKKRobustRegrIVWC
Methods
Public methods
-
InferenceContinKKRobustRegrIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceContinKKRobustRegrIVWC$supports_rand_pval_for_incidence() -
InferenceContinKKRobustRegrIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKRobustRegrIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceContinKKRobustRegrIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKRobustRegrIVWC$supports_rand_pval_for_incidence()
Usage
InferenceContinKKRobustRegrIVWC$supports_rand_pval_for_incidence()
InferenceContinKKRobustRegrIVWC$compute_rand_two_sided_pval()
Usage
InferenceContinKKRobustRegrIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKRobustRegrIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKRobustRegrIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinKKRobustRegrIVWC$new(seq_des)
inf$compute_estimate()
Robust-Regression Combined-Likelihood Inference for KK Designs
Description
Fits a single stacked robust regression over matched-pair differences and reservoir observations for KK matching-on-the-fly designs with continuous responses.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Model. Analogous to
InferenceContinKKOLSOneLik's single
stacked design (matched-pair difference rows plus reservoir rows fit in one
regression, treatment coefficient \beta_T), but fit with a robust
M/MM-estimator (MASS::rlm, or an internal Rcpp IRLS kernel when
use_rcpp = TRUE) instead of ordinary least squares. "MM" (the
default method) starts from an LQS-based high-breakdown fit; "M"
can warm-start from OLS (start_with_ols = TRUE).
Likelihood tier. likelihood_tier = "quasi": the robust
objective is not a normalized likelihood, so only Wald-type asymptotic
inference is available (compute_wald_confidence_interval()/
compute_wald_two_sided_pval(), aliased by the standard
compute_asymp_* names) — no score/gradient/likelihood-ratio testing
types, unlike the OLS one-likelihood sibling.
Assumptions. Continuous response; independent matched pairs and/or independent reservoir subjects; no censoring; a KK matching-on-the-fly design. Robust regression trades some efficiency under exactly-Gaussian errors for resistance to outliers and heavy tails.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceContinKKRobustRegrOneLik
Methods
Public methods
-
InferenceContinKKRobustRegrOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceContinKKRobustRegrOneLik$supports_rand_pval_for_incidence() -
InferenceContinKKRobustRegrOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinKKRobustRegrOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceContinKKRobustRegrOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKRobustRegrOneLik$supports_rand_pval_for_incidence()
Usage
InferenceContinKKRobustRegrOneLik$supports_rand_pval_for_incidence()
InferenceContinKKRobustRegrOneLik$compute_rand_two_sided_pval()
Usage
InferenceContinKKRobustRegrOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceContinKKRobustRegrOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinKKRobustRegrOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential
allocation with higher power and efficiency. Biometrics, 70(2),
378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)
See Also
InferenceContinKKRobustRegrIVWC
for the inverse-variance-weighted-combination alternative to this
one-likelihood combined-fit approach. Analogous Python API:
statsmodels RLM.
Lin (2013) Covariate-Adjusted OLS Inference for Continuous Responses
Description
Fits Lin's (2013) covariate-adjusted linear estimator for continuous
responses: OLS of Y_i on [1, W_i, X_i^c, W_i X_i^c], where
X_i^c = X_i - \bar X are covariates centered at their sample means
and W_i X_i^c are treatment-by-centered-covariate interactions
(omitted when there are no covariates, reducing to plain OLS with
\hat\beta_T the simple mean difference). Centering makes
\hat\beta_T interpretable as the average treatment effect regardless
of whether interactions are included, following the design-based
reinterpretation of Freedman's critique of ANCOVA in randomized
experiments. Standard errors use the HC2 heteroskedasticity-consistent
(Huber-White-type) covariance estimator (ols_hc2_post_fit_cpp),
not classical OLS SEs assuming homoskedasticity — this is the estimator
Lin (2013) recommends since it remains conservative under treatment-effect
heterogeneity, unlike the classical or HC0 sandwich variants.
likelihood_tier = "full": Wald, score, gradient, and
likelihood-ratio tests are all available, with the parametric-likelihood
bootstrap using the OLS Gaussian-errors model
(Y_i \mid x_i \sim N(x_i^\top \beta, \sigma^2)) as the generative
null even though the design-based HC2 standard error does not itself
assume homoskedastic Gaussian errors — the likelihood-ratio/score/gradient
machinery is a secondary, model-based inference path alongside the primary
HC2-Wald and randomization paths. Validity of \hat\beta_T as an
average-treatment-effect estimator relies on randomization (of W),
not on any particular outcome model; the working linear model need not be
correctly specified.
Super class
Inference -> InferenceContinLin
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinLin$new()
Initialize inference for Lin's (2013) covariate-adjusted OLS
estimator (intercept, treatment, centered covariates, and
treatment-by-centered-covariate interactions); see
InferenceContinLin for the model
form. Does not fit the model; the fit is deferred to the first call to
compute_estimate() or a method that requires it.
Usage
InferenceContinLin$new( des_obj, model_formula = NULL, verbose = FALSE, harden = TRUE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a continuous response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
hardenFlag for consistent API.
smart_cold_start_defaultFlag for consistent API.
InferenceContinLin$compute_estimate()
Fits Lin's covariate-adjusted OLS model
(stats's lm.fit on the centered-covariate design
matrix) and returns \hat\beta_T, the estimated average treatment
effect. A design matrix that is rank-deficient or has fewer usable rows
than columns, or a fit with non-finite coefficients, is cached as
nonestimable rather than returned.
Usage
InferenceContinLin$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf
TRUE, skip HC2 variance computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceContinLin$compute_estimate_with_bootstrap_weights()
Refits Lin's model with subject/block-level weights applied
to a weighted least-squares fit (stats::lm.wfit) — Bayesian-bootstrap
or nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights() — and
returns the reweighted estimate \hat\beta_T^{(w)}. When
estimate_only = FALSE, also computes a weighted-residual variance
estimate (not the HC2 estimator used by
compute_asymp_confidence_interval()) for internal bootstrap
diagnostics. Rows with non-finite or non-positive weight, or non-finite
response, are dropped from the weighted fit; if no rows remain, or the
fit fails, the estimate is NA.
Usage
InferenceContinLin$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceContinLin$compute_asymp_confidence_interval()
Wald confidence interval for \beta_T using the HC2
heteroskedasticity-robust standard error
(ols_hc2_post_fit_cpp); see
InferenceAsymp for the shared
t/z interval contract. Fits the model first if not already
cached.
Usage
InferenceContinLin$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level. The default is 0.05.
InferenceContinLin$compute_asymp_two_sided_pval()
Two-sided Wald test of H_0: \beta_T = \code{delta}
using the HC2 heteroskedasticity-robust standard error; see
InferenceAsymp for the shared
t/z test contract. Fits the model first if not already
cached.
Usage
InferenceContinLin$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect. Defaults to 0.
InferenceContinLin$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinLin$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Lin, W. (2013). "Agnostic notes on regression adjustments to experimental data: Reexamining Freedman's critique." The Annals of Applied Statistics, 7(1), 295-318, doi:10.1214/12-AOAS583.
See Also
InferenceContinOLS for the
uncentered, non-interacted OLS estimator this class generalizes.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinLin$new(seq_des)
inf$compute_estimate()
OLS Inference for Continuous Responses
Description
Fits an ordinary least squares regression for continuous responses:
Y_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \epsilon_i, using the
treatment indicator and, optionally, all recorded covariates as predictors
(uncentered, no treatment-covariate interactions — see
InferenceContinLin for the
centered-covariate, interacted variant). \hat\beta_T is a mean
difference on the response's natural scale. likelihood_tier =
"full": Wald, score, gradient, and likelihood-ratio tests are all
available (fast_ols_cpp/fast_ols_with_var_cpp),
with the parametric-likelihood bootstrap using the OLS Gaussian-errors
model as the generative null. Standard errors are the closed-form OLS
variance under homoskedastic errors (unlike
InferenceContinLin's HC2
heteroskedasticity-robust SE). Warm starts are disabled
(fit_warm_start_enabled = FALSE set at construction) because OLS is
a closed-form estimator and gains nothing from an iterative optimizer's
warm-started initial values. Validity requires the usual OLS assumptions:
correctly specified linear predictor, and (for the asymptotic/likelihood
inference path specifically) homoskedastic, approximately normal errors;
the randomization-inference path relies only on randomization of W.
Super class
Inference -> InferenceContinOLS
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinOLS$new()
Initialize inference for the OLS model Y_i = \beta_0 +
\beta_T W_i + X_i^\top \gamma + \epsilon_i; see
InferenceContinOLS for the model
form. Disables warm-started optimizer initial values (not applicable to
this closed-form estimator). Does not fit the model; the fit is
deferred to the first call to compute_estimate() or a method
that requires it.
Usage
InferenceContinOLS$new( des_obj, model_formula = NULL, verbose = FALSE, max_resample_attempts = 50L, harden = TRUE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a continuous response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
max_resample_attemptsMaximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. Default
50L.hardenWhether to apply robustness measures.
smart_cold_start_defaultFlag for consistent API.
InferenceContinOLS$compute_estimate()
Fits the OLS model
(fast_ols_cpp/fast_ols_with_var_cpp) and
returns \hat\beta_T. If not hardened
(private$harden == FALSE), fits directly on the full design
matrix; otherwise uses QR column-dropping hardening to handle
rank-deficient designs. A fit with a non-finite treatment coefficient
is cached as nonestimable rather than returned.
Usage
InferenceContinOLS$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceContinOLS$compute_estimate_with_bootstrap_weights()
Refits the OLS model with subject/block-level weights applied
to a weighted least-squares fit (stats::lm.wfit) — Bayesian-bootstrap
or nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights() — and
returns the reweighted estimate \hat\beta_T^{(w)}. When
estimate_only = FALSE, also computes a weighted-residual variance
estimate for internal bootstrap diagnostics. Rows with non-finite or
non-positive weight, or non-finite response, are dropped from the
weighted fit; if no rows remain, or the fit fails, the estimate is
NA.
Usage
InferenceContinOLS$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceContinOLS$compute_asymp_confidence_interval()
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Usage
InferenceContinOLS$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
InferenceContinOLS$compute_asymp_two_sided_pval()
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
InferenceContinOLS$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. Default is zero.
InferenceContinOLS$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinOLS$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Rosenbaum, P. R. (2002). Observational Studies (2nd ed.). Springer, for the OLS mean-difference estimator's design-based justification under randomization.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinOLS$new(seq_des)
inf$compute_estimate()
Quantile Regression Inference for Continuous Responses
Description
Fits a linear quantile regression, Q_\tau(Y_i \mid x_i) = x_i^\top\beta_\tau,
for a continuous response, estimating \beta_\tau by minimizing the
asymmetric ("pinball" / "check") loss
\hat\beta_\tau = \operatorname*{arg\,min}_\beta \sum_{i=1}^n \rho_\tau(y_i -
x_i^\top\beta), \qquad \rho_\tau(u) = u\left(\tau - \mathbb{1}[u < 0]\right),
via quantreg's simplex method (quantreg::rq/rq.fit(...,
method = "br")). The treatment coefficient is the estimated shift in the
\tau-th conditional quantile of the response attributable to treatment,
holding any other covariates in model_formula fixed; by default
tau = 0.5, so this is median regression (robust to outliers and
distributional skew relative to mean-based estimators, at the cost of losing
the mean-shift interpretation away from \tau = 0.5).
Standard errors use quantreg's Powell-style "nid" (non-i.i.d.,
kernel-based sparsity/local-density estimator) sandwich covariance when
available, with fallback to the simpler "iid" estimator if needed.
Inference (confidence intervals, p-values) is based on the resulting
asymptotic normal approximation, not an exact finite-sample distribution.
This class requires the quantreg package, which is listed under
Suggests and is not installed automatically with EDI.
Install quantreg manually before use.
Super class
Inference -> InferenceContinQuantileRegr
Methods
Public methods
-
InferenceContinQuantileRegr$compute_estimate_with_bootstrap_weights() -
InferenceContinQuantileRegr$compute_asymp_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinQuantileRegr$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a quantile-regression inference object for a completed design with a continuous, uncensored response. Requires the quantreg package to be installed.
Usage
InferenceContinQuantileRegr$new( des_obj, model_formula = NULL, tau = 0.5, verbose = FALSE )
Arguments
des_objA completed
Designobject with a continuous response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.tauThe quantile
\tau \in (0, 1)to estimate (default 0.5, i.e. median regression).verboseWhether to print progress messages. Default
FALSE.
InferenceContinQuantileRegr$compute_estimate()
Computes the treatment coefficient \hat\beta_{T,\tau}
from a check-loss quantile regression fit at quantile tau (see
class documentation for the full model). Rank-deficient covariate
columns are dropped before fitting (see
private$reduce_design_matrix_for_quantile()); returns NA
if the reduced design has no usable treatment column or too few
residual degrees of freedom.
Usage
InferenceContinQuantileRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceContinQuantileRegr$compute_estimate_with_bootstrap_weights()
Recomputes the quantile-regression treatment estimate under
subject/block bootstrap weights (quantreg::rq(..., weights =
row_weights)), used by the Bayesian bootstrap and related
weighted-resampling machinery; see
InferenceBayesianBootstrap.
Unlike $compute_estimate(), this always reduces the design
matrix from scratch (reuse_factorizations = FALSE) rather than
reusing a cached rank-reduction from a prior warm-started fit.
Usage
InferenceContinQuantileRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceContinQuantileRegr$compute_asymp_confidence_interval()
Computes a 1-\alpha level confidence interval for the
quantile-regression treatment coefficient \hat\beta_{T,\tau},
using quantreg's Powell-style "nid" asymptotic standard
error (falling back to "iid" if unavailable — see class
documentation) and residual degrees of freedom n - p. See
InferenceAsymp for the shared
asymptotic confidence-interval contract this delegates to.
Usage
InferenceContinQuantileRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
InferenceContinQuantileRegr$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\beta_{T,\tau} = \code{delta}, from the same quantreg sandwich
standard error and degrees of freedom used by
$compute_asymp_confidence_interval(). See
InferenceAsymp for the shared
asymptotic two-sided p-value contract this delegates to.
Usage
InferenceContinQuantileRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. Default is zero.
InferenceContinQuantileRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinQuantileRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Koenker, R., and Bassett, G. (1978). "Regression Quantiles." Econometrica, 46(1), 33-50, doi:10.2307/1913643, for the check-loss quantile regression estimator; Koenker, R. (2005). Quantile Regression, Cambridge University Press, for the Powell-style sandwich standard error estimators used here.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinQuantileRegr$new(seq_des)
inf$compute_estimate()
Robust (M/MM-Estimator) Regression Inference for Continuous Responses
Description
Fits a robust linear regression — by default via this package's
fast_robust_regression_cpp C++ backend (use_rcpp = TRUE;
see that page for the full M/MM-estimator model, weight functions, and
asymptotic variance formula), or via MASS::rlm when use_rcpp =
FALSE — for a continuous response using the treatment indicator and,
optionally, all recorded covariates as predictors. This provides a
Huber/MM-style robustness upgrade over ordinary least squares when outcomes
are heavy-tailed or outlier-prone, down-weighting large residuals rather
than letting them dominate the fit the way squared-error loss does.
The method argument is passed through to either backend and may be
either "M" (Huber's psi function) or "MM" (Tukey's bisquare
weight, the default — higher breakdown point than "M" at some cost
in asymptotic efficiency under normality). When use_rcpp = TRUE
(the default), coefficient standard errors come from the C++ backend's own
M-estimator asymptotic variance (ssq_b_j); when FALSE, they
come from the coefficient table returned by summary.rlm(). Either
way, confidence intervals and p-values use that standard error with
residual degrees of freedom n - p in a normal-theory Wald
approximation (not an exact finite-sample distribution).
Super class
Inference -> InferenceContinRobustRegr
Methods
Public methods
-
InferenceContinRobustRegr$compute_estimate_with_bootstrap_weights() -
InferenceContinRobustRegr$compute_asymp_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceContinRobustRegr$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a robust-regression inference object for a completed design with a continuous, uncensored response.
Usage
InferenceContinRobustRegr$new( des_obj, model_formula = NULL, method = "MM", use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a continuous response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.methodRobust estimation method,
"M"(Huber) or"MM"(Tukey bisquare, higher breakdown point). Default"MM".use_rcppWhether to use the
fast_robust_regression_cppC++ backend (TRUE, default) instead ofMASS::rlm(FALSE).verboseWhether to print progress messages. Default
FALSE.smart_cold_start_defaultWhether to use smart starting values for the optimizer.
InferenceContinRobustRegr$compute_estimate()
Computes the robust-regression treatment coefficient
\hat\beta_T from an M/MM-estimator fit (see class documentation
for the full model and use_rcpp backend choice). Rank-deficient
covariate columns are dropped before fitting via
private$fit_with_hardened_qr_column_dropping().
Usage
InferenceContinRobustRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceContinRobustRegr$compute_estimate_with_bootstrap_weights()
Recomputes the robust-regression treatment estimate under
subject/block bootstrap weights, used by the Bayesian bootstrap and
related weighted-resampling machinery; see
InferenceBayesianBootstrap.
When use_rcpp = TRUE, reproduces MASS::rlm's default
wt.method = "inv.var" weighting exactly by pre-multiplying
X and y by \sqrt{\text{weight}} and running
unweighted M-estimation on the transformed data via the C++ backend
(rather than adding native weight support to that backend); when
use_rcpp = FALSE, passes weights directly to
MASS::rlm. This variant never populates a standard error or
degrees of freedom (both left NA) — it is estimate-only by
construction, matching the estimate_only default of the fast
path it always uses internally.
Usage
InferenceContinRobustRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceContinRobustRegr$compute_asymp_confidence_interval()
Computes a 1-\alpha level confidence interval for the
robust-regression treatment coefficient \hat\beta_T, using the
M/MM-estimator's asymptotic standard error (see class documentation
for its source depending on use_rcpp) and residual degrees of
freedom n - p in a normal-theory Wald approximation. See
InferenceAsymp for the shared
asymptotic confidence-interval contract this delegates to.
Usage
InferenceContinRobustRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
InferenceContinRobustRegr$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\beta_T = \code{delta}, from the same M/MM-estimator standard error
and residual degrees of freedom used by
$compute_asymp_confidence_interval(). See
InferenceAsymp for the shared
asymptotic two-sided p-value contract this delegates to.
Usage
InferenceContinRobustRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. Default is zero.
InferenceContinRobustRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceContinRobustRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'continuous')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(10))
inf = InferenceContinRobustRegr$new(seq_des)
inf$compute_estimate()
Hurdle Negative Binomial Regression Inference for Count Responses
Description
Fits a hurdle negative binomial regression for count responses: a binary
hurdle submodel P(Y_i > 0) = \mathrm{logit}^{-1}(X_i^{h\top}
\gamma^h) (fit jointly with the count submodel) crossed with a
zero-truncated negative-binomial count submodel for Y_i \mid Y_i > 0:
\log E[Y_i \mid Y_i > 0, w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top
\gamma, \mathrm{Var}(Y_i \mid Y_i > 0) = \mu_i + \mu_i^2 / \theta
(fast_hurdle_negbin_cpp/
fast_hurdle_negbin_with_var_cpp). The hurdle and count
submodels may use different covariate formulas
(model_formula/model_formula_hurdle). The reported treatment
effect is the coefficient from the conditional
(truncated, Y > 0) count component, on the log-rate scale,
conditional on clearing the hurdle: it is not the effect on the
unconditional mean E[Y], which also depends on how treatment shifts
the hurdle-crossing probability. A marginal (unconditional-mean) estimand
is not yet implemented for this class (see
marginal_estimand_report.md). likelihood_tier = "full":
Wald, gradient, and (bootstrap-calibrated) likelihood-ratio tests are
available for the count submodel's treatment coefficient; a plain score
test is not exposed. Jackknife inference is not supported:
delete-one refits of this two-part model with a jointly-estimated
dispersion parameter are numerically unstable, so
compute_jackknife_estimate() and related methods report explicit
non-estimability rather than attempting delete-one refits.
Super class
Inference -> InferenceCountHurdleNegBin
Methods
Public methods
-
InferenceCountHurdleNegBin$compute_asymp_confidence_interval() -
InferenceCountHurdleNegBin$compute_gradient_two_sided_pval() -
InferenceCountHurdleNegBin$compute_gradient_confidence_interval() -
InferenceCountHurdleNegBin$compute_estimate_with_bootstrap_weights() -
InferenceCountHurdleNegBin$compute_jackknife_bias_estimate() -
InferenceCountHurdleNegBin$compute_jackknife_wald_two_sided_pval() -
InferenceCountHurdleNegBin$compute_jackknife_wald_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountHurdleNegBin$new()
Initialize inference for the two-part hurdle negative
binomial model (binary hurdle submodel plus zero-truncated
negative-binomial count submodel); see
InferenceCountHurdleNegBin
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Usage
InferenceCountHurdleNegBin$new( des_obj, model_formula = NULL, model_formula_hurdle = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment.
model_formula_hurdleFormula for the hurdle submodel. If
NULL(default), it uses the same formula asmodel_formula.verboseA flag indicating whether messages should be displayed.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceCountHurdleNegBin$compute_asymp_confidence_interval()
Compute the hurdle negative-binomial asymptotic confidence
interval for the treatment coefficient, using the shared count-likelihood
semantics documented in
InferenceCountLikelihood.
Usage
InferenceCountHurdleNegBin$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level (default 0.05).
InferenceCountHurdleNegBin$compute_asymp_two_sided_pval()
Compute the hurdle negative-binomial asymptotic two-sided
p-value for the treatment coefficient, falling back through the shared
count-likelihood machinery when needed; see
InferenceCountLikelihood.
Usage
InferenceCountHurdleNegBin$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect (default 0).
InferenceCountHurdleNegBin$compute_gradient_two_sided_pval()
Gradient test of H_0: \beta_T = \code{delta} on the
truncated count submodel's treatment coefficient (a score-test variant
using the observed rather than expected information); see
InferenceCountLikelihood
for the shared likelihood-test dispatch.
Usage
InferenceCountHurdleNegBin$compute_gradient_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect (default 0).
InferenceCountHurdleNegBin$compute_gradient_confidence_interval()
Compute a hurdle negative-binomial likelihood-based confidence
interval by inverting the configured likelihood test. See
InferenceCountLikelihood
for related score, likelihood-ratio, and gradient methods.
Usage
InferenceCountHurdleNegBin$compute_gradient_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level (default 0.05).
InferenceCountHurdleNegBin$compute_estimate_with_bootstrap_weights()
Refits the hurdle negative-binomial model with subject/
block-level weights (Bayesian-bootstrap or nonparametric-bootstrap draw
weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
glmmTMB's glmmTMB(family = truncated_nbinom2()) (not the
package's internal C++ solver, which has no weighted variant for this
model), and returns the reweighted conditional-count log-rate-ratio
estimate \hat\beta_T^{(w)}. Requires the glmmTMB package;
errors if unavailable. No standard error is computed
(s_beta_hat_T is always NA). A fit that fails, or whose
fitted treatment coefficient is missing or non-finite, is cached as
nonestimable and returns NA.
Usage
InferenceCountHurdleNegBin$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceCountHurdleNegBin$compute_jackknife_estimate()
Hurdle negative-binomial delete-one refits are unstable for jackknife inference; report explicit non-estimability.
Usage
InferenceCountHurdleNegBin$compute_jackknife_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".
InferenceCountHurdleNegBin$compute_jackknife_bias_estimate()
Report that the jackknife bias estimate is unavailable for
hurdle negative-binomial fits because delete-one refits are unstable; see
InferenceJackknife for the shared
jackknife contract.
Usage
InferenceCountHurdleNegBin$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".
InferenceCountHurdleNegBin$compute_jackknife_std_error()
Report that the jackknife standard error is unavailable for
hurdle negative-binomial fits because delete-one refits are unstable; see
InferenceJackknife for the shared
jackknife contract.
Usage
InferenceCountHurdleNegBin$compute_jackknife_std_error(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".
InferenceCountHurdleNegBin$compute_jackknife_wald_two_sided_pval()
Reports that jackknife-Wald p-values are unavailable here; see
InferenceJackknife.
Usage
InferenceCountHurdleNegBin$compute_jackknife_wald_two_sided_pval( delta = 0, unit = "auto" )
Arguments
deltaNull treatment-effect value. Default 0.
unitDeletion unit. Default
"auto".
InferenceCountHurdleNegBin$compute_jackknife_wald_confidence_interval()
Reports that jackknife-Wald intervals are unavailable here; see
InferenceJackknife.
Usage
InferenceCountHurdleNegBin$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaSignificance level. Default 0.05.
unitDeletion unit. Default
"auto".
InferenceCountHurdleNegBin$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountHurdleNegBin$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Mullahy, J. (1986). "Specification and Testing of Some Modified Count Data Models." Journal of Econometrics, 33(3), 341-365, doi:10.1016/0304-4076(86)90002-3, for the hurdle count-model framework.
See Also
InferenceCountNegBin for
the single-part negative binomial model this class's count submodel
generalizes to two parts.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountHurdleNegBin$new(seq_des, model_formula = ~ x1)
inf$compute_estimate()
Hurdle Poisson Regression Inference for Count Responses
Description
Fits a hurdle Poisson regression for count responses: a binary hurdle
submodel P(Y_i > 0) = \mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) (fit
jointly with the count submodel) crossed with a zero-truncated Poisson
count submodel for Y_i \mid Y_i > 0: \log E[Y_i \mid Y_i > 0,
w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma. The hurdle and count
submodels may use different covariate formulas
(model_formula/model_formula_hurdle). The reported treatment
effect is the coefficient from the conditional (truncated, Y > 0)
count component, on the log-rate scale, conditional on clearing
the hurdle: it is not the effect on the unconditional mean E[Y],
which also depends on how treatment shifts the hurdle-crossing
probability, under the default estimand = "conditional".
likelihood_tier = "full": Wald, gradient, and (bootstrap-calibrated)
likelihood-ratio tests are available for the count submodel's treatment
coefficient under that estimand; a plain score test is not exposed.
Jackknife inference is not supported: delete-one refits of this
two-part model are numerically unstable, so
compute_jackknife_estimate() and related methods report explicit
non-estimability rather than attempting delete-one refits. Unlike
InferenceCountHurdleNegBin,
the count submodel here assumes Poisson (equidispersion) conditional on
clearing the hurdle, with no separate dispersion parameter.
Marginal (unconditional-mean) estimand. Via
set_estimand(), this class
also supports estimand = "marginal_mean_diff" and
"marginal_ratio": the g-computation average, over the empirical
covariate distribution, of the model-implied unconditional mean
E[Y_i \mid w_i, x_i] = (1 - \pi(x_i)) \cdot \lambda(x_i) / (1 -
e^{-\lambda(x_i)}) — the hurdle-crossing probability times the
zero-truncated Poisson mean, E[Y \mid Y>0] = \lambda / (1 -
e^{-\lambda}) (exact for Poisson: truncating at 0 changes the
normalizing constant but not the rate parameter \lambda; see
Cameron and Trivedi, Regression Analysis of Count Data, ch. 4.2) —
at w_i = 1 vs. w_i = 0. A pure post-fit transform of the same
maximum-likelihood fit (no refit), with a delta-method standard error
against the sandwich-robust covariance matrix already used for this
class's conditional Wald inference. Only "wald"-type inference is
available under a marginal estimand.
Super classes
Inference -> InferenceCountZeroAugmentedPoissonAbstract -> InferenceCountHurdlePoisson
Methods
Public methods
+ inherited public methods from InferenceCountZeroAugmentedPoissonAbstract
InferenceCountZeroAugmentedPoissonAbstract$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_jackknife_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_randomization_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_subsampling_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_estimate_with_bootstrap_weights()InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_bias_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_std_error()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_score_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_score_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_sensitivity()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_wald_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_wald_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$get_information_preference()InferenceCountZeroAugmentedPoissonAbstract$get_information_source_used()InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_diagnostics()InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_estimate_diagnostics()InferenceCountZeroAugmentedPoissonAbstract$get_mod()InferenceCountZeroAugmentedPoissonAbstract$get_summary()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_information_preferences()InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_testing_types()InferenceCountZeroAugmentedPoissonAbstract$get_testing_type()InferenceCountZeroAugmentedPoissonAbstract$select_optimal_b_subsampling()InferenceCountZeroAugmentedPoissonAbstract$select_optimal_m_out_of_n_bootstrap()InferenceCountZeroAugmentedPoissonAbstract$set_information_preference()InferenceCountZeroAugmentedPoissonAbstract$set_testing_type()InferenceCountZeroAugmentedPoissonAbstract$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountHurdlePoisson$new()
Initialize inference for the hurdle Poisson model (binary
hurdle submodel plus zero-truncated Poisson count submodel); see
InferenceCountHurdlePoisson
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Usage
InferenceCountHurdlePoisson$new( des_obj, model_formula = NULL, model_formula_hurdle = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for the count submodel.
model_formula_hurdleFormula for the hurdle submodel. If
NULL(default), it uses the same formula asmodel_formula.use_rcppLogical. If
TRUE(default), use the internal Rcpp implementation. IfFALSE, use glmmTMB.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
optimization_algOptimization algorithm. Default is dispatched via policy.
InferenceCountHurdlePoisson$compute_estimate()
Fits the hurdle Poisson model. Under the default
estimand = "conditional", returns \hat\beta_T, the
treatment log-rate coefficient from the zero-truncated count
submodel (conditional on clearing the hurdle). Under
estimand = "marginal_mean_diff" or "marginal_ratio"
(set via set_estimand()), returns the g-computation
marginal mean difference or log-scale marginal ratio of the
unconditional mean instead — see the class-level @details
for the formula. A pure post-fit transform of the same cached
fit, no refit.
Usage
InferenceCountHurdlePoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceCountHurdlePoisson$compute_asymp_confidence_interval()
Asymptotic confidence interval. Under the conditional
estimand, delegates to the shared zero-augmented count-model Wald/
bootstrap-fallback contract; under a marginal estimand, the
delta-method interval computed by compute_estimate(). Calls
self$compute_estimate() first (not private$shared()
directly) so the estimand-aware cache is always current regardless
of call order.
Usage
InferenceCountHurdlePoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level (default 0.05).
InferenceCountHurdlePoisson$compute_asymp_two_sided_pval()
Asymptotic two-sided p-value, dispatched exactly as
compute_asymp_confidence_interval(); see that method's
description for the marginal-estimand path.
Usage
InferenceCountHurdlePoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect under the current estimand (default 0).
InferenceCountHurdlePoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountHurdlePoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Mullahy, J. (1986). "Specification and Testing of Some Modified Count Data Models." Journal of Econometrics, 33(3), 341-365, doi:10.1016/0304-4076(86)90002-3, for the hurdle count-model framework.
See Also
InferenceCountPoisson for
the single-part Poisson model this class's count submodel generalizes to
two parts; InferenceCountHurdleNegBin
for the overdispersion-robust negative-binomial variant (does not
support a marginal estimand — the mean-function derivation here is
Poisson-specific).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountHurdlePoisson$new(seq_des)
inf$compute_estimate()
One-Likelihood Conditional-Poisson Inference for KK Count Designs
Description
Estimates a treatment log-rate-ratio \beta_T for count outcomes
collected under a KK matching-on-the-fly design
(DesignSeqOneByOneKK14 or
subclass) by maximizing a single combined likelihood that couples a
conditional (within-matched-pair, intercept-free) Poisson likelihood for
matched subjects with an ordinary Poisson likelihood for reservoir
subjects, sharing one treatment coefficient across both pieces. This is
the "one-likelihood" alternative to the inverse-variance-weighted
combination
(...IVWC pattern used
elsewhere in the KK family): rather than fitting matched and reservoir
models separately and pooling by inverse-variance weights, the treatment
coefficient here is estimated jointly from the full combined
log-likelihood, and its standard error, score, likelihood-ratio, and
gradient statistics are all "design-conservative" – each is the
pointwise-wider of the model-based asymptotic quantity and a
design-based quantity computed by treating the estimate as a plug-in
statistic under InferenceAsymp's
z/t machinery, so inference never overstates precision
relative to the design alone.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Estimand. \beta_T, the treatment coefficient in a
log-linear (Poisson) mean model E[Y \mid w, x] = \exp(\beta_0 +
\beta_T w + x\beta), interpreted as a log rate ratio (equivalently,
\exp(\hat\beta_T) is the treatment-vs-control incidence rate
ratio).
Model. Matched subjects contribute a conditional-Poisson term
that eliminates the pair-specific nuisance intercept by conditioning on
the pair total count (removing the need to estimate one intercept per
pair); reservoir subjects contribute an ordinary Poisson log-likelihood
with a single shared intercept. Both pieces are summed into one
combined negative log-likelihood and maximized jointly in
(\beta_0, \beta_T, \beta) (see get_cpoisson_combined_hessian_cpp
and fast_cpoisson_combined_with_var_cpp for the backend fitting
contract). likelihood_tier = "full", so likelihood-ratio, score,
and gradient tests and a parametric likelihood bootstrap are all
available in addition to the design-conservative Wald path.
Assumptions. Independence of counts across matched pairs and
reservoir subjects given covariates; correct log-linear mean
specification; a KK matching-on-the-fly design supplying the
matched/reservoir partition. No response censoring is supported (checked
at construction via assertNoCensoring()).
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceCountKKCondPoissonOneLik
Methods
Public methods
-
InferenceCountKKCondPoissonOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceCountKKCondPoissonOneLik$supports_rand_pval_for_incidence() -
InferenceCountKKCondPoissonOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountKKCondPoissonOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceCountKKCondPoissonOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCountKKCondPoissonOneLik$supports_rand_pval_for_incidence()
Usage
InferenceCountKKCondPoissonOneLik$supports_rand_pval_for_incidence()
InferenceCountKKCondPoissonOneLik$compute_rand_two_sided_pval()
Usage
InferenceCountKKCondPoissonOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCountKKCondPoissonOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountKKCondPoissonOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A. and Krieger, A. (2014). "Matching on-the-fly: A group
sequential covariate balanced randomization procedure." arXiv
preprint arXiv:1305.6259. (KK14 in REFERENCES.md.)
See Also
Analogous Python API for count models:
statsmodels
discrete models (ConditionalPoisson, Poisson).
Poisson
regression (orientation).
GLMM Inference for KK Designs with Count Response
Description
Fits a Poisson GLMM for count responses under a KK matching-on-the-fly design. The random intercept per matched pair is integrated out via Gauss-Hermite quadrature.
When use_rcpp = TRUE (default) the likelihood is maximised by an internal
Rcpp routine. Set use_rcpp = FALSE to fall back to glmmTMB.
Model. Y_{ij} \mid b_i \sim \mathrm{Poisson}(\mu_{ij}) with
\log \mu_{ij} = X_{ij}'\beta + \beta_T \cdot W_{ij} + b_i, where
i indexes matched pairs, j \in \{1, 2\} the two subjects within
a pair, W_{ij} is the treatment indicator, and b_i \sim
\mathcal{N}(0, \sigma_b^2) is a pair-level random intercept absorbing
within-pair correlation induced by matching. \beta_T is a log-rate
(log relative risk) treatment effect: \exp(\hat\beta_T) is the
estimated rate ratio. The random effect is integrated out of the marginal
likelihood by adaptive Gauss-Hermite quadrature rather than a Laplace
approximation.
Likelihood tier. likelihood_tier = "full": both Wald
(model-based standard error) and likelihood-ratio testing types are
available. Because the GLMM likelihood alone does not encode the KK design's
matched-pair randomization structure, the likelihood-ratio CI/p-value are
conservatively widened/calibrated against the design-aware Wald result (see
compute_lik_ratio_confidence_interval()/
compute_lik_ratio_two_sided_pval()) so the model-based test is never
anti-conservative relative to the design.
Assumptions. Count response modeled as conditionally Poisson given
the random intercept (equidispersion conditional on b_i); pair-level
random effects independent across pairs; a KK matching-on-the-fly design.
Super class
Inference -> InferenceCountKKGLMM
Methods
Public methods
-
InferenceCountKKGLMM$compute_estimate_with_bootstrap_weights() -
InferenceCountKKGLMM$compute_lik_ratio_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountKKGLMM$new()
Initialize a KK Poisson-GLMM inference object for a matched-pair KK design with a count response and prepare the matched-pair random-intercept likelihood machinery; see the class topic for the model.
Usage
InferenceCountKKGLMM$new( des_obj, model_formula = NULL, use_rcpp = TRUE, optimization_alg = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed KK matching-on-the-fly
Designobject (DesignSeqOneByOneKK14or subclass) with a count response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used.use_rcppLogical. If
TRUE(default), maximize the Gauss-Hermite-quadrature marginal likelihood with the internal Rcpp Poisson-GLMM routine; ifFALSE, fall back to glmmTMB.optimization_algOptimization algorithm passed to the likelihood maximizer. If
NULL(default), an algorithm is dispatched via the package's optimizer policy.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart starting values for the optimizer.
InferenceCountKKGLMM$compute_estimate()
Point estimate of the treatment log-rate coefficient \beta_T from a
Poisson GLMM with a matched-pair random intercept, fit by maximizing the
Gauss-Hermite-quadrature-integrated marginal likelihood (internal Rcpp routine when
use_rcpp = TRUE, else glmmTMB). See the class topic for the model form.
Usage
InferenceCountKKGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf
TRUE, skip variance-component calculations.
Returns
Numeric scalar: the treatment coefficient on the log-rate (link) scale, i.e.
\exp(\hat\beta_T) is a rate ratio.
InferenceCountKKGLMM$compute_estimate_with_bootstrap_weights()
Recomputes the KK Poisson-GLMM treatment estimate under nonparametric/
Bayesian-bootstrap subject-or-block weights, refitting the weighted GLMM
(compute_weighted_glmm_bootstrap_estimate()). Standard error, degrees of
freedom, and the cached summary table are cleared/set to NA/Inf/
NULL since only the point estimate is meaningful under resampling weights.
Falls back to the unweighted point estimate when the weights are effectively
constant.
Usage
InferenceCountKKGLMM$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric vector of nonnegative bootstrap replicate weights, one per subject or per matched block (KK match structure).
estimate_onlyIf
TRUE, compute only the weighted point estimate (this method never computes a weighted standard error regardless of this argument).
Returns
Numeric scalar treatment-effect estimate (log-rate scale) under the given weights.
InferenceCountKKGLMM$compute_wald_confidence_interval()
Wald confidence interval for the treatment log-rate coefficient:
\hat\beta_T \pm t_{1-\alpha/2,\,df}\cdot \hat{se}(\hat\beta_T), using the
model-based GLMM standard error. See
InferenceAsymp for the shared contract.
Usage
InferenceCountKKGLMM$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
Returns
A length-2 numeric vector c(lower, upper) on the log-rate scale.
InferenceCountKKGLMM$compute_wald_two_sided_pval()
Two-sided Wald p-value for H_0: \beta_T = \code{delta} vs.
H_1: \beta_T \neq \code{delta}, using the model-based GLMM standard error.
Usage
InferenceCountKKGLMM$compute_wald_two_sided_pval(delta = 0)
Arguments
deltaThe null value of
\beta_Tto test against; 0 (the default) tests for any treatment effect at all.
Returns
Numeric scalar p-value in [0, 1].
InferenceCountKKGLMM$compute_lik_ratio_confidence_interval()
Likelihood-ratio confidence interval for the treatment log-rate
coefficient, inverting the GLMM's profile likelihood-ratio test against
\chi^2_1. Because the GLMM likelihood does not itself account for the KK
matched-pair design's randomization structure, this interval is conservatively
widened to be at least as wide as the design-aware Wald interval
(compute_wald_confidence_interval()) via .conservative_kk_onelik_ci() –
guarding against the model-based interval being anti-conservative relative to the
design.
Usage
InferenceCountKKGLMM$compute_lik_ratio_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
Returns
A length-2 numeric vector c(lower, upper) on the log-rate scale.
InferenceCountKKGLMM$compute_lik_ratio_two_sided_pval()
Two-sided likelihood-ratio p-value for H_0: \beta_T = \code{delta},
from the GLMM's profile likelihood-ratio test referred to \chi^2_1. As with
compute_lik_ratio_confidence_interval(), this is conservatively calibrated
(via .conservative_kk_onelik_pval()) against the design-aware Wald p-value so
the model-based test cannot be anti-conservative relative to the KK matched-pair
design.
Usage
InferenceCountKKGLMM$compute_lik_ratio_two_sided_pval(delta = 0)
Arguments
deltaThe null value of
\beta_Tto test against; 0 (the default) tests for any treatment effect at all.
Returns
Numeric scalar p-value in [0, 1].
InferenceCountKKGLMM$compute_asymp_confidence_interval()
Asymptotic confidence interval, dispatching to
compute_wald_confidence_interval() or compute_lik_ratio_confidence_interval()
depending on self$get_testing_type() (defaults to Wald if the testing type is
neither). See InferenceAsymp for the shared
testing-type dispatch contract.
Usage
InferenceCountKKGLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
Returns
A length-2 numeric vector c(lower, upper) on the log-rate scale.
InferenceCountKKGLMM$compute_asymp_two_sided_pval()
Asymptotic two-sided p-value, dispatching to
compute_wald_two_sided_pval() or compute_lik_ratio_two_sided_pval()
depending on self$get_testing_type() (defaults to Wald if the testing type is
neither). See InferenceAsymp for the shared
testing-type dispatch contract.
Usage
InferenceCountKKGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null value of
\beta_Tto test against; 0 (the default) tests for any treatment effect at all.
Returns
Numeric scalar p-value in [0, 1].
InferenceCountKKGLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountKKGLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A. and Krieger, A. M. (2014). Matching on-the-fly: Sequential
allocation with higher power and efficiency. Biometrics, 70(2),
378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)
See Also
Analogous Python API for Poisson/count GLMs: statsmodels discrete models. Generalized linear model and Gauss-Hermite quadrature (orientation).
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountKKGLMM$new(seq_des)
inf$compute_estimate()
KK Hurdle Poisson IVWC Inference for Count Responses
Description
Inverse-variance weighted combined inference for count responses under a KK matching-on-the-fly design. The matched-pair component is fit with a hurdle-Poisson mixed model using pair random intercepts, and the reservoir component is fit with an ordinary Poisson log-link regression. The reported treatment effect is on the log-rate scale.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceCountKKHurdlePoissonIVWC
Methods
Public methods
-
InferenceCountKKHurdlePoissonIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceCountKKHurdlePoissonIVWC$supports_rand_pval_for_incidence() -
InferenceCountKKHurdlePoissonIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountKKHurdlePoissonIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceCountKKHurdlePoissonIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCountKKHurdlePoissonIVWC$supports_rand_pval_for_incidence()
Usage
InferenceCountKKHurdlePoissonIVWC$supports_rand_pval_for_incidence()
InferenceCountKKHurdlePoissonIVWC$compute_rand_two_sided_pval()
Usage
InferenceCountKKHurdlePoissonIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCountKKHurdlePoissonIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountKKHurdlePoissonIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
KK Hurdle-Poisson Combined-Likelihood Inference for Count Responses
Description
Fits a two-part hurdle-Poisson model to a KK matching-on-the-fly count
design by maximizing a single combined likelihood over the matched pairs
and reservoir subjects jointly, rather than fitting matched and reservoir
submodels separately and combining them afterward (contrast with the
inverse-variance-weighted-combination sibling,
InferenceCountKKHurdlePoissonIVWC).
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Model. For subject i with count response Y_i \ge 0, a
hurdle model factors the likelihood into (1) a binary zero-vs-positive part
P(Y_i = 0) = 1 - \pi_i, \mathrm{logit}(\pi_i) = X_i'\gamma, and
(2) a zero-truncated Poisson part for the positive counts,
Y_i \mid Y_i > 0 \sim \text{Poisson}_{+}(\lambda_i),
\log \lambda_i = X_i'\beta + \beta_T W_i, where W_i is the
treatment indicator and \beta_T is the treatment log-rate coefficient
for the positive-count submodel (the estimand returned by
compute_estimate()). Unlike a standard hurdle model fit by maximum
likelihood on i.i.d. rows, this class's negative log-likelihood combines
the matched-pair rows and reservoir rows of a KK design into one objective
(see private$fit_combined_hurdle()), so the fitted \beta_T
and its curvature already reflect the design's matched/reservoir
structure rather than treating all subjects as exchangeable.
Likelihood tier. likelihood_tier = "full": Wald, score,
likelihood-ratio, and gradient testing types are all available (see
get_testing_type()). Because a hurdle-Poisson combined likelihood
does not by itself encode the KK design's finite-sample matched-pair
randomization distribution, the score/likelihood-ratio/gradient
confidence intervals and p-values are computed twice — once from this
class's design-aware asymptotic variance (the same Wald-type calculation
used by compute_wald_confidence_interval()) and once from the
generic likelihood-based calculation inherited from
InferenceAsympLik — and the wider
interval / larger p-value of the two is returned (see
.conservative_kk_onelik_ci()/.conservative_kk_onelik_pval()),
so the model-based test is never anti-conservative relative to the design.
The Wald confidence interval and p-value fall back to the
BayesianBootstrap component's
bootstrap distribution when the model-based standard error is unavailable
or non-finite (e.g. a boundary/separation fit).
Assumptions. Independence across matched pairs and reservoir
subjects conditional on covariates; correct specification of the
logistic hurdle and log-linear positive-count submodels; a KK
matching-on-the-fly design (DesignSeqOneByOneKK14
or subclass) supplying the matched/reservoir partition. No response
censoring is supported (checked at construction via
assertNoCensoring()).
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceCountKKHurdlePoissonOneLik
Methods
Public methods
-
InferenceCountKKHurdlePoissonOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceCountKKHurdlePoissonOneLik$supports_rand_pval_for_incidence() -
InferenceCountKKHurdlePoissonOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountKKHurdlePoissonOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceCountKKHurdlePoissonOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCountKKHurdlePoissonOneLik$supports_rand_pval_for_incidence()
Usage
InferenceCountKKHurdlePoissonOneLik$supports_rand_pval_for_incidence()
InferenceCountKKHurdlePoissonOneLik$compute_rand_two_sided_pval()
Usage
InferenceCountKKHurdlePoissonOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCountKKHurdlePoissonOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountKKHurdlePoissonOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Mullahy, J. (1986). "Specification and Testing of Some Modified Count
Data Models." Journal of Econometrics, 33(3), 341-365.
doi:10.1016/0304-4076(86)90002-3. (Mullahy1986 in REFERENCES.md.)
See Also
Analogous Python API for hurdle/zero-truncated count models:
statsmodels
discrete models (HurdleCountModel).
Poisson
regression (orientation).
Count-Specific Likelihood Inference
Description
Component source (CountLikelihoodPlumbingSource) for the
count-based likelihood families (Poisson, Negative Binomial, Zero-Inflated,
Hurdle): centralizes count-specific parameter packing, warm starts, and
likelihood dispatch. Composed by every count class through the registered
CountLikelihoodPlumbing component.
Computes the treatment-effect estimate using the underlying
count likelihood model. Concrete subclasses fit a Poisson, negative
binomial, zero-inflated, hurdle, or combined likelihood model and cache
the treatment coefficient for related
InferenceCountLikelihood
p-value and confidence-interval methods.
Usage
CountLikelihoodPlumbingSource
Negative Binomial Regression Inference for Count Responses
Description
Fits a negative binomial regression for count responses:
Y_i \mid w_i, x_i \sim \mathrm{NegBin}(\mu_i, \theta), \log
\mu_i = \beta_0 + \beta_T w_i + x_i^\top \gamma, \mathrm{Var}(Y_i) =
\mu_i + \mu_i^2 / \theta, jointly maximizing over the regression
coefficients and the dispersion parameter \theta
(fast_neg_bin_cpp/fast_neg_bin_with_var_cpp).
\hat\beta_T is a log-rate-ratio: \exp(\hat\beta_T) is the
estimated rate ratio. Unlike InferenceCountPoisson,
the negative-binomial model allows overdispersion (\mathrm{Var}(Y_i) >
E[Y_i]) via \theta; smaller \theta indicates more
overdispersion, and the model converges to Poisson as \theta \to
\infty. likelihood_tier = "full": Wald, score, gradient, and
likelihood-ratio tests are all available, plus parametric-likelihood
bootstrap calibration of the likelihood-ratio test (simulating new
responses from \mathrm{NegBin}(\hat\mu_i, \hat\theta) under the null).
Jackknife inference is not supported: delete-one refits of a
jointly-estimated dispersion parameter are numerically unstable, so
compute_jackknife_estimate() and related methods report explicit
non-estimability rather than attempting delete-one refits. Validity
requires the negative-binomial mean-variance relationship to hold and the
usual correctly-specified-linear-predictor-on-the-log-scale assumption.
Super class
Inference -> InferenceCountNegBin
Methods
Public methods
-
InferenceCountNegBin$compute_estimate_with_bootstrap_weights() -
InferenceCountNegBin$compute_jackknife_wald_two_sided_pval() -
InferenceCountNegBin$compute_jackknife_wald_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountNegBin$new()
Initialize inference for the negative binomial regression
model Y_i \mid w_i, x_i \sim \mathrm{NegBin}(\mu_i, \theta),
\log \mu_i = \beta_0 + \beta_T w_i + x_i^\top \gamma; see
InferenceCountNegBin for the
model form. Does not fit the model; the fit is deferred to the first
call to compute_estimate() or a method that requires it.
Usage
InferenceCountNegBin$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart optimizer start values by default.
optimization_algOptimization algorithm to use. Default is dispatched via policy.
InferenceCountNegBin$compute_estimate_with_bootstrap_weights()
Refits the negative binomial model with subject/block-level
weights applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_neg_bin_weighted_cpp, and returns the reweighted
log-rate-ratio estimate \hat\beta_T^{(w)}. If the weighted
negative-binomial fit fails to converge, falls back to a weighted
Poisson GLM (stats::glm(family = poisson())) as an
estimating-equation-consistent point estimate of the same mean
structure (this fallback does not itself estimate \theta, so no
standard error is computed in that path); no standard error is computed
in either path (s_beta_hat_T is always NA).
Usage
InferenceCountNegBin$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceCountNegBin$compute_jackknife_estimate()
Negative-binomial delete-one refits are unstable for jackknife inference; report explicit non-estimability.
Usage
InferenceCountNegBin$compute_jackknife_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".
InferenceCountNegBin$compute_jackknife_bias_estimate()
Report that the jackknife bias estimate is unavailable for
negative-binomial fits when delete-one refits are not stable; see
InferenceJackknife for the shared
jackknife contract.
Usage
InferenceCountNegBin$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".
InferenceCountNegBin$compute_jackknife_std_error()
Report that the jackknife standard error is unavailable for
negative-binomial fits when delete-one refits are not stable; see
InferenceJackknife for the shared
jackknife contract.
Usage
InferenceCountNegBin$compute_jackknife_std_error(unit = "auto")
Arguments
unitDeletion unit. Default
"auto".
InferenceCountNegBin$compute_jackknife_wald_two_sided_pval()
Reports that jackknife-Wald p-values are unavailable here; see
InferenceJackknife.
Usage
InferenceCountNegBin$compute_jackknife_wald_two_sided_pval( delta = 0, unit = "auto" )
Arguments
deltaNull treatment-effect value. Default 0.
unitDeletion unit. Default
"auto".
InferenceCountNegBin$compute_jackknife_wald_confidence_interval()
Reports that jackknife-Wald intervals are unavailable here; see
InferenceJackknife.
Usage
InferenceCountNegBin$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaSignificance level. Default 0.05.
unitDeletion unit. Default
"auto".
InferenceCountNegBin$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountNegBin$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Cameron, A. C., and Trivedi, P. K. (2013). Regression Analysis of Count Data (2nd ed.). Cambridge University Press, for the negative binomial regression model and its maximum-likelihood theory.
See Also
Comparable Python API:
statsmodels
discrete models (NegativeBinomial). See also:
Negative
binomial distribution (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountNegBin$new(seq_des)
inf$compute_estimate()
Poisson Regression Inference for Count Responses
Description
Fits a Poisson log-link regression for count responses:
Y_i \mid w_i, x_i \sim \mathrm{Poisson}(\mu_i), \log \mu_i =
\beta_0 + \beta_T w_i + x_i^\top \gamma
(fast_poisson_regression_cpp/
fast_poisson_regression_with_var_cpp). \hat\beta_T is a
log-rate-ratio: \exp(\hat\beta_T) is the estimated rate ratio.
likelihood_tier = "full": Wald, score, gradient, and
likelihood-ratio tests are all available, plus parametric-likelihood
bootstrap calibration of the likelihood-ratio test (simulating new Poisson
responses under the null).
Design-conservative testing. Every asymptotic/likelihood test
method on this class does not report the raw model-based result directly.
Instead, it also computes a design-based jackknife-Wald test
(compute_jackknife_wald_two_sided_pval()/
compute_jackknife_wald_confidence_interval(), which do not assume
the Poisson mean-variance relationship) and combines the two
conservatively: p-values report \max of the model-based and
design-based p-values, and confidence intervals report the union of
the model-based and design-based intervals. This guards against the
model-based test being anti-conservative when the Poisson equidispersion
assumption (\mathrm{Var}(Y_i) = E[Y_i]) fails — a real risk for count
data, which is frequently overdispersed (see
InferenceCountNegBin for a model
that estimates dispersion directly instead). If either component is
unavailable, the available one is used alone; if neither is available, the
result is NA. Validity requires the usual correctly-specified
linear predictor on the log scale; unlike the raw Poisson likelihood
alone, this class's actual reported inference degrades gracefully (rather
than becoming anti-conservative) under mean-variance misspecification.
Estimand. Composes
MarginalEstimand
(set_estimand()/get_estimand()/get_supported_estimands()).
Under the default estimand = "conditional", \hat\beta_T is the
log-rate-ratio above. Under estimand = "marginal_mean_diff", the
reported quantity is instead the g-computation marginal rate difference
\frac{1}{n}\sum_i \{\exp(\hat\beta_0 + \hat\beta_T + X_i^\top
\hat\gamma) - \exp(\hat\beta_0 + X_i^\top \hat\gamma)\}. Under
estimand = "marginal_ratio", the log of the corresponding marginal
rate ratio — which is numerically identical to the conditional
\hat\beta_T for this family: because the log link is linear in
w_i with no treatment-by-covariate interaction term, every subject's
treated-vs-control mean ratio is \exp(\hat\beta_0 + \hat\beta_T +
X_i^\top\hat\gamma) / \exp(\hat\beta_0 + X_i^\top\hat\gamma) =
\exp(\hat\beta_T) exactly, so averaging over subjects before or after
taking the ratio makes no difference. "marginal_ratio" is provided
for estimand-API consistency with the other model families, not because it
differs numerically from "conditional" here; "marginal_mean_diff"
is the estimand where g-computation actually changes the reported number
for a Poisson GLM, since a difference (unlike a ratio) does not collapse
under a nonlinear (log) mean function. Because there is no latent submodel
for this family (unlike e.g.
InferenceCountZeroInflatedPoisson's
excess-zero mixture), the marginal mean function is exactly the model's
own fitted mean; no separate standardization step beyond the g-computation
average is needed. Standard errors under a marginal estimand use the delta
method against the model's coefficient covariance (degrees of freedom
Inf), including in the design-conservative union/max combination
above (the design-based jackknife-Wald component also refits under the
active estimand); testing_type is restricted to "wald"
whenever the estimand is non-conditional. The underlying model fit is
identical regardless of estimand — switching estimand is a pure
post-fit transform, never a refit.
Super class
Inference -> InferenceCountPoisson
Methods
Public methods
-
InferenceCountPoisson$compute_lik_ratio_confidence_interval() -
InferenceCountPoisson$compute_gradient_confidence_interval() -
InferenceCountPoisson$compute_lik_ratio_bootstrap_two_sided_pval() -
InferenceCountPoisson$compute_lik_ratio_bootstrap_confidence_interval() -
InferenceCountPoisson$compute_estimate_with_bootstrap_weights()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountPoisson$new()
Initialize inference for the Poisson regression model
Y_i \mid w_i, x_i \sim \mathrm{Poisson}(\mu_i), \log \mu_i =
\beta_0 + \beta_T w_i + x_i^\top \gamma; see
InferenceCountPoisson for the
model form and the design-conservative testing mechanism. Does not fit
the model; the fit is deferred to the first call to
compute_estimate() or a method that requires it.
Usage
InferenceCountPoisson$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL, harden = TRUE )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
hardenWhether to apply robustness measures.
InferenceCountPoisson$compute_estimate()
Fits the Poisson regression model by maximum likelihood.
Under the default estimand = "conditional", returns
\hat\beta_T, the treatment log-rate-ratio. Under
estimand = "marginal_mean_diff"/"marginal_ratio" (set
via set_estimand()), returns the g-computation marginal
rate difference/log-rate-ratio instead — see the class-level
@details for the formula. The underlying model fit is
identical either way (a pure post-fit transform of the same
cached fit, no refit).
Usage
InferenceCountPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceCountPoisson$compute_asymp_confidence_interval()
Design-conservative confidence interval for \beta_T
using whichever test type is configured
(private$testing_type: "wald", "score",
"gradient", or "lik_ratio"); see
InferenceCountPoisson for the
union-with-jackknife-Wald combination rule.
Usage
InferenceCountPoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceCountPoisson$compute_asymp_two_sided_pval()
Design-conservative two-sided p-value for H_0: \beta_T =
\code{delta} using whichever test type is configured
(private$testing_type); see
InferenceCountPoisson for the
max-with-jackknife-Wald combination rule.
Usage
InferenceCountPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
InferenceCountPoisson$compute_wald_confidence_interval()
Wald confidence interval for \beta_T using the fitted
Poisson model's Fisher-information-based standard error, unioned with
the design-based jackknife-Wald interval; see
InferenceCountPoisson for the
combination rule and InferenceAsymp
for the underlying Wald contract.
Usage
InferenceCountPoisson$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceCountPoisson$compute_wald_two_sided_pval()
Wald test of H_0: \beta_T = \code{delta} using the
fitted Poisson model's Fisher-information-based standard error, taking
the max with the design-based jackknife-Wald p-value; see
InferenceCountPoisson for the
combination rule.
Usage
InferenceCountPoisson$compute_wald_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
InferenceCountPoisson$compute_score_confidence_interval()
Score-test confidence interval for \beta_T (inverting
the Poisson score test at each candidate null, no full re-fit needed at
the observed information), unioned with the design-based jackknife-Wald
interval; see InferenceCountPoisson
for the combination rule.
Usage
InferenceCountPoisson$compute_score_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceCountPoisson$compute_score_two_sided_pval()
Score test of H_0: \beta_T = \code{delta}, taking the
max with the design-based jackknife-Wald p-value; see
InferenceCountPoisson for the
combination rule.
Usage
InferenceCountPoisson$compute_score_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
InferenceCountPoisson$compute_lik_ratio_confidence_interval()
Likelihood-ratio-test confidence interval for \beta_T
(test inversion, requiring a null refit at each candidate value),
unioned with the design-based jackknife-Wald interval; see
InferenceCountPoisson for the
combination rule.
Usage
InferenceCountPoisson$compute_lik_ratio_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceCountPoisson$compute_lik_ratio_two_sided_pval()
Likelihood-ratio test of H_0: \beta_T = \code{delta},
taking the max with the design-based jackknife-Wald p-value; see
InferenceCountPoisson for the
combination rule.
Usage
InferenceCountPoisson$compute_lik_ratio_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
InferenceCountPoisson$compute_gradient_confidence_interval()
Gradient-test confidence interval for \beta_T (a
score-test variant using the observed rather than expected
information), unioned with the design-based jackknife-Wald interval;
see InferenceCountPoisson for
the combination rule.
Usage
InferenceCountPoisson$compute_gradient_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default 0.05.
InferenceCountPoisson$compute_gradient_two_sided_pval()
Gradient test of H_0: \beta_T = \code{delta}, taking
the max with the design-based jackknife-Wald p-value; see
InferenceCountPoisson for the
combination rule.
Usage
InferenceCountPoisson$compute_gradient_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect. Default 0.
InferenceCountPoisson$compute_lik_ratio_bootstrap_two_sided_pval()
Parametric-likelihood-bootstrap-calibrated likelihood-ratio
test of H_0: \beta_T = \code{delta} (simulating new Poisson
responses from the null-constrained fit to calibrate the LR statistic's
null distribution), taking the max with the design-based jackknife-Wald
p-value; see InferenceCountPoisson
for the combination rule.
Usage
InferenceCountPoisson$compute_lik_ratio_bootstrap_two_sided_pval( delta = 0, B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L )
Arguments
deltaNull treatment effect. Default 0.
BNumber of bootstrap replicates.
show_progressWhether to show progress.
min_number_usable_samplesMinimum usable bootstrap samples.
max_attempts_per_replicateMaximum attempts per replicate.
InferenceCountPoisson$compute_lik_ratio_bootstrap_confidence_interval()
Parametric-likelihood-bootstrap-calibrated likelihood-ratio
confidence interval for \beta_T (test inversion using the
bootstrap-calibrated null distribution), unioned with the design-based
jackknife-Wald interval; see
InferenceCountPoisson for the
combination rule.
Usage
InferenceCountPoisson$compute_lik_ratio_bootstrap_confidence_interval( alpha = 0.05, B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L, root_tolerance = NULL, max_root_iterations = 8L )
Arguments
alphaSignificance level. Default 0.05.
BNumber of bootstrap replicates.
show_progressWhether to show progress.
min_number_usable_samplesMinimum usable bootstrap samples.
max_attempts_per_replicateMaximum attempts per replicate.
root_toleranceRoot tolerance.
max_root_iterationsMaximum root iterations.
InferenceCountPoisson$compute_estimate_with_bootstrap_weights()
Refits the Poisson model with subject/block-level weights
applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_poisson_regression_weighted_cpp, and returns the
reweighted log-rate-ratio estimate \hat\beta_T^{(w)}. Uses the
same QR column-dropping hardening as the unweighted fit; a hardened fit
with a non-finite treatment coefficient is cached as nonestimable and
returns NA. The weighted refit always targets the conditional
treatment coefficient, whatever the active estimand (it is not
estimand-aware).
Usage
InferenceCountPoisson$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsRow weights for the bootstrap sample.
estimate_onlyIf TRUE, skip variance calculations.
InferenceCountPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Cameron, A. C., and Trivedi, P. K. (2013). Regression Analysis of Count Data (2nd ed.). Cambridge University Press, for the Poisson regression model and its maximum-likelihood theory.
See Also
Comparable Python API:
statsmodels
discrete models (Poisson). See also:
Poisson
regression (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountPoisson$new(seq_des)
inf$compute_estimate()
inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)
GEE Inference for KK Designs with Count Response
Description
Fits a Generalized Estimating Equations (GEE) model with a Poisson family and
log link, \log E[Y_i \mid x_i] = x_i^\top\beta, for count responses
under a KK matching-on-the-fly design, using an exchangeable working
correlation structure where each cluster is either a matched pair (2 members)
or a reservoir singleton (1 member) — see
$compute_estimate()'s method-level documentation for the full fitting
contract (internal Rcpp solver vs. geepack fallback, hardening/retry
behavior). GEE is used here purely to fit one marginal model jointly across
matched-pair and reservoir subjects while accounting for the within-pair
correlation the matching induces, not as a longitudinal/repeated-measures
tool. Inference is quasi-likelihood/estimating-equation based
(likelihood_tier = "quasi"): standard errors are GEE sandwich (robust)
standard errors, not model-likelihood-based.
Super class
Inference -> InferenceCountPoissonKKGEE
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountPoissonKKGEE$new()
Initialize KK count-response GEE inference, validate the
matched/reservoir design, and prepare the exchangeable-working-correlation
Poisson (log-link) GEE fitting machinery used by
InferenceCountPoissonKKGEE.
Usage
InferenceCountPoissonKKGEE$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for covariate adjustment.
use_rcppWhether to use the internal Rcpp GEE solver (
TRUE, default) with automatic fallback togeepack::geeglmon failure, or always usegeepack::geeglmdirectly (FALSE).verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceCountPoissonKKGEE$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountPoissonKKGEE$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the GEE estimating-equation framework and sandwich variance estimator used here.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountPoissonKKGEE$new(seq_des)
inf$compute_estimate()
Quasi-Poisson Regression Inference for Count Responses
Description
Fits a Poisson log-link mean model, \log E[Y_i \mid x_i] =
x_i^\top\beta, for count responses using the treatment indicator and,
optionally, all recorded covariates as predictors, via
fast_quasipoisson_regression_with_var_cpp — see that page for
the full model and the Pearson-dispersion-scaled ("quasi-Poisson") variance
formula, \widehat{\mathrm{Var}}(\hat\beta_k) = \hat\phi\,[(X^\top
\hat{W}X)^{-1}]_{kk}, which corrects standard errors for overdispersion
(\mathrm{Var}(Y_i) > E[Y_i]) relative to the strict Poisson assumption
without changing the point estimate \hat\beta. This class has no
likelihood-ratio/score/gradient testing capability
(likelihood_tier = "quasi"): the dispersion-scaled quasi-likelihood is
not a normalized model likelihood, so only Wald inference is available.
Rank-deficient covariate columns are dropped automatically before fitting
(via private$fit_with_hardened_qr_column_dropping()).
Estimand. Composes
MarginalEstimand
(set_estimand()/get_estimand()/get_supported_estimands()).
Under the default estimand = "conditional", \hat\beta_T is the
treatment log-rate-ratio. Under "marginal_mean_diff" it is the
g-computed difference in the average fitted count under treatment vs. control,
\frac{1}{n}\sum_i \{\exp(x_{i1}^\top\hat\beta) -
\exp(x_{i0}^\top\hat\beta)\}, with every subject plugged in at treatment 1
and 0. "marginal_ratio" is the log of the corresponding ratio, which for
this log-link family equals the conditional \hat\beta_T exactly (there is
no treatment-by-covariate term), so it is offered for estimand-API consistency;
its delta-method SE equals the conditional SE. Under a marginal estimand the
standard error is the delta-method SE against the dispersion-scaled coefficient
covariance \hat\phi (X^\top \hat W X)^{-1}, i.e. the plain-Poisson
marginal SE inflated by \sqrt{\hat\phi}, with a normal reference
(degrees of freedom Inf). Switching the estimand is a pure post-fit
transform of the cached fit, never a refit. The Bayesian-bootstrap weighted
refit is not estimand-aware: it always targets the conditional coefficient.
Super class
Inference -> InferenceCountQuasiPoisson
Methods
Public methods
-
InferenceCountQuasiPoisson$compute_estimate_with_bootstrap_weights() -
InferenceCountQuasiPoisson$compute_asymp_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountQuasiPoisson$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a quasi-Poisson regression inference object for a completed design with a count, uncensored response.
Usage
InferenceCountQuasiPoisson$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL, harden = TRUE )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
hardenWhether to apply robustness measures.
InferenceCountQuasiPoisson$compute_estimate()
Computes the quasi-Poisson point estimate via
fast_quasipoisson_regression_with_var_cpp (see class
documentation for the full model). Under the default
estimand = "conditional" this is the treatment coefficient
\hat\beta_T; under "marginal_mean_diff" or
"marginal_ratio" (set via set_estimand()) it is the
g-computed marginal mean difference / log ratio, a pure post-fit
transform of the same cached fit (no refit). Rank-deficient covariate
columns are dropped before fitting.
Usage
InferenceCountQuasiPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
InferenceCountQuasiPoisson$compute_estimate_with_bootstrap_weights()
Recomputes the Poisson-mean-model treatment estimate under
subject/block bootstrap weights (via
fast_poisson_regression_weighted_cpp), used by the
Bayesian bootstrap and related weighted-resampling machinery; see
InferenceBayesianBootstrap.
When estimate_only = FALSE, also computes a weighted
Pearson-dispersion-scaled standard error (fixed 2026-09-07 –
previously always NA regardless of estimate_only,
which starved the Bayesian-bootstrap studentized/BCa variants of a
per-replicate SE and left them NA on the large majority of calls).
The weighted refit always targets the conditional treatment
coefficient, whatever the active estimand (it is not estimand-aware).
Usage
InferenceCountQuasiPoisson$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip the dispersion-correction computation.
InferenceCountQuasiPoisson$compute_asymp_confidence_interval()
Computes a 1-\alpha level Wald confidence interval for the
active estimand. Under the default conditional estimand this is the
quasi-Poisson treatment coefficient \hat\beta_T with the
Pearson-dispersion-scaled standard error from
fast_quasipoisson_regression_with_var_cpp (see class
documentation); under a marginal estimand it is the g-computed functional
with its delta-method standard error. Both use a normal reference. See
InferenceAsymp for the shared
asymptotic confidence-interval contract this delegates to.
Usage
InferenceCountQuasiPoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaConfidence level.
InferenceCountQuasiPoisson$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\beta_T = \code{delta} under the conditional estimand, or that the
active marginal functional equals delta under a marginal
estimand, from the same standard error used by
$compute_asymp_confidence_interval() (normal reference). See
InferenceAsymp for the shared
asymptotic two-sided p-value contract this delegates to.
Usage
InferenceCountQuasiPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect value.
InferenceCountQuasiPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountQuasiPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountQuasiPoisson$new(seq_des)
inf$compute_estimate()
Robust (Sandwich-Variance) Poisson Regression Inference for Count Responses
Description
Fits the same Poisson log-link mean model as
InferenceCountPoisson (point
estimate via fast_poisson_regression_cpp, maximum likelihood),
but computes standard errors via a Huber-White (Eicker-Huber-White)
sandwich estimator instead of the model-based Poisson Fisher information or
the quasi-Poisson dispersion scaling used by
InferenceCountQuasiPoisson:
\widehat{\mathrm{Var}}(\hat\beta) = B\,M\,B, with "bread" B =
(X^\top \hat W X)^{-1} (the Poisson Fisher information at \hat\beta)
and "meat" M = X^\top \mathrm{diag}((y_i-\hat\mu_i)^2) X (the empirical
score outer product), via robust_sandwich_variance_from_xtwx(). This
is robust to arbitrary mean-variance misspecification (not just
proportional overdispersion), at the cost of somewhat higher variance in the
SE estimate itself for small samples. This class has no likelihood-ratio/
score/gradient testing capability (likelihood_tier = "quasi"): only
Wald inference is available. Rank-deficient covariate columns are dropped
automatically before fitting.
Super class
Inference -> InferenceCountRobustPoisson
Methods
Public methods
-
InferenceCountRobustPoisson$compute_estimate_with_bootstrap_weights() -
InferenceCountRobustPoisson$compute_asymp_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountRobustPoisson$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a robust (sandwich-variance) Poisson regression inference object for a completed design with a count, uncensored response.
Usage
InferenceCountRobustPoisson$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart starting values for the optimizer.
InferenceCountRobustPoisson$compute_estimate()
Computes the Poisson treatment coefficient \hat\beta_T
via fast_poisson_regression_cpp (see class documentation
for the sandwich-variance model). Rank-deficient covariate columns are
dropped before fitting.
Usage
InferenceCountRobustPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
InferenceCountRobustPoisson$compute_estimate_with_bootstrap_weights()
Recomputes the Poisson-mean-model treatment estimate under
subject/block bootstrap weights (via
fast_poisson_regression_weighted_cpp), used by the
Bayesian bootstrap and related weighted-resampling machinery; see
InferenceBayesianBootstrap.
When estimate_only = FALSE, also computes a weighted
Huber-White sandwich standard error (fixed 2026-09-07 – previously
always NA regardless of estimate_only, which starved
the Bayesian-bootstrap studentized/BCa variants of a per-replicate
SE and left them NA on the large majority of calls).
Usage
InferenceCountRobustPoisson$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip the sandwich-variance computation.
InferenceCountRobustPoisson$compute_asymp_confidence_interval()
Computes a 1-\alpha level confidence interval for the
robust Poisson treatment coefficient \hat\beta_T, using the
Huber-White sandwich standard error (see class documentation). See
InferenceAsymp for the shared
asymptotic confidence-interval contract this delegates to.
Usage
InferenceCountRobustPoisson$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaConfidence level.
InferenceCountRobustPoisson$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\beta_T = \code{delta}, from the same Huber-White sandwich standard
error used by $compute_asymp_confidence_interval(). See
InferenceAsymp for the shared
asymptotic two-sided p-value contract this delegates to.
Usage
InferenceCountRobustPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect value.
InferenceCountRobustPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountRobustPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountRobustPoisson$new(seq_des)
inf$compute_estimate()
Zero-Inflated Negative Binomial Regression Inference for Count Responses
Description
Fits a zero-inflated negative binomial regression for count responses: a
binary excess-zero submodel P(\text{structural zero}_i) =
\mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) mixed with a (non-truncated)
negative-binomial count submodel \log E[Y_i \mid \text{not structural
zero}, w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma,
\mathrm{Var}(Y_i \mid \text{not structural zero}) = \mu_i + \mu_i^2 /
\theta. Unlike a hurdle model, zero counts can arise from either the
structural-zero mechanism or from an ordinary negative-binomial draw of
0. The hurdle and count submodels may use different covariate
formulas (model_formula/model_formula_zero). The reported
treatment effect is the coefficient from the conditional count component,
on the log-rate scale, conditional on the response coming from the
count process, not the excess-zero-inflation mechanism: it is not the
effect on the unconditional mean E[Y], which also depends on how
treatment shifts the excess-zero probability. A marginal
(unconditional-mean) estimand is not yet implemented for this class (see
marginal_estimand_report.md). likelihood_tier = "full":
Wald, gradient, score, and (bootstrap-calibrated) likelihood-ratio tests
are all available for the count submodel's treatment coefficient (unlike
the Poisson variant, this class's private
get_supported_testing_types_impl() includes "score").
Jackknife inference is not supported: delete-one refits of this
two-part mixture model with a jointly-estimated dispersion parameter are
numerically unstable, so compute_jackknife_estimate() and related
methods report explicit non-estimability rather than attempting
delete-one refits.
Super classes
Inference -> InferenceCountZeroAugmentedPoissonAbstract -> InferenceCountZeroInflatedNegBin
Methods
Public methods
+ inherited public methods from InferenceCountZeroAugmentedPoissonAbstract
InferenceCountZeroAugmentedPoissonAbstract$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_jackknife_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_randomization_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_subsampling_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$compute_asymp_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_asymp_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_estimate_with_bootstrap_weights()InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_bias_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_std_error()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_score_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_score_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_sensitivity()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_wald_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_wald_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$get_information_preference()InferenceCountZeroAugmentedPoissonAbstract$get_information_source_used()InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_diagnostics()InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_estimate_diagnostics()InferenceCountZeroAugmentedPoissonAbstract$get_mod()InferenceCountZeroAugmentedPoissonAbstract$get_summary()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_information_preferences()InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_testing_types()InferenceCountZeroAugmentedPoissonAbstract$get_testing_type()InferenceCountZeroAugmentedPoissonAbstract$select_optimal_b_subsampling()InferenceCountZeroAugmentedPoissonAbstract$select_optimal_m_out_of_n_bootstrap()InferenceCountZeroAugmentedPoissonAbstract$set_information_preference()InferenceCountZeroAugmentedPoissonAbstract$set_testing_type()InferenceCountZeroAugmentedPoissonAbstract$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountZeroInflatedNegBin$new()
Initialize inference for the zero-inflated negative
binomial model (binary excess-zero submodel mixed with a
negative-binomial count submodel); see
InferenceCountZeroInflatedNegBin
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Usage
InferenceCountZeroInflatedNegBin$new( des_obj, model_formula = NULL, model_formula_zero = NULL, use_rcpp = TRUE, verbose = FALSE, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for covariate adjustment.
model_formula_zeroFormula for the zero-inflation submodel. If
NULL(default), it uses the same formula asmodel_formula.use_rcppLogical. If
TRUE(default), use our internal Rcpp implementation. IfFALSE, use glmmTMB.verboseWhether to print progress messages.
optimization_algOptimization algorithm. Default is dispatched via policy.
InferenceCountZeroInflatedNegBin$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountZeroInflatedNegBin$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Lambert, D. (1992). "Zero-Inflated Poisson Regression, with an Application to Defects in Manufacturing." Technometrics, 34(1), 1-14, doi:10.2307/1269547, for the zero-inflated count-model framework.
See Also
InferenceCountNegBin for
the single-part negative binomial model this class's count submodel
generalizes;
InferenceCountZeroInflatedPoisson
for the Poisson (equidispersed) variant.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountZeroInflatedNegBin$new(seq_des, model_formula = ~ x1)
inf$compute_estimate()
Zero-Inflated Poisson Regression Inference for Count Responses
Description
Fits a zero-inflated Poisson regression for count responses: a binary
excess-zero submodel P(\text{structural zero}_i) =
\mathrm{logit}^{-1}(X_i^{h\top} \gamma^h) mixed with a (non-truncated)
Poisson count submodel \log E[Y_i \mid \text{not structural zero},
w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma. Unlike a hurdle
model, zero counts can arise from either the structural-zero mechanism or
from an ordinary Poisson draw of 0, so the two mixture components are
not identified by disjoint support. The hurdle and count submodels may use
different covariate formulas (model_formula/model_formula_zero).
The reported treatment effect is the coefficient from the conditional count
component, on the log-rate scale, conditional on the response
coming from the count process, not the excess-zero-inflation mechanism:
it is not the effect on the unconditional mean E[Y], which also
depends on how treatment shifts the excess-zero probability, under the
default estimand = "conditional". likelihood_tier = "full":
Wald, gradient, and (bootstrap-calibrated) likelihood-ratio tests are
available for the count submodel's treatment coefficient under that
estimand; a plain score test is not exposed. Jackknife inference
is not supported: delete-one refits of this two-part mixture model are
numerically unstable, so compute_jackknife_estimate() and related
methods report explicit non-estimability rather than attempting
delete-one refits.
Marginal (unconditional-mean) estimand. Via
set_estimand(), this class
also supports estimand = "marginal_mean_diff" and
"marginal_ratio": the g-computation average, over the empirical
covariate distribution, of the model-implied unconditional mean
E[Y_i \mid w_i, x_i] = (1 - \pi(x_i)) \lambda(x_i) (the untruncated
Poisson mean weighted by the non-structural-zero probability) at
w_i = 1 vs. w_i = 0 — a mean difference or, on the log scale,
a mean ratio. This is a pure post-fit transform of the same maximum-
likelihood fit (no refit), with a delta-method standard error computed
against the sandwich-robust covariance matrix already used for this
class's conditional Wald inference. Only "wald"-type inference is
available under a marginal estimand (no likelihood-ratio/score/gradient
test, since the marginal quantity is a functional of the fitted
parameters, not itself a likelihood).
Super classes
Inference -> InferenceCountZeroAugmentedPoissonAbstract -> InferenceCountZeroInflatedPoisson
Methods
Public methods
-
InferenceCountZeroInflatedPoisson$compute_asymp_confidence_interval() -
InferenceCountZeroInflatedPoisson$compute_asymp_two_sided_pval()
+ inherited public methods from InferenceCountZeroAugmentedPoissonAbstract
InferenceCountZeroAugmentedPoissonAbstract$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_jackknife_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_randomization_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$approximate_subsampling_distribution_beta_hat_T()InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_bayesian_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_estimate_with_bootstrap_weights()InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_gradient_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_bias_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_std_error()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_jackknife_wald_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bartlett_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_lik_ratio_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_estimate()InferenceCountZeroAugmentedPoissonAbstract$compute_param_bootstrap_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_bootstrap_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_rand_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_score_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_score_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_sensitivity()InferenceCountZeroAugmentedPoissonAbstract$compute_subsampling_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$compute_wald_confidence_interval()InferenceCountZeroAugmentedPoissonAbstract$compute_wald_two_sided_pval()InferenceCountZeroAugmentedPoissonAbstract$get_information_preference()InferenceCountZeroAugmentedPoissonAbstract$get_information_source_used()InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_diagnostics()InferenceCountZeroAugmentedPoissonAbstract$get_last_param_bootstrap_estimate_diagnostics()InferenceCountZeroAugmentedPoissonAbstract$get_mod()InferenceCountZeroAugmentedPoissonAbstract$get_summary()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bayesian_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_information_preferences()InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_ci_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_rand_bootstrap_pval_types()InferenceCountZeroAugmentedPoissonAbstract$get_supported_testing_types()InferenceCountZeroAugmentedPoissonAbstract$get_testing_type()InferenceCountZeroAugmentedPoissonAbstract$select_optimal_b_subsampling()InferenceCountZeroAugmentedPoissonAbstract$select_optimal_m_out_of_n_bootstrap()InferenceCountZeroAugmentedPoissonAbstract$set_information_preference()InferenceCountZeroAugmentedPoissonAbstract$set_testing_type()InferenceCountZeroAugmentedPoissonAbstract$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCountZeroInflatedPoisson$new()
Initialize inference for the zero-inflated Poisson model
(binary excess-zero submodel mixed with a Poisson count submodel); see
InferenceCountZeroInflatedPoisson
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Usage
InferenceCountZeroInflatedPoisson$new( des_obj, model_formula = NULL, model_formula_zero = NULL, use_rcpp = TRUE, verbose = FALSE, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a count response.model_formulaOptional formula for the count submodel.
model_formula_zeroFormula for the zero-inflation submodel. If
NULL(default), it uses the same formula asmodel_formula.use_rcppLogical. If
TRUE(default), use our internal Rcpp implementation. IfFALSE, use glmmTMB.verboseWhether to print progress messages.
optimization_algOptimization algorithm. Default is dispatched via policy.
InferenceCountZeroInflatedPoisson$compute_estimate()
Fits the zero-inflated Poisson model. Under the default
estimand = "conditional", returns \hat\beta_T, the
treatment log-rate coefficient from the conditional count submodel
(see the class-level caveat that this is conditional on the response
coming from the count process, not an unconditional-mean effect).
Under estimand = "marginal_mean_diff" or
"marginal_ratio" (set via set_estimand()), returns the
g-computation marginal mean difference or log-scale marginal ratio
of the unconditional mean E[Y \mid w, x] = (1-\pi(x))\lambda(x)
instead — a pure post-fit transform of the same cached fit, no refit.
Usage
InferenceCountZeroInflatedPoisson$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceCountZeroInflatedPoisson$compute_asymp_confidence_interval()
Asymptotic confidence interval. Under the conditional
estimand, delegates to the shared zero-augmented count-model Wald/
bootstrap-fallback contract; under a marginal estimand, the
delta-method interval computed by compute_estimate(). Calls
self$compute_estimate() first (not private$shared()
directly) so the estimand-aware cache is always current regardless
of call order.
Usage
InferenceCountZeroInflatedPoisson$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaThe significance level (default 0.05).
InferenceCountZeroInflatedPoisson$compute_asymp_two_sided_pval()
Asymptotic two-sided p-value, dispatched exactly as
compute_asymp_confidence_interval(); see that method's
description for the marginal-estimand path.
Usage
InferenceCountZeroInflatedPoisson$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect under the current estimand (default 0).
InferenceCountZeroInflatedPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCountZeroInflatedPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Lambert, D. (1992). "Zero-Inflated Poisson Regression, with an Application to Defects in Manufacturing." Technometrics, 34(1), 1-14, doi:10.2307/1269547, for the zero-inflated count-model framework.
See Also
InferenceCountPoisson for
the single-part Poisson model this class's count submodel generalizes;
InferenceCountHurdlePoisson
for the related hurdle (disjoint-support) variant;
InferenceCountZeroInflatedNegBin
for the overdispersion-robust negative-binomial variant (does not
support a marginal estimand — the mean-function derivation here is
Poisson-specific).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'count')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rpois(10, 2))
inf = InferenceCountZeroInflatedPoisson$new(seq_des)
inf$compute_estimate()
Internal base for user-defined asymptotic inference extensions
Description
InferenceCustomAsymp is intentionally not exported. Extension packages
may retrieve it with getFromNamespace("InferenceCustomAsymp", "EDI")
while this API is experimental.
Subclasses implement a public fit(estimate_only = FALSE) method and
return a named list with the custom-fit result contract:
estimateRequired numeric scalar treatment-effect estimate.
seOptional numeric scalar standard error. Required for Wald confidence intervals and asymptotic p-values unless
estimate_onlyisTRUE.dfOptional numeric scalar degrees of freedom. Use
NA_real_for z inference.modelOptional fitted model object retained for
get_mod()andget_summary().nonestimable_reasonOptional character scalar. When supplied with a non-finite estimate or standard error, EDI records the result as explicitly non-estimable.
Subclasses should use public accessors such as get_analysis_data(),
get_response(), get_treatment(), and get_covariates()
rather than EDI private fields.
Value
A two-sided p-value.
Super class
Inference -> InferenceCustomAsymp
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCustomAsymp$compute_rand_two_sided_pval()
Computes a randomization two-sided p-value. Delegates to the 'RandomizationCI'-provided dispatch (Zhang incidence support, type/args_for_type) since 'NonparametricBootstrap' pulls in both 'RandomizationTest' and 'RandomizationCI', and the two provide conflicting 'compute_rand_two_sided_pval' implementations.
Usage
InferenceCustomAsymp$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, type = NULL, args_for_type = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors.
deltaNull treatment effect value.
transform_responsesResponse transformation to apply during the test. For survival responses the default
"log"multiplies the recorded times of the units treated under each reference allocation bye^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); seecompute_rand_confidence_interval()for the assumptions.na.rmWhether to remove non-finite simulated statistics.
show_progressWhether to show progress.
permutationsOptional pre-generated assignment draws.
typeOptional incidence-specific exact randomization type.
args_for_typeOptional arguments keyed by
type.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCustomAsymp$fit()
Calls the user-defined fit callback for this custom inference path; see
InferenceCustomAsymp.
Usage
InferenceCustomAsymp$fit(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
Returns
A list with fit results.
InferenceCustomAsymp$compute_estimate()
Compute the treatment-effect estimate by delegating to the
user-supplied custom estimator. See
InferenceCustomRand,
InferenceCustomAsymp, and
InferenceCustomBoot for related
extension classes.
Usage
InferenceCustomAsymp$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
Returns
The treatment estimate.
InferenceCustomAsymp$compute_asymp_confidence_interval()
Compute asymptotic confidence interval.
Usage
InferenceCustomAsymp$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level.
Returns
Confidence interval.
InferenceCustomAsymp$compute_asymp_two_sided_pval()
Compute asymptotic p-value.
Usage
InferenceCustomAsymp$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect.
Returns
P-value.
InferenceCustomAsymp$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCustomAsymp$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Internal base for user-defined bootstrap inference extensions
Description
This class uses the same fit() result contract as
InferenceCustomAsymp, but only promises estimate/bootstrap behavior.
Value
A two-sided p-value.
Super class
Inference -> InferenceCustomBoot
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCustomBoot$compute_rand_two_sided_pval()
Computes a randomization two-sided p-value. Delegates to the 'RandomizationCI'-provided dispatch (Zhang incidence support, type/args_for_type) since 'NonparametricBootstrap' pulls in both 'RandomizationTest' and 'RandomizationCI', and the two provide conflicting 'compute_rand_two_sided_pval' implementations.
Usage
InferenceCustomBoot$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, type = NULL, args_for_type = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors.
deltaNull treatment effect value.
transform_responsesResponse transformation to apply during the test. For survival responses the default
"log"multiplies the recorded times of the units treated under each reference allocation bye^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); seecompute_rand_confidence_interval()for the assumptions.na.rmWhether to remove non-finite simulated statistics.
show_progressWhether to show progress.
permutationsOptional pre-generated assignment draws.
typeOptional incidence-specific exact randomization type.
args_for_typeOptional arguments keyed by
type.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCustomBoot$fit()
Calls the user-defined fit callback for this custom inference path; see
InferenceCustomAsymp.
Usage
InferenceCustomBoot$fit(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
Returns
A list with fit results.
InferenceCustomBoot$compute_estimate()
Compute the treatment-effect estimate by delegating to the
user-supplied custom asymptotic estimator. See
InferenceCustomAsymp and
InferenceAsymp.
Usage
InferenceCustomBoot$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
Returns
The treatment estimate.
InferenceCustomBoot$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCustomBoot$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Internal base for user-defined randomization inference extensions
Description
This class uses the same fit() result contract as
InferenceCustomAsymp, but only promises estimate/randomization/
randomization-CI behavior.
Value
A two-sided p-value.
Super class
Inference -> InferenceCustomRand
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceCustomRand$compute_rand_two_sided_pval()
Computes a randomization two-sided p-value. Delegates to the 'RandomizationCI'-provided dispatch (Zhang incidence support, type/args_for_type) since 'RandomizationCI' pulls in 'RandomizationTest', and the two provide conflicting 'compute_rand_two_sided_pval' implementations – the same conflict 'InferenceCustomAsymp' and 'InferenceCustomBoot' resolve the same way.
Usage
InferenceCustomRand$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, type = NULL, args_for_type = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors.
deltaNull treatment effect value.
transform_responsesResponse transformation to apply during the test. For survival responses the default
"log"multiplies the recorded times of the units treated under each reference allocation bye^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); seecompute_rand_confidence_interval()for the assumptions.na.rmWhether to remove non-finite simulated statistics.
show_progressWhether to show progress.
permutationsOptional pre-generated assignment draws.
typeOptional incidence-specific exact randomization type.
args_for_typeOptional arguments keyed by
type.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceCustomRand$fit()
Calls the user-defined fit callback for this custom inference path; see
InferenceCustomAsymp.
Usage
InferenceCustomRand$fit(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
Returns
A list with fit results.
InferenceCustomRand$compute_estimate()
Compute the treatment-effect estimate by delegating to the
user-supplied custom bootstrap estimator. See
InferenceCustomBoot and
InferenceNonParamBootstrap.
Usage
InferenceCustomRand$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance calculations.
Returns
The treatment estimate.
InferenceCustomRand$clone()
The objects of this class are cloneable with this method.
Usage
InferenceCustomRand$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Binomial Identity Risk Difference Inference for Incidence Responses
Description
Fits a binomial regression with the identity link for binary
(incidence) responses: P(Y_i = 1) = \beta_0 + \beta_T W_i + X_i^\top
\gamma, where W_i is the treatment indicator and X_i are
optional recorded covariates, by maximum likelihood
(fast_identity_binomial_regression_cpp/
fast_identity_binomial_regression_weighted_cpp). Because the
link is the identity rather than the logit, \hat\beta_T is directly a
risk difference on the probability scale, not a log-odds-ratio —
the class name and estimand differ from
InferenceIncidLogRegr for exactly
this reason. likelihood_tier = "full": likelihood-ratio, score,
gradient, and Wald tests are all available when the model converges, plus
parametric-likelihood-bootstrap calibration of the likelihood-ratio test.
Because the identity link does not constrain fitted probabilities to
[0,1], fits are hardened by QR column-dropping and rejected as
nonestimable when the fitted linear predictor produces implausible
coefficients (see private$is_identity_binomial_fit_reasonable());
this is a real practical limitation of the identity link relative to
logit/probit, not a bug. Validity requires the additive risk-difference
model to be correctly specified over the covariate range actually observed
(an identity-link fit can be well-behaved in-sample yet imply
out-of-range probabilities for other covariate values).
Estimand. Composes
MarginalEstimand
(set_estimand()/get_estimand()/get_supported_estimands()).
estimand = "marginal_mean_diff" is supported for API consistency
with the other GLM families, but it is algebraically identical to
the default estimand = "conditional" for this class: the
g-computation marginal risk difference is \frac{1}{n}\sum_i
\{(\hat\beta_0 + \hat\beta_T + X_i^\top \hat\gamma) - (\hat\beta_0 +
X_i^\top \hat\gamma)\}, which simplifies to exactly \hat\beta_T for
every subject (not merely on average) because the identity link is linear
in W_i with no treatment-by-covariate interaction term — the
per-subject treated-minus-control difference \hat\beta_T does not
depend on X_i at all, so standardizing over the covariate
distribution changes nothing. Contrast with
InferenceCountPoisson's
"marginal_ratio" (also a collapsing case, for the same
no-interaction reason) and "marginal_mean_diff" (which does not
collapse, since a difference does not distribute through the nonlinear
log-link mean). This collapsing is a genuine property of the identity-link
model, not a wiring bug — it is documented here so a user comparing
estimands for this class is not surprised the two never differ.
Standard errors are computed independently for each estimand (the
marginal path uses the delta method against the model's coefficient
covariance; the conditional path uses the model information matrix), so
while the two point estimates coincide exactly, their standard errors may
differ slightly by construction even though both are asymptotically valid.
Super class
Inference -> InferenceIncidBinomialIdentityRiskDiff
Methods
Public methods
-
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_confidence_interval() -
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidBinomialIdentityRiskDiff$compute_estimate()
Fits the identity-link binomial regression model by
maximum likelihood. Under the default estimand =
"conditional", returns \hat\beta_T, the risk difference
coefficient. Under estimand = "marginal_mean_diff" (set via
set_estimand()), returns the g-computation marginal risk
difference — see the class-level @details for why this is
algebraically identical to the conditional estimate for this
family. The underlying model fit is identical either way (a pure
post-fit transform of the same cached fit, no refit).
Usage
InferenceIncidBinomialIdentityRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_confidence_interval()
Wald confidence interval, dispatched by
testing_type for the conditional estimand; under a
marginal estimand testing_type is always "wald" (the
only value set_estimand() permits there). Calls
self$compute_estimate() first (not private$shared()
directly) so the estimand-aware cache is always current
regardless of call order.
Usage
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_two_sided_pval()
Wald two-sided p-value, dispatched by
testing_type exactly as
compute_asymp_confidence_interval(); see that method's
description for the marginal-estimand always-Wald note.
Usage
InferenceIncidBinomialIdentityRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment-effect value under the current estimand (both scales coincide for this family — see the class-level
@details).
InferenceIncidBinomialIdentityRiskDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidBinomialIdentityRiskDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the binomial GLM family and identity-link risk-difference parameterization.
See Also
InferenceIncidLogRegr
(logit link, log-odds-ratio estimand),
InferenceIncidLogBinomial
(log link, log-risk-ratio estimand) for alternative link/estimand
choices on the same response type. Comparable Python API:
statsmodels GLM
(family=Binomial(link=identity())). See also:
Generalized
linear model (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidBinomialIdentityRiskDiff$new(seq_des)
inf$compute_estimate()
CMH Blocked Incidence Inference
Description
Unadjusted blocked-design incidence inference using the simple mean-difference point estimate with a randomization-based standard error.
Legacy inference class. This class is retained for backwards compatibility and is not comprehensively tested by the package comprehensive-test harness.
Internally, this class recodes treatment assignments to w_i \in \{-1, +1\}
(the package-wide convention is \{0,1\}; see Design). For a balanced
design the treatment-effect estimator is \hat\tau = (2/n)\,\mathbf{y}'\mathbf{w},
and since E_w[\mathbf{y}'\mathbf{w}] = 0 for any balanced randomization the
standard error is
SE(\hat\tau) = \frac{2}{n}\sqrt{\frac{\sum_k (\mathbf{y}'\mathbf{w}_k)^2}{K}}
where K draws \mathbf{w}_1,\ldots,\mathbf{w}_K come from the design's
reference distribution. Centering at the known zero mean (rather than the sample mean)
makes the denominator K rather than K-1.
For blocking designs the expectation is evaluated exactly:
SE(\hat\tau) = \frac{2}{n}\sqrt{\sum_b \frac{n_{1b}\,n_{0b}}{n_B - 1}}
where n_{1b}, n_{0b} are the numbers of positive and negative responses in
block b and n_B is the (common) block size. This equals
2\sqrt{V_{\rm CMH}} where V_{\rm CMH} is the CMH variance from
Azriel et al. (2026), Equation 3.
For non-blocking designs, the "balanced design" precondition above requires
the observed treatment allocation to be exactly balanced
(n_T = n_C), not merely drawn from a prob\_T = 0.5 mechanism –
e.g. plain Bernoulli randomization has prob\_T = 0.5 but does not
guarantee an exactly balanced realized allocation. A warning (not an error)
is issued once, the first time the standard error is actually computed
(i.e. on the first confidence-interval / p-value / standard-error request,
not at construction or for estimate-only use), when this is violated –
erroring would make this class unusable with Bernoulli-style non-blocking
designs entirely; the warning tells the caller the reported standard error
may be miscalibrated.
Super class
Inference -> InferenceIncidCMH
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidCMH$compute_asymp_confidence_interval()
Uses the randomization-CI layer's two-sided p-value contract
(InferenceRandCI's version, not InferenceRand's): for
incidence responses this dispatches to the Zhang exact randomization
test where applicable rather than refusing outright, matching this
class's pre-migration old-ladder behavior (it inherited from
InferenceAllSimpleAverageDiff, whose own pin was already
corrected to InferenceRandCI – see that file's identical
rationale). This class independently composes the same components
rather than truly inheriting InferenceAllSimpleAverageDiff, so it
had its own stale copy of the old InferenceRand pin, which
silently regressed Zhang dispatch for the non-blocking balanced-design
path – found via
test-incid-cmh-extended-robins-migration-golden.R's
randomization_pval case going from '"ok"' to '"unsupported"'.
Wald confidence interval for the balanced-design/CMH risk-difference
estimate \hat\tau, using the randomization-based (blocking-design: exact CMH
variance formula; non-blocking design: Monte Carlo over se_est_num_vectors
design draws) standard error documented in the class @details. See
InferenceAsymp for the shared Wald contract.
Usage
InferenceIncidCMH$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
Returns
A length-2 numeric vector c(lower, upper) on the risk-difference scale.
InferenceIncidCMH$compute_asymp_two_sided_pval()
Two-sided Wald p-value for H_0: \tau = \code{delta} vs.
H_1: \tau \neq \code{delta}, using the same randomization-based standard error
as compute_asymp_confidence_interval().
Usage
InferenceIncidCMH$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null value of
\tauto test against; 0 (the default) tests for any treatment effect at all.
Returns
Numeric scalar p-value in [0, 1].
InferenceIncidCMH$new()
Initialize Cochran-Mantel-Haenszel incidence inference,
validate the stratified binary-response design, and prepare the
stratum-adjusted test used by
InferenceIncidCMH.
Usage
InferenceIncidCMH$new( des_obj, model_formula = NULL, se_est_num_vectors = 5000L, verbose = FALSE )
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment.
se_est_num_vectorsFor non-block designs, the number of randomization vectors drawn from the design to estimate the standard error. Default
1000L.verboseLogical. Whether to print progress messages.
Returns
A new InferenceIncidCMH object.
InferenceIncidCMH$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidCMH$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneRandomBlockSize$new(n = 20, response_type = 'incidence',
strata_cols = 'x1')
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(
data.frame(x1 = factor(rep(1:2, 10)[i], levels=1:2)))
}
seq_des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidCMH$new(seq_des)
inf$compute_estimate()
Exact Binomial (McNemar-Type) Incidence Inference for Matched-Pair Designs
Description
Performs exact matched-pair inference for binary (incidence) outcomes using
only discordant matched pairs — pairs where the treated and control
member's outcomes differ — the same reduction classical McNemar's test makes.
Writing d_+ for the count of discordant pairs where the treated
subject had the event and the control did not, and d_- for the reverse,
the point estimate is the Haldane-Anscombe continuity-corrected log odds
ratio \log\left((d_+ + 0.5)/(d_- + 0.5)\right); the confidence interval
inverts the exact (Clopper-Pearson) binomial confidence interval for
d_+ / (d_+ + d_-) against 1/2 (via stats::binom.test) onto
the log-odds scale; and the two-sided p-value is an exact binomial test of
d_+ vs. d_- (via zhang_exact_binom_pval_cpp) against a
null log odds ratio. This class is available for
DesignFixedBinaryMatch and KK matching-on-the-fly designs. For KK
designs, only the matched-pair data are used and the reservoir is ignored.
If there are no matched pairs, or no discordant pairs, the relevant
quantities are reported as non-estimable rather than as NaN/Inf.
Initialize exact matched-pair binomial inference for
incidence outcomes. Requires des_obj to be
DesignFixedBinaryMatch or a KK matching-on-the-fly-capable
design; errors otherwise. Requires an uncensored incidence response.
Computes the Haldane-Anscombe continuity-corrected
matched-pair log odds ratio \log\left((d_+ + 0.5)/(d_- + 0.5)\right)
from the discordant matched-pair counts (see class documentation for
the full model). NA if there are no matched pairs.
Value
A new InferenceIncidExactBinomial object.
The treatment estimate.
Super class
Inference -> InferenceIncidExactBinomial
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidExactBinomial$new()
Usage
InferenceIncidExactBinomial$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceIncidExactBinomial$compute_estimate()
Usage
InferenceIncidExactBinomial$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIgnored for this estimator (the exact statistic is always cheap to compute; there is no separate variance step to skip).
InferenceIncidExactBinomial$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidExactBinomial$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidExactBinomial$new(seq_des)
inf$compute_estimate()
Exact Fisher (Conditional Hypergeometric) Incidence Inference
Description
Performs exact conditional inference for binary (incidence) outcomes via
Fisher's exact test on one or more 2x2 (treated/control by case/noncase)
tables. When the design provides no stratification structure (e.g. an
unstructured or iBCRD design), a single overall 2x2 table is built and
fisher.test is used directly, giving the conditional
MLE odds ratio and its exact confidence interval/p-value. When the design
has blocking structure (DesignFixedBlocking,
DesignSeqOneByOneSPBR, DesignSeqOneByOneRandomBlockSize), a
separate 2x2 table is built per block-defining covariate stratum. When the
design has matched-pair structure (KK matching-on-the-fly designs),
each matched pair becomes its own 2x2 table, with any reservoir (unmatched)
subjects pooled into one additional stratum table. In either stratified
case, mantelhaen.test (exact conditional test) is used
instead, giving the common odds ratio across strata; stratified inference
only supports testing/estimating against a null odds ratio of 1 (log odds
ratio 0) — a non-zero null shift is rejected with an error. Strata with no
cases or no noncases in either arm are dropped before analysis; if no
informative strata remain, this errors rather than returning a degenerate
result.
Initialize exact Fisher inference for incidence outcomes. Requires an uncensored incidence response; the design's structure (unstructured, blocked, or matched) determines the stratification used at estimation time (see class documentation).
Computes the log of the (conditional MLE, or common-odds-ratio
if stratified) odds ratio from fisher.test or
mantelhaen.test (see class documentation for which
applies and why).
Value
A new InferenceIncidExactFisher object.
The treatment estimate.
Super class
Inference -> InferenceIncidExactFisher
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidExactFisher$new()
Usage
InferenceIncidExactFisher$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceIncidExactFisher$compute_estimate()
Usage
InferenceIncidExactFisher$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIgnored for this estimator (the exact statistic is always cheap to compute; there is no separate variance step to skip).
InferenceIncidExactFisher$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidExactFisher$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Fisher, R. A. (1935). "The Logic of Inductive Inference."
Journal of the Royal Statistical Society, 98(1), 39-82,
doi:10.2307/2342435, for the exact conditional test underlying
fisher.test; Mantel, N., and Haenszel, W. (1959).
"Statistical Aspects of the Analysis of Data from Retrospective Studies
of Disease." Journal of the National Cancer Institute, 22(4),
719-748, for the stratified common-odds-ratio test used when the design
provides multiple strata.
Examples
des = DesignFixediBCRD$new(n = 20, response_type = 'incidence')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(20)))
des$assign_w_to_all_subjects()
des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidExactFisher$new(des)
inf$compute_estimate()
inf$compute_exact_two_sided_pval_for_treatment_effect()
Exact Zhang Combined-Test Incidence Inference
Description
Performs exact inference for a binary (incidence) outcome that
combines two exact component tests when the design has both
matched-pair and reservoir (unmatched) subjects — an internal-to-this-package
method (not drawn from external literature) analogous in spirit to
InferenceIncidExactBinomial
(matched pairs) and InferenceIncidExactFisher
(unmatched 2x2 table), fused into one combined exact test rather than a
Wald-style variance combination. The point estimate is always the
Haldane-Anscombe continuity-corrected log odds ratio \log\left((n_{11} +
0.5)(n_{00} + 0.5) / \left((n_{10}+0.5)(n_{01}+0.5)\right)\right) from the
pooled 2\times2 table across all subjects (matched and reservoir
together). For p-values and confidence intervals, the two
subsets are tested separately (an exact matched-pairs binomial test, as in
InferenceIncidExactBinomial, on discordant pairs; an exact Fisher test
on the reservoir 2\times2 table, as in InferenceIncidExactFisher),
and their p-values are combined via combination_method: "Fisher"
(default; -2(\log p_M + \log p_R) \sim \chi^2_4 under independence),
"Stouffer" (averaged z-scores), or "min_p" (Šidák-style
1-(1-\min(p_M,p_R))^2). If only one of the two subsets is informative
(e.g. a pure-Bernoulli design with no matching, or no discordant pairs), the
combined p-value degenerates to that one component's p-value. Confidence
intervals are obtained by numerically inverting (bisection) the combined
p-value as a function of the hypothesized log odds ratio, starting from a
normal-approximation (Haldane-Anscombe MLE) interval as the search bracket.
Requires a Bernoulli-capable or matching-capable design.
Initialize exact Zhang combined-test incidence inference.
Requires des_obj to be Bernoulli-capable or
matching-capable, an uncensored incidence response.
Computes the Haldane-Anscombe continuity-corrected log odds
ratio \log\left((n_{11}+0.5)(n_{00}+0.5) /
\left((n_{10}+0.5)(n_{01}+0.5)\right)\right) from the pooled
2\times2 table across all subjects (matched and reservoir
combined) — see class documentation for the full combined-test model.
Computes an exact confidence interval for the log odds ratio by bisection-inverting the combined matched-pairs + Fisher-exact p-value (see class documentation for the full combination methodology).
Value
A new InferenceIncidExactZhang object.
The treatment estimate.
A confidence interval.
Super class
Inference -> InferenceIncidExactZhang
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidExactZhang$new()
Usage
InferenceIncidExactZhang$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceIncidExactZhang$compute_estimate()
Usage
InferenceIncidExactZhang$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIgnored for this estimator (the exact statistic is always cheap to compute; there is no separate variance step to skip).
InferenceIncidExactZhang$compute_exact_confidence_interval()
Usage
InferenceIncidExactZhang$compute_exact_confidence_interval( alpha = 0.05, pval_epsilon = 0.005, type = NULL, args_for_type = NULL )
Arguments
alphaSignificance level.
pval_epsilonBisection tolerance for the inversion routine.
typeExact inference type; only
"Zhang"(the default) is supported.args_for_typeOptional arguments keyed by exact type; recognizes
combination_method("Fisher"(default),"Stouffer", or"min_p") inside the"Zhang"entry.
InferenceIncidExactZhang$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidExactZhang$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = 'incidence')
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidExactZhang$new(seq_des)
inf$compute_estimate()
inf$compute_exact_two_sided_pval_for_treatment_effect()
Extended Robins Blocked Incidence Inference
Description
Unadjusted blocked-design incidence inference using the simple mean-difference point estimate with a block-stratified standard error.
Legacy inference class. This class is retained for backwards compatibility and is not comprehensively tested by the package comprehensive-test harness.
Super class
Inference -> InferenceIncidExtendedRobins
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidExtendedRobins$compute_asymp_confidence_interval()
Uses the randomization-CI layer's two-sided p-value contract
(InferenceRandCI's version, not InferenceRand's): for
incidence responses this dispatches to the Zhang exact randomization
test where applicable rather than refusing outright, matching this
class's pre-migration old-ladder behavior (it inherited from
InferenceAllSimpleAverageDiff, whose own pin was already
corrected to InferenceRandCI – see that file's identical
rationale). This class independently composes the same components
rather than truly inheriting InferenceAllSimpleAverageDiff, so it
had its own stale copy of the old InferenceRand pin (same bug
as InferenceIncidWald/InferenceIncidCMH, fixed
alongside them even though this class's own golden test's design
doesn't happen to trigger the Zhang-eligible path that would have
caught it).
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Usage
InferenceIncidExtendedRobins$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaNumeric. Significance level (default 0.05).
InferenceIncidExtendedRobins$compute_asymp_two_sided_pval()
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
InferenceIncidExtendedRobins$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNumeric. Null treatment effect value (default 0).
InferenceIncidExtendedRobins$new()
Initialize Extended Robins blocked-design incidence inference.
Usage
InferenceIncidExtendedRobins$new( des_obj, model_formula = NULL, verbose = FALSE )
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment.
verboseLogical. Whether to print progress messages.
Returns
A new InferenceIncidExtendedRobins object.
InferenceIncidExtendedRobins$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidExtendedRobins$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneRandomBlockSize$new(n = 20, response_type = 'incidence',
strata_cols = 'x1')
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(
data.frame(x1 = factor(rep(1:2, 10)[i], levels=1:2)))
}
seq_des$add_all_subject_responses(rbinom(20, 1, 0.5))
inf = InferenceIncidExtendedRobins$new(seq_des)
inf$compute_estimate()
G-Computation Risk-Difference Inference for Binary Responses
Description
Fits a logistic working model, \mathrm{logit}\,\Pr(Y_i=1\mid x_i) =
x_i^\top\hat\beta, for an incidence outcome using treatment and, optionally,
all recorded covariates, then estimates the marginal (standardized) risk
difference \mathrm{RD} = \overline{\mathrm{risk}}_1 -
\overline{\mathrm{risk}}_0 by G-computation: setting every subject's
treatment indicator to 1 (respectively 0) while holding their other observed
covariates fixed, averaging the model-implied risk over the empirical
covariate distribution under each counterfactual, and differencing — see
gcomp_logistic_point_estimate_cpp for the exact standardization
formula. Inference is nonparametric-bootstrap/randomization/jackknife-based
(likelihood_tier = "none"): no closed-form asymptotic standard error
is used.
Uses the shared nonparametric bootstrap distribution contract; see
InferenceNonParamBootstrap.
Computes a bootstrap confidence interval for the treatment effect.
Computes a bootstrap two-sided p-value for the treatment effect.
Computes a Bayesian-bootstrap two-sided p-value for the treatment effect.
Computes a Bayesian-bootstrap confidence interval for the treatment effect.
Computes a jackknife-Wald two-sided p-value for the treatment effect.
Computes a jackknife-Wald confidence interval for the treatment effect.
Computes a PRW subsampling two-sided p-value for the treatment effect.
Computes a PRW subsampling confidence interval for the treatment effect.
Computes an m-out-of-n bootstrap two-sided p-value for the treatment effect.
Computes an m-out-of-n bootstrap confidence interval for the treatment effect.
Value
A numeric vector of bootstrap estimates.
Super class
Inference -> InferenceIncidGCompRiskDiff
Methods
Public methods
-
InferenceIncidGCompRiskDiff$approximate_bootstrap_distribution_beta_hat_T() -
InferenceIncidGCompRiskDiff$compute_bootstrap_confidence_interval() -
InferenceIncidGCompRiskDiff$compute_bootstrap_two_sided_pval() -
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_two_sided_pval() -
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_confidence_interval() -
InferenceIncidGCompRiskDiff$compute_jackknife_wald_two_sided_pval() -
InferenceIncidGCompRiskDiff$compute_jackknife_wald_confidence_interval() -
InferenceIncidGCompRiskDiff$compute_subsampling_two_sided_pval() -
InferenceIncidGCompRiskDiff$compute_subsampling_confidence_interval() -
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_two_sided_pval() -
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidGCompRiskDiff$approximate_bootstrap_distribution_beta_hat_T()
Usage
InferenceIncidGCompRiskDiff$approximate_bootstrap_distribution_beta_hat_T( B = 501, show_progress = TRUE, debug = FALSE, bootstrap_type = NULL )
Arguments
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
debugWhether to return diagnostics.
bootstrap_typeOptional resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskDiff$compute_bootstrap_confidence_interval()
Usage
InferenceIncidGCompRiskDiff$compute_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
InferenceIncidGCompRiskDiff$compute_bootstrap_two_sided_pval()
Usage
InferenceIncidGCompRiskDiff$compute_bootstrap_two_sided_pval( delta = NULL, B = 501, type = "symmetric", na.rm = FALSE, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_two_sided_pval()
Usage
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_two_sided_pval( delta = NULL, B = 501, type = NULL, na.rm = FALSE, show_progress = TRUE, min_number_usable_samples = 5L, weighting_unit_type = NULL )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
weighting_unit_typeOptional resampling unit override.
weighting_unit_typeOptional resampling unit override.
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_confidence_interval()
Usage
InferenceIncidGCompRiskDiff$compute_bayesian_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, min_number_usable_samples = 5L, weighting_unit_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
weighting_unit_typeOptional resampling unit override.
weighting_unit_typeOptional resampling unit override.
InferenceIncidGCompRiskDiff$compute_jackknife_wald_two_sided_pval()
Usage
InferenceIncidGCompRiskDiff$compute_jackknife_wald_two_sided_pval( delta = NULL, unit = "auto" )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceIncidGCompRiskDiff$compute_jackknife_wald_confidence_interval()
Usage
InferenceIncidGCompRiskDiff$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceIncidGCompRiskDiff$compute_subsampling_two_sided_pval()
Usage
InferenceIncidGCompRiskDiff$compute_subsampling_two_sided_pval( delta = NULL, B = 501, b = NULL, type = "centered", show_progress = TRUE, min_number_usable_samples = 5L, subsampling_type = NULL )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
subsampling_typeOptional empirical-resampling scheme.
subsampling_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskDiff$compute_subsampling_confidence_interval()
Usage
InferenceIncidGCompRiskDiff$compute_subsampling_confidence_interval( alpha = 0.05, B = 501, b = NULL, type = "basic", show_progress = TRUE, min_number_usable_samples = 5L, subsampling_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
subsampling_typeOptional empirical-resampling scheme.
subsampling_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_two_sided_pval()
Usage
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_two_sided_pval( delta = NULL, B = 501, m = NULL, type = "centered", show_progress = TRUE, min_number_usable_samples = 5L, bootstrap_type = NULL )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
bootstrap_typeOptional resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_confidence_interval()
Usage
InferenceIncidGCompRiskDiff$compute_m_out_of_n_bootstrap_confidence_interval( alpha = 0.05, B = 501, m = NULL, type = "basic", show_progress = TRUE, min_number_usable_samples = 5L, bootstrap_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
bootstrap_typeOptional resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidGCompRiskDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
See Also
InferenceIncidGCompRiskRatio
for the risk-ratio version of this same standardized logistic working model.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidGCompRiskDiff$new(seq_des)
inf$compute_estimate()
G-Computation Risk-Ratio Inference for Binary Responses
Description
Fits a logistic working model, \mathrm{logit}\,\Pr(Y_i=1\mid x_i) =
x_i^\top\hat\beta, for an incidence outcome using treatment and, optionally,
all recorded covariates, then estimates the marginal (standardized) risk
ratio \mathrm{RR} = \overline{\mathrm{risk}}_1 / \overline{\mathrm{risk}}_0
by G-computation: setting every subject's treatment indicator to 1
(respectively 0) while holding their other observed covariates fixed,
averaging the model-implied risk over the empirical covariate distribution
under each counterfactual, and taking the ratio — see
gcomp_logistic_point_estimate_cpp for the exact standardization
formula (mean1/mean0). Bootstrap/jackknife inference on this
estimand is generally done on the log risk-ratio scale internally
(see $compute_bootstrap_confidence_interval(),
$compute_bayesian_bootstrap_confidence_interval(), and the
jackknife-Wald methods, whose "basic"/"wald" interval types
route through log-scale-specific helpers for this estimand), then
back-transformed, since ratio estimators are typically closer to normally
distributed on the log scale. Inference is nonparametric-bootstrap/
randomization/jackknife-based (likelihood_tier = "none"): no
closed-form asymptotic standard error is used.
Uses the shared nonparametric bootstrap distribution contract; see
InferenceNonParamBootstrap.
Computes a bootstrap confidence interval for the treatment effect.
Computes a bootstrap two-sided p-value for the treatment effect.
Computes a Bayesian-bootstrap two-sided p-value for the treatment effect.
Computes a Bayesian-bootstrap confidence interval for the treatment effect.
Computes a jackknife-Wald two-sided p-value for the treatment effect.
Computes a jackknife-Wald confidence interval for the treatment effect.
Computes a PRW subsampling two-sided p-value for the treatment effect.
Computes a PRW subsampling confidence interval for the treatment effect.
Computes an m-out-of-n bootstrap two-sided p-value for the treatment effect.
Computes an m-out-of-n bootstrap confidence interval for the treatment effect.
Value
A numeric vector of bootstrap estimates.
Super class
Inference -> InferenceIncidGCompRiskRatio
Methods
Public methods
-
InferenceIncidGCompRiskRatio$approximate_bootstrap_distribution_beta_hat_T() -
InferenceIncidGCompRiskRatio$compute_bootstrap_confidence_interval() -
InferenceIncidGCompRiskRatio$compute_bootstrap_two_sided_pval() -
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_two_sided_pval() -
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_confidence_interval() -
InferenceIncidGCompRiskRatio$compute_jackknife_wald_two_sided_pval() -
InferenceIncidGCompRiskRatio$compute_jackknife_wald_confidence_interval() -
InferenceIncidGCompRiskRatio$compute_subsampling_two_sided_pval() -
InferenceIncidGCompRiskRatio$compute_subsampling_confidence_interval() -
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_two_sided_pval() -
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidGCompRiskRatio$approximate_bootstrap_distribution_beta_hat_T()
Usage
InferenceIncidGCompRiskRatio$approximate_bootstrap_distribution_beta_hat_T( B = 501, show_progress = TRUE, debug = FALSE, bootstrap_type = NULL )
Arguments
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
debugWhether to return diagnostics.
bootstrap_typeOptional resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskRatio$compute_bootstrap_confidence_interval()
Usage
InferenceIncidGCompRiskRatio$compute_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
InferenceIncidGCompRiskRatio$compute_bootstrap_two_sided_pval()
Usage
InferenceIncidGCompRiskRatio$compute_bootstrap_two_sided_pval( delta = NULL, B = 501, type = "symmetric", na.rm = FALSE, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_two_sided_pval()
Usage
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_two_sided_pval( delta = NULL, B = 501, type = NULL, na.rm = FALSE, show_progress = TRUE, min_number_usable_samples = 5L, weighting_unit_type = NULL )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
weighting_unit_typeOptional resampling unit override.
weighting_unit_typeOptional resampling unit override.
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_confidence_interval()
Usage
InferenceIncidGCompRiskRatio$compute_bayesian_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, min_number_usable_samples = 5L, weighting_unit_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
weighting_unit_typeOptional resampling unit override.
weighting_unit_typeOptional resampling unit override.
InferenceIncidGCompRiskRatio$compute_jackknife_wald_two_sided_pval()
Usage
InferenceIncidGCompRiskRatio$compute_jackknife_wald_two_sided_pval( delta = NULL, unit = "auto" )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceIncidGCompRiskRatio$compute_jackknife_wald_confidence_interval()
Usage
InferenceIncidGCompRiskRatio$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
unitDeletion unit. Default
"auto".unitDeletion unit. Default
"auto".
InferenceIncidGCompRiskRatio$compute_subsampling_two_sided_pval()
Usage
InferenceIncidGCompRiskRatio$compute_subsampling_two_sided_pval( delta = NULL, B = 501, b = NULL, type = "centered", show_progress = TRUE, min_number_usable_samples = 5L, subsampling_type = NULL )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
subsampling_typeOptional empirical-resampling scheme.
subsampling_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskRatio$compute_subsampling_confidence_interval()
Usage
InferenceIncidGCompRiskRatio$compute_subsampling_confidence_interval( alpha = 0.05, B = 501, b = NULL, type = "basic", show_progress = TRUE, min_number_usable_samples = 5L, subsampling_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval.bSubsample size. See
InferenceNonParamBootstrap$compute_subsampling_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
subsampling_typeOptional empirical-resampling scheme.
subsampling_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_two_sided_pval()
Usage
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_two_sided_pval( delta = NULL, B = 501, m = NULL, type = "centered", show_progress = TRUE, min_number_usable_samples = 5L, bootstrap_type = NULL )
Arguments
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaThe null treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
deltaNull treatment effect. Defaults to 0 for RD and 1 for RR.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
bootstrap_typeOptional resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_confidence_interval()
Usage
InferenceIncidGCompRiskRatio$compute_m_out_of_n_bootstrap_confidence_interval( alpha = 0.05, B = 501, m = NULL, type = "basic", show_progress = TRUE, min_number_usable_samples = 5L, bootstrap_type = NULL )
Arguments
alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
alphaSignificance level. Default
0.05.alphaSignificance level. Default 0.05.
alphaSignificance level. Default 0.05.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of Bayesian-bootstrap samples.
BNumber of subsamples.
BNumber of subsamples.
BNumber of resamples.
BNumber of resamples.
mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval.mResample size. See
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval.typeBootstrap CI type. See
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval.typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.typeBayesian-bootstrap p-value type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval.typeBayesian-bootstrap CI type. See
InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval.typeP-value type.
typeConfidence-interval type.
typeP-value type.
typeConfidence-interval type.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite subsampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
min_number_usable_samplesMinimum number of finite resampled estimates required.
bootstrap_typeOptional resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
bootstrap_typeOptional empirical-resampling scheme.
InferenceIncidGCompRiskRatio$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidGCompRiskRatio$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
See Also
InferenceIncidGCompRiskDiff
for the risk-difference version of this same standardized logistic working model.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidGCompRiskRatio$new(seq_des)
inf$compute_estimate()
Conditional Logistic Plus GLMM IVWC Inference for KK Designs
Description
Fits a combined conditional-logit-plus-random-intercept-GLMM likelihood for
incidence responses under a KK matching-on-the-fly design, where
reservoir (unmatched) subjects are excluded from the GLMM
component (private$combine_reservoir_into_glmm() == FALSE): only
concordant matched pairs contribute their random-intercept GLMM likelihood
alongside the discordant-pair conditional-logit term, both sharing a single
treatment coefficient \beta_T. See
InferencePropKKGLMM for the full
model form (conditional-logit-on-discordant plus random-intercept-GLMM,
jointly maximized) and
InferenceAbstractKKCondLogitGLMM
for the shared fitting/caching contract. Contrast with the sibling
InferenceIncidKKCondLogitGLMMOneLik,
which instead includes reservoir subjects in the GLMM component
(combine_reservoir_into_glmm() == TRUE) — this class's naming
("IVWC") reflects that reservoir information, when used, is intended to be
combined with this fit's estimate via inverse-variance weighting rather
than folded into the same likelihood.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Super classes
Inference -> InferenceAbstractKKCondLogitGLMM -> InferenceIncidKKCondLogitGLMMIVWC
Methods
Public methods
+ inherited public methods from InferenceAbstractKKCondLogitGLMM
InferenceAbstractKKCondLogitGLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_estimate()InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKCondLogitGLMM$compute_gradient_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_gradient_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_jackknife_bias_estimate()InferenceAbstractKKCondLogitGLMM$compute_jackknife_estimate()InferenceAbstractKKCondLogitGLMM$compute_jackknife_std_error()InferenceAbstractKKCondLogitGLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_estimate()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_pval()InferenceAbstractKKCondLogitGLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_rand_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_rand_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_score_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_score_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_subsampling_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_subsampling_sensitivity()InferenceAbstractKKCondLogitGLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_wald_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_wald_two_sided_pval()InferenceAbstractKKCondLogitGLMM$get_information_preference()InferenceAbstractKKCondLogitGLMM$get_information_source_used()InferenceAbstractKKCondLogitGLMM$get_last_param_bootstrap_diagnostics()InferenceAbstractKKCondLogitGLMM$get_last_param_bootstrap_estimate_diagnostics()InferenceAbstractKKCondLogitGLMM$get_mod()InferenceAbstractKKCondLogitGLMM$get_summary()InferenceAbstractKKCondLogitGLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_information_preferences()InferenceAbstractKKCondLogitGLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_testing_types()InferenceAbstractKKCondLogitGLMM$get_testing_type()InferenceAbstractKKCondLogitGLMM$initialize()InferenceAbstractKKCondLogitGLMM$select_optimal_b_subsampling()InferenceAbstractKKCondLogitGLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKCondLogitGLMM$set_information_preference()InferenceAbstractKKCondLogitGLMM$set_testing_type()InferenceAbstractKKCondLogitGLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKCondLogitGLMMIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKCondLogitGLMMIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKCondLogitGLMMIVWC$new(seq_des)
inf$compute_estimate()
Conditional Logistic Plus GLMM Combined-Likelihood Inference for KK Designs
Description
Fits a combined conditional-logit-plus-random-intercept-GLMM likelihood for
incidence responses under a KK matching-on-the-fly design, where
reservoir (unmatched) subjects are included in the GLMM component
(private$combine_reservoir_into_glmm() == TRUE), so all subjects
(discordant matched pairs, concordant matched pairs, and reservoir) enter
one joint likelihood with a single treatment coefficient \beta_T. See
InferencePropKKGLMM for the full
model form (conditional-logit-on-discordant plus random-intercept-GLMM,
jointly maximized) and
InferenceAbstractKKCondLogitGLMM
for the shared fitting/caching contract. Contrast with the sibling
InferenceIncidKKCondLogitGLMMIVWC,
which excludes reservoir subjects from the GLMM component.
Super classes
Inference -> InferenceAbstractKKCondLogitGLMM -> InferenceIncidKKCondLogitGLMMOneLik
Methods
Public methods
+ inherited public methods from InferenceAbstractKKCondLogitGLMM
InferenceAbstractKKCondLogitGLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_estimate()InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKCondLogitGLMM$compute_gradient_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_gradient_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_jackknife_bias_estimate()InferenceAbstractKKCondLogitGLMM$compute_jackknife_estimate()InferenceAbstractKKCondLogitGLMM$compute_jackknife_std_error()InferenceAbstractKKCondLogitGLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_estimate()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_pval()InferenceAbstractKKCondLogitGLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_rand_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_rand_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_score_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_score_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_subsampling_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_subsampling_sensitivity()InferenceAbstractKKCondLogitGLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_wald_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_wald_two_sided_pval()InferenceAbstractKKCondLogitGLMM$get_information_preference()InferenceAbstractKKCondLogitGLMM$get_information_source_used()InferenceAbstractKKCondLogitGLMM$get_last_param_bootstrap_diagnostics()InferenceAbstractKKCondLogitGLMM$get_last_param_bootstrap_estimate_diagnostics()InferenceAbstractKKCondLogitGLMM$get_mod()InferenceAbstractKKCondLogitGLMM$get_summary()InferenceAbstractKKCondLogitGLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_information_preferences()InferenceAbstractKKCondLogitGLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_testing_types()InferenceAbstractKKCondLogitGLMM$get_testing_type()InferenceAbstractKKCondLogitGLMM$select_optimal_b_subsampling()InferenceAbstractKKCondLogitGLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKCondLogitGLMM$set_information_preference()InferenceAbstractKKCondLogitGLMM$set_testing_type()InferenceAbstractKKCondLogitGLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKCondLogitGLMMOneLik$new()
Initialize inference for the combined conditional-logit
(discordant matched pairs) plus random-intercept-GLMM (concordant pairs
and reservoir subjects) incidence model; see
InferenceIncidKKCondLogitGLMMOneLik
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Usage
InferenceIncidKKCondLogitGLMMOneLik$new( des_obj, model_formula = NULL, max_abs_reasonable_coef = 50, max_abs_reasonable_se = 1.25, max_abs_log_sigma = 8, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with an incidence response.model_formulaOptional formula for covariate adjustment.
max_abs_reasonable_coefCap for reasonable coefficient estimates.
max_abs_reasonable_seCap for reasonable treatment standard errors.
max_abs_log_sigmaCap for reasonable log random effect variance.
verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart optimizer start values.
optimization_algCharacter. Optimization algorithm (default "lbfgs").
InferenceIncidKKCondLogitGLMMOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKCondLogitGLMMOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKCondLogitGLMMOneLik$new(seq_des)
inf$compute_estimate()
Conditional Logistic IVWC Inference (KK Designs, Binary Response)
Description
Inverse-variance-weighted combination (IVWC) of two independently fit
conditional-likelihood pieces for KK matched-pair-plus-reservoir binary
designs: matched pairs are analyzed with exact conditional logistic
regression (conditional_logit_fit_matched_pairs(), which
conditions out the pair-specific nuisance intercept and estimates only the
treatment log-odds-ratio \beta_T from discordant pairs, or the joint
clogit-style likelihood when covariates are present), and reservoir
subjects are analyzed with ordinary logistic regression
(conditional_logit_fit_reservoir()). If \hat\beta_m,
\hat\sigma^2_m and \hat\beta_r, \hat\sigma^2_r are the matched-pair
and reservoir estimates and their variances, the combined estimate is the
variance-weighted average
\hat\beta_T = w^\star \hat\beta_m + (1-w^\star) \hat\beta_r, \quad
w^\star = \frac{\hat\sigma^2_r}{\hat\sigma^2_r + \hat\sigma^2_m},
with combined variance \hat\sigma^2_m \hat\sigma^2_r / (\hat\sigma^2_m
+ \hat\sigma^2_r). This is the classical fixed-effects inverse-variance
meta-analysis pooling formula (see Cochrane Handbook / DerSimonian-Laird),
applied here to combine the two conditionally-independent likelihood
contributions of a KK design rather than to pool separate studies. When
only one of the two components is estimable the combined estimate falls
back to that component alone. Contrast this with
InferenceIncidKKCondLogitOneLik, which instead fits a single joint
likelihood over both pieces (see that class's documentation) –
likelihood_tier = "partial" here reflects that the matched-pair
piece is a genuine conditional (partial) likelihood, but the two-piece
combination itself is a closed-form Wald/meta-analytic step, not a further
likelihood evaluation.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceIncidKKCondLogitIVWC
Methods
Public methods
-
InferenceIncidKKCondLogitIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceIncidKKCondLogitIVWC$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKCondLogitIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceIncidKKCondLogitIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceIncidKKCondLogitIVWC$supports_rand_pval_for_incidence()
Usage
InferenceIncidKKCondLogitIVWC$supports_rand_pval_for_incidence()
InferenceIncidKKCondLogitIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKCondLogitIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Fleiss, J.L., Levin, B., Paik, M.C. (2003). Statistical Methods for Rates and Proportions, 3rd ed. Wiley. (conditional logistic regression for matched pairs)
See Also
InferenceIncidKKCondLogitOneLik
for the one-likelihood alternative combining strategy.
One-Likelihood Conditional-Logistic Inference for KK Binary Designs
Description
Estimates a treatment log-odds-ratio \beta_T for binary
(incidence) outcomes collected under a KK matching-on-the-fly design
(DesignSeqOneByOneKK14 or
subclass) by maximizing one combined likelihood that couples a
conditional-logistic (intercept-free, within-matched-pair) likelihood
for matched subjects with an ordinary logistic likelihood for reservoir
subjects, sharing a single treatment coefficient across both pieces.
This is the "one-likelihood" counterpart to
InferenceIncidKKCondLogitIVWC,
which instead fits the matched and reservoir pieces separately and
pools them by inverse-variance weighting; here the treatment coefficient
is a single joint MLE, and likelihood_tier = "full" exposes
likelihood-ratio, score, and gradient inference plus a parametric
likelihood bootstrap in addition to Wald.
Estimand. \beta_T, the treatment coefficient of a
logistic mean model \mathrm{logit}(P(Y=1 \mid w,x)) = \beta_0 +
\beta_T w + x\beta; \exp(\hat\beta_T) is the treatment-vs-control
odds ratio.
Model. Matched pairs contribute McFadden-style conditional
logistic likelihood terms that condition away the pair-specific nuisance
intercept (see build_matching_combined_clogit_design_cpp/
collect_discordant_pairs_cpp); reservoir subjects contribute an
ordinary logistic likelihood with one shared intercept. The combined
negative log-likelihood is minimized jointly in
(\beta_0, \beta_T, \beta) via fast_logistic_regression_cpp/
fast_logistic_regression_with_var_cpp. When
get_testing_type() != "wald", asymptotic CI/p-value calls are
routed through InferenceAsympLik's
generic score/likelihood-ratio/gradient dispatch instead of the
design's own Wald machinery.
Assumptions. Independence across matched pairs and reservoir
subjects given covariates; correct logistic mean specification; a KK
matching-on-the-fly design supplying the matched/reservoir partition.
No response censoring is supported (checked at construction via
assertNoCensoring()).
Super class
Inference -> InferenceIncidKKCondLogitOneLik
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKCondLogitOneLik$compute_rand_two_sided_pval()
Computes a randomization-based two-sided p-value for the
treatment effect, preflighting the observed combined-likelihood
treatment statistic (see the class-header note above) before
delegating to InferenceRandCI's
Zhang-dispatch-aware implementation.
Usage
InferenceIncidKKCondLogitOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, type = NULL, args_for_type = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization (permutation) draws.
deltaThe null treatment effect. Default 0.
transform_responsesOptional response transform applied before the randomization statistic is computed. Default
"none".na.rmWhether to remove non-finite permutation replicates.
show_progressWhether to show a progress bar.
permutationsOptional pre-generated permutation matrix/list to reuse instead of drawing new permutations.
typeOptional randomization-statistic type override.
args_for_typeOptional list of extra arguments for
type.zero_one_logit_clampClamp applied to responses at the 0/1 boundary before a logit-scale transform, to avoid infinite values. Default
.Machine$double.eps.
InferenceIncidKKCondLogitOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKCondLogitOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A., and Krieger, A. M. (2014). "Matching on-the-fly: Sequential
allocation with higher power and efficiency." Biometrics, 70(2),
378-388. doi:10.1111/biom.12148. (KK14 in REFERENCES.md.)
See Also
Analogous Python API for conditional logistic regression:
statsmodels
discrete models (ConditionalLogit).
Logistic
regression (orientation).
G-Computation Risk-Difference Inference for KK Designs with Binary Responses
Description
Fits an all-subject logistic working model \mathrm{logit}\,P(Y=1\mid w,x) =
\beta_0 + \beta_T w + \beta_X^\top x for a KK incidence outcome using treatment
w and, optionally, all recorded covariates x, then estimates the
marginal (standardized, g-computation) risk difference
\hat\theta = n^{-1}\sum_i \{\hat p(1, x_i) - \hat p(0, x_i)\} by averaging
the fitted-model predicted risks under all-treated and all-control assignments
over the empirical covariate distribution (Robins 1986). Matched pairs are
treated as clusters and reservoir subjects are treated as singletons when
computing the sandwich covariance of the standardized estimator (the
delta-method variance of the empirical mean of the two counterfactual-risk
contrasts, not the naive logistic-regression coefficient variance).
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Details
This estimator has likelihood_tier = "none": the fitted logistic model
is a working model for standardization only, and reported inference is
sandwich/bootstrap-based, not likelihood-based. compute_estimate() fits
the model and returns \hat\theta on the risk-difference (probability)
scale; compute_estimate_with_bootstrap_weights() refits under
Bayesian-bootstrap subject weights for compute_bayesian_bootstrap_confidence_interval().
Jackknife deletes one cluster (matched pair or singleton reservoir subject) at
a time. If the working model fails to converge or the design has no
treatment-arm variation, the estimate is marked non-estimable via
is_nonestimable().
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceIncidKKGCompRiskDiff
Methods
Public methods
-
InferenceIncidKKGCompRiskDiff$approximate_randomization_distribution_beta_hat_T() -
InferenceIncidKKGCompRiskDiff$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKGCompRiskDiff$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceIncidKKGCompRiskDiff$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceIncidKKGCompRiskDiff$supports_rand_pval_for_incidence()
Usage
InferenceIncidKKGCompRiskDiff$supports_rand_pval_for_incidence()
InferenceIncidKKGCompRiskDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKGCompRiskDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period. Mathematical Modelling, 7(9-12), 1393-1512. doi:10.1016/0270-0255(86)90088-6
See Also
InferenceIncidKKGCompRiskRatio
for the risk-ratio analog on the same standardization machinery.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKGCompRiskDiff$new(seq_des)
inf$compute_estimate()
G-Computation Risk-Ratio Inference for KK Designs with Binary Responses
Description
Fits the same all-subject logistic working model as
InferenceIncidKKGCompRiskDiff
for a KK incidence outcome using treatment and, optionally, all recorded
covariates, then estimates the marginal (standardized, g-computation) risk
ratio \hat\theta = \left(n^{-1}\sum_i \hat p(1, X_i)\right) /
\left(n^{-1}\sum_i \hat p(0, X_i)\right) by averaging fitted-model predicted
risks under all-treated and all-control assignments over the empirical
covariate distribution (Robins 1986). Matched pairs are treated as clusters
and reservoir subjects are treated as singletons when computing the sandwich
covariance; the delta method is applied on the log-risk-ratio scale to keep
the reported ratio and its confidence interval positive, then
back-transformed for reporting.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Details
This estimator has likelihood_tier = "none". If the working model
fails to converge, has no treatment-arm variation, or the all-control
standardized risk is zero (undefined ratio), the estimate is marked
non-estimable via is_nonestimable().
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceIncidKKGCompRiskRatio
Methods
Public methods
-
InferenceIncidKKGCompRiskRatio$approximate_randomization_distribution_beta_hat_T() -
InferenceIncidKKGCompRiskRatio$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKGCompRiskRatio$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceIncidKKGCompRiskRatio$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceIncidKKGCompRiskRatio$supports_rand_pval_for_incidence()
Usage
InferenceIncidKKGCompRiskRatio$supports_rand_pval_for_incidence()
InferenceIncidKKGCompRiskRatio$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKGCompRiskRatio$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Robins, J. (1986). A new approach to causal inference in mortality studies with a sustained exposure period. Mathematical Modelling, 7(9-12), 1393-1512. doi:10.1016/0270-0255(86)90088-6
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKGCompRiskRatio$new(seq_des)
inf$compute_estimate()
GEE Inference for KK Designs with Binary Response
Description
Fits a Generalized Estimating Equations (GEE) model with a binomial family
and logit link, \mathrm{logit}\,\Pr(Y_i = 1 \mid x_i) = x_i^\top\beta,
for binary (incidence) responses under a KK matching-on-the-fly design, using
an exchangeable working correlation structure where each cluster is
either a matched pair (2 members) or a reservoir singleton (1 member) — see
$compute_estimate()'s method-level documentation for the full fitting
contract (internal Rcpp solver vs. geepack fallback, hardening/retry
behavior). GEE is used here purely to fit one marginal model jointly across
matched-pair and reservoir subjects while accounting for the within-pair
correlation the matching induces, not as a longitudinal/repeated-measures
tool. Inference is quasi-likelihood/estimating-equation based
(likelihood_tier = "quasi"): standard errors are GEE sandwich (robust)
standard errors, not model-likelihood-based.
Super class
Inference -> InferenceIncidKKGEE
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKGEE$new()
Initialize KK binary-response GEE inference, validate the
matched/reservoir design, and prepare the exchangeable-working-correlation
binomial (logit-link) GEE fitting machinery used by
InferenceIncidKKGEE.
Usage
InferenceIncidKKGEE$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with an incidence response.model_formulaOptional formula for covariate adjustment.
use_rcppWhether to use the internal Rcpp GEE solver (
TRUE, default) with automatic fallback togeepack::geeglmon failure, or always usegeepack::geeglmdirectly (FALSE).verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceIncidKKGEE$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKGEE$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the GEE estimating-equation framework and sandwich variance estimator used here.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKGEE$new(seq_des)
inf$compute_estimate()
Modified-Poisson Inference for KK Designs with Binary Responses
Description
Fits Zou's (2004) modified-Poisson working model for binary incidence
outcomes under a KK matching-on-the-fly design: a log-link Poisson model
\log E[Y_i \mid w_i, x_i] = \beta_0 + \beta_T w_i + x_i^\top \gamma
is fit to the binary (0/1) response by ordinary Poisson maximum likelihood
(a working, misspecified likelihood — the true response is Bernoulli, not
Poisson), and the coefficient standard errors are corrected by a
cluster-robust sandwich covariance rather than the (invalid, for a
misspecified likelihood) model-based Poisson information. Matched pairs
are treated as clusters (2 members) and reservoir subjects as singleton
clusters when computing the sandwich covariance, so the matched-pair
correlation induced by the design is accounted for even though the
modified-Poisson working model itself does not encode it directly.
\exp(\hat\beta_T) is the estimated risk ratio, directly interpretable
unlike a logistic regression's odds ratio (which only approximates the
risk ratio when the outcome is rare). likelihood_tier = "none"
(the sandwich-corrected inference is not a normalized model likelihood):
only Wald inference is exposed. See
InferenceAbstractKKMarginalIncid
for the shared marginal-incidence fitting contract.
Super classes
Inference -> InferenceAbstractKKMarginalIncid -> InferenceAbstractKKModifiedPoisson -> InferenceIncidKKModifiedPoisson
Methods
Public methods
+ inherited public methods from InferenceAbstractKKModifiedPoisson
+ inherited public methods from InferenceAbstractKKMarginalIncid
InferenceAbstractKKMarginalIncid$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKMarginalIncid$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_gradient_confidence_interval()InferenceAbstractKKMarginalIncid$compute_gradient_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_jackknife_bias_estimate()InferenceAbstractKKMarginalIncid$compute_jackknife_estimate()InferenceAbstractKKMarginalIncid$compute_jackknife_std_error()InferenceAbstractKKMarginalIncid$compute_jackknife_wald_confidence_interval()InferenceAbstractKKMarginalIncid$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bartlett_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_confidence_interval()InferenceAbstractKKMarginalIncid$compute_lik_ratio_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_param_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_param_bootstrap_estimate()InferenceAbstractKKMarginalIncid$compute_param_bootstrap_pval()InferenceAbstractKKMarginalIncid$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKMarginalIncid$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_rand_confidence_interval()InferenceAbstractKKMarginalIncid$compute_rand_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_score_confidence_interval()InferenceAbstractKKMarginalIncid$compute_score_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_subsampling_confidence_interval()InferenceAbstractKKMarginalIncid$compute_subsampling_sensitivity()InferenceAbstractKKMarginalIncid$compute_subsampling_two_sided_pval()InferenceAbstractKKMarginalIncid$compute_wald_confidence_interval()InferenceAbstractKKMarginalIncid$compute_wald_two_sided_pval()InferenceAbstractKKMarginalIncid$get_information_preference()InferenceAbstractKKMarginalIncid$get_information_source_used()InferenceAbstractKKMarginalIncid$get_last_param_bootstrap_diagnostics()InferenceAbstractKKMarginalIncid$get_last_param_bootstrap_estimate_diagnostics()InferenceAbstractKKMarginalIncid$get_mod()InferenceAbstractKKMarginalIncid$get_summary()InferenceAbstractKKMarginalIncid$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKMarginalIncid$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKMarginalIncid$get_supported_bootstrap_ci_types()InferenceAbstractKKMarginalIncid$get_supported_bootstrap_pval_types()InferenceAbstractKKMarginalIncid$get_supported_information_preferences()InferenceAbstractKKMarginalIncid$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKMarginalIncid$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKMarginalIncid$get_supported_testing_types()InferenceAbstractKKMarginalIncid$get_testing_type()InferenceAbstractKKMarginalIncid$initialize()InferenceAbstractKKMarginalIncid$select_optimal_b_subsampling()InferenceAbstractKKMarginalIncid$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKMarginalIncid$set_information_preference()InferenceAbstractKKMarginalIncid$set_testing_type()InferenceAbstractKKMarginalIncid$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidKKModifiedPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidKKModifiedPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Zou, G. (2004). "A Modified Poisson Regression Approach to Prospective Studies with Binary Data." American Journal of Epidemiology, 159(7), 702-706, doi:10.1093/aje/kwh090; Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for.
See Also
InferenceIncidModifiedPoisson
for the non-KK analog.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKModifiedPoisson$new(seq_des)
inf$compute_estimate()
KK Newcombe Risk-Difference IVWC Inference for Binary Responses
Description
Initialize KK Newcombe risk-difference IVWC inference and
prepare matched/reservoir paired-binomial components used by
InferenceIncidKKNewcombeRiskDiff.
Computes the compound Newcombe risk-difference point estimate
\hat\theta = w_1 \hat\theta_1 + w_2 \hat\theta_2 on the risk-difference
(probability) scale, where \hat\theta_1 is the paired-Newcombe
discordant-pair estimate from matched pairs and \hat\theta_2 is the
independent-Newcombe estimate from reservoir subjects, combined by
inverse-variance weighting w_j \propto 1/\widehat{\mathrm{Var}}(\hat\theta_j)
(falls back to the single available component when one has zero subjects).
Caches intermediate match/reservoir statistics for reuse by
compute_asymp_confidence_interval() and
compute_asymp_two_sided_pval().
Usage
KKNewcombeRiskDiffIVWCSource
Details
Implements a compound Newcombe risk-difference estimator for KK designs. This class pools information from matched pairs (using the Paired Newcombe method) and the reservoir (using the Independent Newcombe method) via inverse-variance weighted combination (IVWC).
The matched-pair component applies the paired Newcombe (Method-10-style
Wilson-score) interval to the discordant pairs to estimate the treatment
effect and its variance (Newcombe 1998). The reservoir component applies the
independent-samples Newcombe interval, treating unmatched subjects as two
independent binomial samples. The two component estimates
\hat\theta_1, \hat\theta_2 are combined by inverse-variance weighting,
\hat\theta = w_1 \hat\theta_1 + w_2 \hat\theta_2, w_j =
(1/\hat V_j) / \sum_k (1/\hat V_k), the standard IVWC framework used
throughout the package's KK inference classes. likelihood_tier =
"none": this is a closed-form Wilson-score-type estimator, not a fitted
likelihood model, so no likelihood-ratio or parametric-bootstrap methods are
exposed. If a design has no matched pairs or no reservoir subjects, the
single available component is used directly rather than combined.
References
Newcombe, R. G. (1998). Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods. Statistics in Medicine, 17(8), 873-890. doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidKKNewcombeRiskDiff$new(seq_des)
inf$compute_estimate()
Log-Binomial Regression Inference for Incidence Responses
Description
Fits a binomial regression with the log link for binary
(incidence) responses: \log P(Y_i = 1) = \beta_0 + \beta_T W_i +
X_i^\top \gamma, where W_i is the treatment indicator and X_i
are optional recorded covariates, by maximum likelihood
(fast_log_binomial_regression_cpp/
fast_log_binomial_regression_weighted_cpp). \hat\beta_T
is a log risk ratio: \exp(\hat\beta_T) is the estimated
treatment risk ratio (relative risk) directly, unlike the log-odds-ratio
from InferenceIncidLogRegr's logit
link. likelihood_tier = "full": Wald, score, gradient, and
likelihood-ratio tests are all available when the model converges, plus
parametric-likelihood-bootstrap calibration of the likelihood-ratio test.
Because the log link does not constrain fitted probabilities to
[0,1] (only to [0,\infty)), fits are hardened by QR
column-dropping and a coefficient-magnitude cap
(max_abs_reasonable_coef) and rejected as nonestimable when the fit
is implausible — the same practical limitation as the identity-link
sibling
InferenceIncidBinomialIdentityRiskDiff, here applying to the upper rather
than both tails of the probability scale. Validity requires the
multiplicative log-linear risk model to be correctly specified over the
covariate range observed.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceIncidLogBinomial
Methods
Public methods
-
InferenceIncidLogBinomial$approximate_randomization_distribution_beta_hat_T() -
InferenceIncidLogBinomial$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidLogBinomial$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceIncidLogBinomial$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceIncidLogBinomial$supports_rand_pval_for_incidence()
Usage
InferenceIncidLogBinomial$supports_rand_pval_for_incidence()
InferenceIncidLogBinomial$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidLogBinomial$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the binomial GLM family and log-link relative-risk parameterization.
See Also
InferenceIncidLogRegr
(logit link, log-odds-ratio estimand),
InferenceIncidBinomialIdentityRiskDiff (identity link, risk-difference
estimand) for alternative link/estimand choices on the same response
type. Comparable Python API:
statsmodels GLM
(family=Binomial(link=log())). See also:
Generalized
linear model (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidLogBinomial$new(seq_des)
inf$compute_estimate()
Logistic Regression Inference for Incidence Responses
Description
Fits a logistic regression model for binary (incidence) responses:
\mathrm{logit}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma,
where W_i is the treatment indicator and X_i are optional
recorded covariates, by maximum likelihood
(fast_logistic_regression_cpp/
fast_logistic_regression_weighted_cpp). \hat\beta_T is
a log-odds-ratio: \exp(\hat\beta_T) is the estimated treatment odds
ratio. likelihood_tier = "full": Wald, score, gradient, and
likelihood-ratio tests are all available (via the shared
StandardModelCache model-caching contract), plus
parametric-likelihood-bootstrap calibration of the likelihood-ratio test.
A fit whose coefficients exceed
max_abs_reasonable_coef in magnitude (a proxy for near-perfect
separation) is cached as nonestimable rather than returned. Validity
requires the usual logistic-regression assumptions: correctly specified
linear predictor on the logit scale, independence across subjects
conditional on covariates, and no perfect/quasi-complete separation.
Estimand. Composes
MarginalEstimand
(set_estimand()/get_estimand()/get_supported_estimands()).
Under the default estimand = "conditional", \hat\beta_T is the
log-odds-ratio above. Under estimand = "marginal_mean_diff", the
reported quantity is instead the g-computation marginal risk difference
\frac{1}{n}\sum_i \{\mathrm{plogis}(\hat\beta_0 + \hat\beta_T + X_i^\top
\hat\gamma) - \mathrm{plogis}(\hat\beta_0 + X_i^\top \hat\gamma)\} — every
subject's covariates plugged in once under treatment and once under
control, averaged over the empirical covariate distribution. Under
estimand = "marginal_ratio", the log of the corresponding marginal
risk ratio. Because there is no latent submodel for this family (unlike
e.g. InferencePropZeroOneInflatedBetaRegr's
zero/one-inflation mixture), the marginal mean function is exactly the
model's own fitted mean; no separate standardization step beyond the
g-computation average is needed. Standard errors under a marginal estimand
use the delta method against the model's coefficient covariance (degrees
of freedom Inf); testing_type is restricted to "wald"
whenever the estimand is non-conditional (set_testing_type() errors
otherwise). The underlying model fit is identical regardless of estimand —
switching estimand is a pure post-fit transform, never a refit.
Super class
Inference -> InferenceIncidLogRegr
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidLogRegr$compute_estimate()
Fits the logistic regression model by maximum
likelihood. Under the default estimand = "conditional",
returns \hat\beta_T, the treatment log-odds-ratio. Under
estimand = "marginal_mean_diff"/"marginal_ratio"
(set via set_estimand()), returns the g-computation
marginal risk difference/log-risk-ratio instead — see the
class-level @details for the formula. The underlying
model fit is identical either way (a pure post-fit transform of
the same cached fit, no refit).
Usage
InferenceIncidLogRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceIncidLogRegr$compute_asymp_confidence_interval()
Wald confidence interval, dispatched by
testing_type for the conditional estimand (score/gradient/
likelihood-ratio/Bartlett available); under a marginal estimand
testing_type is always "wald" (the only value
set_estimand() permits there), so this always resolves to
the delta-method interval. Calls self$compute_estimate()
first (not private$shared() directly) so the
estimand-aware cache is always current regardless of call order.
Usage
InferenceIncidLogRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferenceIncidLogRegr$compute_asymp_two_sided_pval()
Wald two-sided p-value, dispatched by
testing_type exactly as
compute_asymp_confidence_interval(); see that method's
description for the marginal-estimand always-Wald note.
Usage
InferenceIncidLogRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment-effect value under the current estimand (conditional log-odds-ratio, or marginal risk difference/ log-risk-ratio).
InferenceIncidLogRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidLogRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the logistic regression model and its maximum-likelihood theory.
See Also
Comparable Python API: statsmodels GLM. See also: Logistic regression (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidLogRegr$new(seq_des)
inf$compute_estimate()
inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)
Miettinen-Nurminen Risk-Difference Inference for Binary Responses
Description
Fits the classical Miettinen-Nurminen score method for the risk difference in a
two-arm binary trial. The point estimate is the observed risk difference, while
confidence intervals and p-values are obtained by inverting the constrained
score test under the null p_T - p_C = \delta.
This class is intentionally unadjusted. It operates on the 2 \times 2
table induced by treatment assignment and incidence response (counts
x_T, x_C of events among n_T, n_C treated/control subjects), and
is therefore the natural classical binary-endpoint complement to the
regression-based incidence methods already in the package. The point
estimate is the plain risk difference \hat p_T - \hat p_C; the
Miettinen-Nurminen confidence interval and p-value invert a
restricted-maximum-likelihood score test of H_0: p_T - p_C =
\delta (via mn_ci_cpp/mn_pvalue_cpp) — under the null, the two
arms' event probabilities are re-estimated subject to the constraint
\hat p_T - \hat p_C = \delta, and the resulting score statistic is
compared to its asymptotic normal distribution, with a small-sample bias
correction factor (n_T+n_C)/(n_T+n_C-1) applied to the naive Wald
variance used elsewhere (e.g. in $compute_estimate_with_bootstrap_weights(),
which never gets this correction since it skips the score-test path
entirely). This score-based interval generally has better small-sample
coverage than the naive normal-approximation Wald interval on the risk
difference.
Super class
Inference -> InferenceIncidMiettinenNurminenRiskDiff
Methods
Public methods
-
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate_with_bootstrap_weights() -
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_confidence_interval() -
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidMiettinenNurminenRiskDiff$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a Miettinen-Nurminen risk-difference inference object for a completed design with an incidence response.
Usage
InferenceIncidMiettinenNurminenRiskDiff$new( des_obj, model_formula = NULL, verbose = FALSE )
Arguments
des_objA completed
DesignSeqOneByOneobject with an incidence response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "incidence")
for (i in 1:20) {
x_i = data.frame(x1 = rnorm(1), x2 = rnorm(1))
w_i = seq_des$add_one_subject_to_experiment_and_assign(x_i)
p_i = plogis(-0.8 + 0.5 * w_i)
seq_des$add_one_subject_response(i, rbinom(1, 1, p_i))
}
seq_des_inf = InferenceIncidMiettinenNurminenRiskDiff$new(seq_des)
seq_des_inf$compute_estimate()
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate()
Computes the observed (unadjusted) risk-difference estimate
\hat p_T - \hat p_C (see class documentation for the full
Miettinen-Nurminen inference model). NA if either arm is empty.
Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate_with_bootstrap_weights()
Recomputes the risk-difference estimate under subject/block
bootstrap weights: the weighted event proportions \hat p_T^w =
\sum_i r_i y_i \mathbb{1}[w_i=1] / \sum_i r_i \mathbb{1}[w_i=1] (and
analogously for control), differenced. Used by the Bayesian bootstrap
and related weighted-resampling machinery; see
InferenceBayesianBootstrap.
Always leaves the standard error and degrees of freedom unavailable
(NA) regardless of estimate_only — this weighted path
never computes the Miettinen-Nurminen score-based variance.
Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyPresent for interface parity; this method never computes variance components regardless of its value.
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_confidence_interval()
Computes a 1-\alpha Miettinen-Nurminen restricted-MLE
score confidence interval for the risk difference (see class
documentation for the full method), by bisection-inverting the score
test (mn_ci_cpp) to pval_epsilon tolerance.
Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_confidence_interval( alpha = 0.05, pval_epsilon = 1e-07 )
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.pval_epsilonBisection tolerance for CI bounds.
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_two_sided_pval()
Computes a two-sided Miettinen-Nurminen restricted-MLE
score p-value (mn_pvalue_cpp) testing H_0: p_T - p_C =
\code{delta} (see class documentation for the full method).
Usage
InferenceIncidMiettinenNurminenRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect on the risk-difference scale.
InferenceIncidMiettinenNurminenRiskDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidMiettinenNurminenRiskDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Miettinen, O., and Nurminen, M. (1985). "Comparative Analysis of Two Rates." Statistics in Medicine, 4(2), 213-226, doi:10.1002/sim.4780040211, for the restricted-maximum-likelihood score method used here.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidMiettinenNurminenRiskDiff$new(seq_des)
inf$compute_estimate()
## ------------------------------------------------
## Method `InferenceIncidMiettinenNurminenRiskDiff$new()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "incidence")
for (i in 1:20) {
x_i = data.frame(x1 = rnorm(1), x2 = rnorm(1))
w_i = seq_des$add_one_subject_to_experiment_and_assign(x_i)
p_i = plogis(-0.8 + 0.5 * w_i)
seq_des$add_one_subject_response(i, rbinom(1, 1, p_i))
}
seq_des_inf = InferenceIncidMiettinenNurminenRiskDiff$new(seq_des)
seq_des_inf$compute_estimate()
Modified Poisson Regression Inference for Incidence Responses
Description
Fits Zou's (2004) modified Poisson regression for binary (incidence)
responses: \log E[Y_i \mid w_i, x_i] = \beta_0 + \beta_T w_i +
x_i^\top \gamma, fit by maximizing the ordinary Poisson log-likelihood
treating the binary Y_i as if it were Poisson-distributed (a valid
estimating equation for the conditional mean regardless of the true
outcome distribution, exactly as
InferencePropFractionalLogit's
quasi-binomial fit is for fractional responses). \hat\beta_T is a
log risk ratio: \exp(\hat\beta_T) is the estimated treatment
relative risk, the same estimand as
InferenceIncidLogBinomial's
log-binomial model, but modified Poisson never produces a fit failure from
the [0,1]-probability constraint that a genuine binomial log-link
model can hit. Caveat: this implementation's standard error comes
from the ordinary (model-based) Poisson Fisher information
(fast_poisson_regression_with_var_cpp's ssq_b_j), not
a robust/sandwich correction — Zou's (2004) original proposal specifically
pairs the misspecified Poisson working model with a robust sandwich
variance estimator to obtain valid standard errors under the resulting
overdispersion; users needing the fully robust modified-Poisson variance
should treat this class's standard errors/CIs/p-values as approximate.
likelihood_tier = "full" metadata is set for component-composition
purposes, but private$supports_likelihood_tests() is hard
FALSE — only Wald inference is exposed
(get_supported_testing_types_impl() returns "wald" only), not
likelihood-ratio/score/gradient tests. Fits with implausible coefficients
or fitted linear predictors (checked via
private$is_modified_poisson_fit_reasonable(), capped by
max_abs_reasonable_coef/max_abs_reasonable_linear_predictor)
are cached as nonestimable rather than returned.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceIncidModifiedPoisson
Methods
Public methods
-
InferenceIncidModifiedPoisson$approximate_randomization_distribution_beta_hat_T() -
InferenceIncidModifiedPoisson$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidModifiedPoisson$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceIncidModifiedPoisson$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceIncidModifiedPoisson$supports_rand_pval_for_incidence()
Usage
InferenceIncidModifiedPoisson$supports_rand_pval_for_incidence()
InferenceIncidModifiedPoisson$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidModifiedPoisson$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Zou, G. (2004). "A Modified Poisson Regression Approach to Prospective Studies with Binary Data." American Journal of Epidemiology, 159(7), 702-706, doi:10.1093/aje/kwh090.
See Also
InferenceIncidLogBinomial
for the genuine log-binomial alternative with the same log-risk-ratio
estimand. Comparable Python API:
statsmodels
discrete models (Poisson family on binary data). See also:
Poisson
regression (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidModifiedPoisson$new(seq_des)
inf$compute_estimate()
Newcombe Risk-Difference Inference for Binary Responses
Description
Fits the Newcombe hybrid score method (Method 10) for the risk difference in a two-arm binary trial. This method constructs a confidence interval for the difference between two independent proportions by combining Wilson score intervals for each group.
This class is unadjusted and assumes independent samples (e.g. from a
Bernoulli design). It ignores any matched-pair structure if present; for
matched data, use InferenceIncidKKNewcombeRiskDiff. The point
estimate is the plain risk difference \hat p_T - \hat p_C. The
confidence interval (newcombe_independent_ci_cpp) is Newcombe's
"Method 10" hybrid score interval: separate Wilson score intervals
[\ell_T, u_T] and [\ell_C, u_C] are computed for each arm's
proportion individually, then combined into a difference interval via
[\hat p_T - \hat p_C - \sqrt{(\hat p_T - \ell_T)^2 + (u_C - \hat
p_C)^2},\ \hat p_T - \hat p_C + \sqrt{(u_T - \hat p_T)^2 + (\hat p_C -
\ell_C)^2}] — this avoids the boundary/coverage problems of the naive
Wald interval on a risk difference while remaining closed-form (no
iterative score-test inversion, unlike the Miettinen-Nurminen method in
InferenceIncidMiettinenNurminenRiskDiff).
The two-sided p-value has no closed form here: it is obtained by
numerically inverting the confidence interval (bisection via
stats::uniroot) to find the significance level at which delta
falls exactly on the interval boundary.
Super class
Inference -> InferenceIncidNewcombeRiskDiff
Methods
Public methods
-
InferenceIncidNewcombeRiskDiff$compute_estimate_with_bootstrap_weights() -
InferenceIncidNewcombeRiskDiff$compute_asymp_confidence_interval() -
InferenceIncidNewcombeRiskDiff$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidNewcombeRiskDiff$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a Newcombe risk-difference inference object for a completed design with an uncensored incidence response.
Usage
InferenceIncidNewcombeRiskDiff$new( des_obj, model_formula = NULL, verbose = FALSE )
Arguments
des_objA completed
DesignSeqOneByOneobject with an incidence response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
InferenceIncidNewcombeRiskDiff$compute_estimate()
Computes the observed (unadjusted) risk-difference estimate
\hat p_T - \hat p_C (see class documentation for the full
Newcombe interval method).
Usage
InferenceIncidNewcombeRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceIncidNewcombeRiskDiff$compute_estimate_with_bootstrap_weights()
Recomputes the risk-difference estimate under subject/block
bootstrap weights: the weighted event proportions \hat p_T^w =
\sum_i r_i y_i \mathbb{1}[w_i=1] / \sum_i r_i \mathbb{1}[w_i=1] (and
analogously for control), differenced. Used by the Bayesian bootstrap
and related weighted-resampling machinery; see
InferenceBayesianBootstrap.
Always leaves the standard error and degrees of freedom unavailable
(NA) regardless of estimate_only — the Newcombe interval
method has no separate variance quantity to compute on this path.
Usage
InferenceIncidNewcombeRiskDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyPresent for interface parity; this method never computes variance components regardless of its value.
InferenceIncidNewcombeRiskDiff$compute_asymp_confidence_interval()
Computes a 1-\alpha Newcombe hybrid Wilson-score
confidence interval for the risk difference (see class documentation
for the full formula), via newcombe_independent_ci_cpp.
Usage
InferenceIncidNewcombeRiskDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level.
InferenceIncidNewcombeRiskDiff$compute_asymp_two_sided_pval()
Computes a two-sided p-value testing H_0: p_T - p_C =
\code{delta} by numerically finding (stats::uniroot) the
significance level \alpha at which delta falls exactly on
the boundary of the Newcombe confidence interval (see class
documentation) — there is no closed-form p-value for this method.
Returns 1 if no root is found in (10^{-10}, 1-10^{-10})
(interpreted as delta being far inside the interval at every
plausible \alpha).
Usage
InferenceIncidNewcombeRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null risk difference.
InferenceIncidNewcombeRiskDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidNewcombeRiskDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Newcombe, R. G. (1998). "Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods." Statistics in Medicine, 17(8), 873-890, doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I, for "Method 10", the hybrid Wilson-score interval used here.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidNewcombeRiskDiff$new(seq_des)
inf$compute_estimate()
Probit Regression Inference for Incidence Responses
Description
Fits a probit regression model for binary (incidence) responses:
\Phi^{-1}(P(Y_i = 1)) = \beta_0 + \beta_T W_i + X_i^\top \gamma,
where \Phi is the standard normal CDF, W_i is the treatment
indicator, and X_i are optional recorded covariates, by maximum
likelihood (fast_probit_regression_cpp/
fast_probit_regression_weighted_cpp). Unlike
InferenceIncidLogRegr's logit
link, \hat\beta_T here is not an odds-ratio scale parameter: it is
the treatment's additive effect on the latent standard-normal index
underlying the binary outcome. likelihood_tier = "full": Wald,
score, gradient, and likelihood-ratio tests are all available when the
model converges, plus parametric-likelihood-bootstrap calibration of the
likelihood-ratio test. A fit whose coefficients exceed
max_abs_reasonable_coef in magnitude (a proxy for near-perfect
separation) is cached as nonestimable rather than returned. Validity
requires the usual probit assumptions: correctly specified linear
predictor on the latent-normal scale, independence across subjects
conditional on covariates, and no perfect/quasi-complete separation.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceIncidProbitRegr
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidProbitRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceIncidProbitRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceIncidProbitRegr$supports_rand_pval_for_incidence()
Usage
InferenceIncidProbitRegr$supports_rand_pval_for_incidence()
InferenceIncidProbitRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidProbitRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
McCullagh, P., and Nelder, J. A. (1989). Generalized Linear Models (2nd ed.). Chapman and Hall/CRC, for the binomial GLM family and probit link.
See Also
InferenceIncidLogRegr
for the logit-link alternative with a log-odds-ratio estimand.
Comparable Python API:
statsmodels GLM
(family=Binomial(link=probit())). See also:
Probit model
(Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidProbitRegr$new(seq_des)
inf$compute_estimate()
Risk Difference Inference for Incidence Responses
Description
Fits a linear probability model E[Y \mid w, x] = \beta_0 + \beta_T w +
\beta_X^\top x via ordinary least squares for binary (incidence) responses
Y \in \{0,1\}, using the treatment indicator w and, optionally,
all recorded covariates x as predictors. \hat\beta_T is reported
directly as the risk-difference estimate: because Y is 0/1, the OLS fit
coincides with a saturated/linear model for the conditional risk
P(Y=1\mid w,x), so the coefficient on w is already on the
risk-difference (probability) scale with no back-transformation needed. This
is a misspecified working model for a binary response (heteroskedastic,
errors not Gaussian: \mathrm{Var}(Y \mid w,x) = P(w,x)(1-P(w,x)) varies
by treatment arm and covariates, so a single pooled residual variance is
the wrong variance model), so likelihood_tier = "none": standard
errors and the Wald CI use the Huber-White (HC0) sandwich
variance of \hat\beta_T – (X'X)^{-1} X'\,\mathrm{diag}(e_i^2)\,X\,
(X'X)^{-1}, from the OLS residuals e_i – not the classical
homoskedastic OLS variance and not a binomial likelihood, and no
likelihood-ratio or parametric-bootstrap methods are exposed.
Super class
Inference -> InferenceIncidRiskDiff
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidRiskDiff$new()
Uses the randomization-CI layer's two-sided p-value contract
(InferenceRandCI's version, not
InferenceRand's): for incidence responses it dispatches to the
Zhang exact randomization test rather than refusing outright, matching
this class's pre-migration old-ladder behavior. This deliberately
differs from the InferenceAllSimpleAverageDiff-family precedent of
pinning InferenceRand's version, which would have regressed the
working Zhang dispatch this class had on the old ladder.
Initialize a risk-difference inference object.
Usage
InferenceIncidRiskDiff$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with an incidence response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
rNumber of randomization vectors. @param delta Null difference.
transform_responsesTransformation. @param na.rm Remove NAs.
show_progressShow progress. @param permutations Pre-computed permutations.
typeOptional exact-inference type for incidence dispatch.
args_for_typeOptional arguments for
type.zero_one_logit_clampClamp for exact 0/1 values when logging.
InferenceIncidRiskDiff$compute_estimate()
Fits the OLS linear-probability model and returns the
risk-difference point estimate \hat\beta_T, the coefficient on
treatment. On a hardened design (private$harden), or when
estimate_only = FALSE, delegates to fast_ols_with_var_cpp()
via the shared model cache so the variance is available for later
confidence-interval/p-value calls without refitting.
Usage
InferenceIncidRiskDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf
TRUE, skip the variance/degrees-of-freedom computation and return only\hat\beta_T(cheaper for simulation/randomization callers that never request inference).
InferenceIncidRiskDiff$compute_asymp_confidence_interval()
Wald confidence interval for the risk difference,
\hat\beta_T \pm z_{1-\alpha/2}\, \hat s(\hat\beta_T), using the
Huber-White (HC0) sandwich standard error from the cached
model fit (see the class-level documentation) and a normal-quantile
multiplier – fast_ols_with_var_cpp() doesn't report residual
degrees of freedom, so private$compute_z_or_t_ci_from_s_and_df()
always falls back to its z branch here, never a t
reference, regardless of n. Interval bounds are not clamped to
[-1, 1]; a linear-probability model can produce out-of-range
endpoints near the boundary of the covariate space.
Usage
InferenceIncidRiskDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaTwo-sided miscoverage rate; the returned interval has nominal coverage
1-\alpha.
InferenceIncidRiskDiff$compute_asymp_two_sided_pval()
Two-sided Wald test of H_0: \beta_T = \delta vs.
H_1: \beta_T \neq \delta, via the Wald statistic
(\hat\beta_T - \delta) / \hat s(\hat\beta_T) (\hat s the
Huber-White sandwich SE) referred to a standard normal distribution
– same z, not t, reference as
$compute_asymp_confidence_interval(), for the same reason
(no residual degrees of freedom are ever cached for this class).
Usage
InferenceIncidRiskDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull risk-difference value under
H_0(default0, no treatment effect).
InferenceIncidRiskDiff$compute_estimate_with_bootstrap_weights()
Refits the linear-probability model by weighted least squares
(stats::lm.wfit()) under subject/block resampling weights
(nonparametric-bootstrap replicate weights or Bayesian-bootstrap Dirichlet
weights, both expanded to row weights via
expand_subject_or_block_weights_to_row_weights()) and returns the
re-estimated treatment coefficient. Rows with non-finite or non-positive
weight are excluded; if too few positive-weight rows remain to identify
the design, the replicate estimate is NA_real_.
Usage
InferenceIncidRiskDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric weights, one per subject or per resampling block (matched pair/cluster), as produced by the bootstrap or Bayesian-bootstrap resampling machinery.
estimate_onlyAccepted for interface compatibility; standard errors are never computed for a single bootstrap replicate regardless of this flag (only the point estimate is used to build the bootstrap distribution).
InferenceIncidRiskDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidRiskDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidRiskDiff$new(seq_des)
inf$compute_estimate()
Wald Incidence Inference
Description
Unadjusted incidence inference using the empirical risk difference
\hat\theta = \bar Y_T - \bar Y_C (sample proportions in the treatment
and control arms) together with the standard unpooled Wald standard error
\hat s(\hat\theta) = \sqrt{\bar Y_T(1-\bar Y_T)/n_T +
\bar Y_C(1-\bar Y_C)/n_C} and normal-approximation confidence interval /
hypothesis test \hat\theta \pm z_{1-\alpha/2}\,\hat s(\hat\theta). This
is the classical two-proportion Wald interval (e.g. Wald 1943; see
InferenceIncidNewcombeRiskDiff
for a small-sample-robust alternative). likelihood_tier = "none": no
likelihood is fit, so no likelihood-ratio or parametric-bootstrap methods are
exposed; the reservoir/covariate structure is ignored, unlike
InferenceIncidRiskDiff's covariate-
adjusted linear-probability model. Non-estimable if either arm has zero
subjects.
Super class
Inference -> InferenceIncidWald
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceIncidWald$new()
Uses the randomization-CI layer's two-sided p-value contract
(InferenceRandCI's version, not InferenceRand's): for
incidence responses this dispatches to the Zhang exact randomization
test where applicable rather than refusing outright, matching this
class's pre-migration old-ladder behavior (it inherited from
InferenceAllSimpleAverageDiff, whose own pin was already
corrected to InferenceRandCI – see that file's identical
rationale). This class independently composes the same components
rather than truly inheriting InferenceAllSimpleAverageDiff, so it
had its own stale copy of the old InferenceRand pin, which
silently regressed Zhang dispatch ('compute_rand_two_sided_pval()'
started throwing "Randomization tests are not supported for
incidence" for the same designs the old ladder handled correctly) –
found via test-incid-wald-migration-golden.R's
randomization_pval case going from '"ok"' to '"unsupported"'.
Pins asymptotic CI dispatch to the composed Wald
component's implementation (InferenceAsymp), which uses
private$get_standard_error() (this class's two-proportion
Wald SE) and private$get_degrees_of_freedom(). Without this
explicit pin, the SimpleMeanDifference component – composed
after Wald in this class's components list – silently
wins the assembly-order collision and dispatches its own Welch's
t-test on raw y instead, making the documented Wald formula
dead code. See class documentation.
Pins asymptotic p-value dispatch to the composed
Wald component's implementation; see
$compute_asymp_confidence_interval() for the rationale.
Initialize Wald risk-difference incidence inference and
prepare the treatment/control binomial summaries used by
InferenceIncidWald and related
InferenceIncidRiskDiff
methods.
Usage
InferenceIncidWald$new(des_obj, model_formula = NULL, verbose = FALSE)
Arguments
des_objA completed design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.deltaNull treatment effect value.
Returns
A new InferenceIncidWald object.
InferenceIncidWald$clone()
The objects of this class are cloneable with this method.
Usage
InferenceIncidWald$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'incidence')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(rbinom(10, 1, 0.5))
inf = InferenceIncidWald$new(seq_des)
inf$compute_estimate()
Jackknife-based Inference
Description
Abstract class for delete-1 jackknife estimate correction and jackknife-Wald inference layered on top of bootstrap-capable inference classes.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife
Methods
Public methods
-
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T() -
InferenceJackknife$compute_jackknife_wald_confidence_interval()
+ inherited public methods from InferenceBayesianBootstrap
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()
Returns the leave-one-out jackknife estimate distribution.
Usage
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T(unit = "auto")
Arguments
unitDeletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.
Returns
A numeric vector of jackknife replicate estimates.
InferenceJackknife$compute_jackknife_estimate()
Computes the delete-1 jackknife bias-corrected treatment estimate.
For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.
Usage
InferenceJackknife$compute_jackknife_estimate(unit = "auto")
Arguments
unitDeletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.
Returns
A numeric jackknife bias-corrected treatment estimate.
InferenceJackknife$compute_jackknife_bias_estimate()
Computes the jackknife bias estimate.
Usage
InferenceJackknife$compute_jackknife_bias_estimate(unit = "auto")
Arguments
unitDeletion unit. Default '\"auto\"'.
Returns
A numeric jackknife bias estimate.
InferenceJackknife$compute_jackknife_std_error()
Computes the delete-1 jackknife standard error.
For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.
Usage
InferenceJackknife$compute_jackknife_std_error(unit = "auto")
Arguments
unitDeletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.
Returns
A numeric jackknife standard error.
InferenceJackknife$compute_jackknife_wald_two_sided_pval()
Computes a two-sided Wald p-value using the jackknife estimate and jackknife standard error.
For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.
Usage
InferenceJackknife$compute_jackknife_wald_two_sided_pval( delta = 0, unit = "auto" )
Arguments
deltaNull treatment-effect value. Default 0.
unitDeletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.
Returns
A two-sided jackknife-Wald p-value.
InferenceJackknife$compute_jackknife_wald_confidence_interval()
Computes a normal-approximation confidence interval using the jackknife estimate and jackknife standard error.
For blocking designs, this uses leave-one-block-out deletion units. For matching designs, it uses leave-match-out deletion units. For KK designs, it uses leave-match-out for matched pairs and leave-one-out for reservoir subjects.
Usage
InferenceJackknife$compute_jackknife_wald_confidence_interval( alpha = 0.05, unit = "auto" )
Arguments
alphaSignificance level. Default 0.05.
unitDeletion unit. Default '\"auto\"', which chooses a design-aware unit automatically.
Returns
A jackknife-Wald confidence interval.
InferenceJackknife$clone()
The objects of this class are cloneable with this method.
Usage
InferenceJackknife$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Internal Base Class for KK Matching-on-the-Fly Designs
Description
Internal method. An abstract R6 class that provides relevant methods when the designs are KK matching-on-the-fly.
Initialize
Recomputes the class-specific treatment estimate under bootstrap weights; see
InferenceBayesianBootstrap.
Usage
kk_passthrough_compound_host_public
Inference for A Sequential Design
Description
An abstract R6 Class that provides asymptotic tests and intervals for a treatment effect in a sequential design where the common denominator is a summary table from a glm.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable
Methods
Public methods
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()InferenceJackknife$compute_jackknife_bias_estimate()InferenceJackknife$compute_jackknife_estimate()InferenceJackknife$compute_jackknife_std_error()InferenceJackknife$compute_jackknife_wald_confidence_interval()InferenceJackknife$compute_jackknife_wald_two_sided_pval()
+ inherited public methods from InferenceBayesianBootstrap
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceMLEorKMSummaryTable$compute_estimate()
Computes the appropriate estimate for mean difference
Usage
InferenceMLEorKMSummaryTable$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
The setting-appropriate (see description) numeric estimate of the treatment effect
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous") seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10]) seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43)) seq_des_inf = InferenceContinOLS$new(seq_des) seq_des_inf$compute_estimate()
InferenceMLEorKMSummaryTable$compute_asymp_confidence_interval()
Computes a 1-alpha level frequentist confidence interval
Usage
InferenceMLEorKMSummaryTable$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
Returns
A (1 - alpha)-sized frequentist confidence interval for the treatment effect
InferenceMLEorKMSummaryTable$compute_asymp_two_sided_pval()
Compute a two-sided p-value for model-summary-table
inference by using the cached treatment estimate and standard error from
the fitted model or Kaplan-Meier summary. See
InferenceMLEorKMSummaryTable
and InferenceAsymp.
Usage
InferenceMLEorKMSummaryTable$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. For any treatment effect at all this is set to zero (the default).
Returns
The approximate frequentist p-value
InferenceMLEorKMSummaryTable$clone()
The objects of this class are cloneable with this method.
Usage
InferenceMLEorKMSummaryTable$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
## ------------------------------------------------
## Method `InferenceMLEorKMSummaryTable$compute_estimate()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "continuous")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(c(4.71, 1.23, 4.78, 6.11, 5.95, 8.43))
seq_des_inf = InferenceContinOLS$new(seq_des)
seq_des_inf$compute_estimate()
Marginal vs. Conditional Estimand Switch
Description
Component scaffold providing set_estimand()/
get_estimand()/get_supported_estimands() for classes whose
reported treatment coefficient is conditional on a latent mixture
component – e.g. the interior beta submodel in zero/one-inflated beta
regression, or the count-process submodel in zero-augmented/hurdle
Poisson – rather than the unconditional response mean E[Y]. See
marginal_estimand_report.md for the full design discussion.
Mirrors the testing_type switch (InferenceAsympLik) on its
own, orthogonal axis: default "conditional" (today's behavior,
fully backward compatible – every class that does not compose this
component is implicitly conditional-only), with "marginal_mean_diff"
and "marginal_ratio" available to classes that declare support for
them via their own get_supported_estimands_impl() override (the
same override pattern already used for
get_supported_testing_types_impl()).
Scope note (2026-08-18): this component provides only the
get/set/supported-values switch and the cache-key helper. It does not
itself compute any marginal estimate – no class currently composes it.
Wiring a class's own compute_estimate() to consult
self$get_estimand() and, for a marginal estimand, call into a
family-specific model-implied mean function plus a shared
g-computation-average/delta-method-gradient helper, is
marginal_estimand_report.md → TODO-4/5/9 – deliberately deferred
until their target classes (still on the legacy deep-hierarchy ladder as
of this writing) migrate to the shallow hierarchy under
fix_inference_hierarchy.md's Full-Likelihood Estimators remainder.
compute_estimate() itself stays 100 percent class-owned either way
– this component never overrides or wraps it, so no
allowed_host_overrides declaration is needed.
Methods
Public methods
InferenceMarginalEstimand$set_estimand()
Sets the target estimand for this inference object.
Usage
InferenceMarginalEstimand$set_estimand(estimand)
Arguments
estimandOne of
get_supported_estimands(). Accepts the canonical values ("conditional","marginal_mean_diff","marginal_ratio") case-insensitively.
Details
If this object also composes LikelihoodTests (checked
via self$supports("likelihood_tests"), the sanctioned
capability query – see
marginal_estimand_report.md → TODO-6), switching to a
non-"conditional" estimand shrinks the set of supported
testing types to "wald" only. If the currently configured
testing_type is no longer in that shrunk set, this errors
loudly and leaves the estimand unchanged, rather than silently
leaving the object in an inconsistent state – the same guarantee
holds regardless of which of set_testing_type()/
set_estimand() is called first.
Returns
The inference object, invisibly.
InferenceMarginalEstimand$get_estimand()
Gets the current target estimand.
Usage
InferenceMarginalEstimand$get_estimand()
Returns
A character scalar, one of get_supported_estimands().
InferenceMarginalEstimand$get_supported_estimands()
Gets the estimands supported by this inference object.
Usage
InferenceMarginalEstimand$get_supported_estimands()
Returns
A character vector. Always includes "conditional".
Canonicalizes a requested estimand value, rejecting anything not in
the fixed set of recognized spellings – unlike
'get_supported_estimands_impl()' below (host-overridable, varies by
class), this recognizes syntax, not per-class support.
Default: every class implicitly supports only the conditional
estimand until it declares otherwise. Concrete classes override this
private method (via ‘define_inference_class()'’s 'overrides'
argument, the same pattern 'get_supported_testing_types_impl()'
already uses) once they wire a marginal mean function.
Cache-key fragment for the current estimand, generalizing
‘likelihood_test_delta_key()'’s testing_type/delta pattern to this
orthogonal axis. Any cache keyed partly by estimand should prefix or
combine with this so a cache entry built under one estimand is never
reused under another.
InferenceMarginalEstimand$clone()
The objects of this class are cloneable with this method.
Usage
InferenceMarginalEstimand$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Bootstrap-based Inference
Description
Abstract class for bootstrap-based inference.
The default m = NULL rule is a cheap deterministic
intermediate sequence: m \to \infty and m / n \to 0, as
required by the standard m-out-of-n bootstrap asymptotic setup
(Bickel, Gotze, and van Zwet; Bickel and Sakov). The exponent 0.7 is a
pragmatic interior point in (0, 1); it is not a silver-bullet
optimal choice. Use select_optimal_m_out_of_n_bootstrap() for
data-adaptive minimum-volatility selection.
The default m = NULL follows the intermediate-sequence
convention from the m-out-of-n bootstrap literature:
m \to \infty and m / n \to 0. The deterministic exponent
0.7 is a first-pass default; for unstable paths prefer the
minimum-volatility selector.
The NULL default is grounded in the standard m-out-of-n
asymptotic condition m \to \infty and m / n \to 0. The
minimum-volatility selector is available when a fixed deterministic
exponent is too brittle for a specific estimator/design path.
This implements the same minimum-volatility idea used in PTE's m-selection workflow: scan admissible intermediate sizes and choose a stable region of the target statistic rather than assuming one exponent is uniformly optimal.
The default b = NULL rule is a cheap deterministic
intermediate sequence: b \to \infty and b / n \to 0, as
required by the Politis, Romano, and Wolf subsampling framework. The
exponent 0.7 is a pragmatic interior point in (0, 1); it is not a
universal optimum. Use select_optimal_b_subsampling() for
data-adaptive minimum-volatility selection.
The default b = NULL follows the intermediate-sequence
convention from the Politis/Romano/Wolf subsampling literature:
b \to \infty and b / n \to 0. The deterministic exponent
0.7 is a first-pass default; for unstable paths prefer the
minimum-volatility selector.
The NULL default is grounded in the standard
Politis/Romano/Wolf asymptotic condition b \to \infty and
b / n \to 0. The minimum-volatility selector is available when a
fixed deterministic exponent is too brittle for a specific
estimator/design path.
This implements the same minimum-volatility idea used in PTE's m-selection workflow, applied to the PRW block/subsample-size choice: scan admissible intermediate sizes and choose a stable region of the target statistic rather than assuming one exponent is uniformly optimal.
Design-specific validity caveats for the nonparametric bootstrap
Nonparametric bootstrap methods are not supported for DesignSeqOneByOne
designs and their subclasses, except the concrete
DesignSeqOneByOneBernoulli class. The restriction includes m-out-of-n
bootstrap and subsampling methods registered under the same capability.
Calling restricted methods reports that the method is not supported.
For supported designs, the nonparametric bootstrap resamples experimental units with replacement from their
empirical distribution, carrying each unit's realized (x, w, y) into the
replicate, and recomputes the estimator. Its validity rests on the resampled units
being (approximately) iid draws from the design's unit-level superpopulation. Since
covariate-adaptive designs induce dependence among the assignments w_i (and
between w and X), the appropriate resampling unit and the fidelity with
which the design's dependence is replicated differ by design. In all cases below the
inference is asymptotic, never finite-sample exact (for exact finite-sample inference
under the design's actual randomization mechanism, use the randomization tests and
randomization confidence intervals where their assumptions hold). Calibration
depends on the design and estimator; omitted dependence does not in general
guarantee conservative inference.
DesignFixedBernoulliandDesignSeqOneByOneBernoulli-
Assignments are independent coin flips that do not use covariates or past assignments. If subjects and their potential outcomes arrive iid, with a noninformative sample size, observed rows are iid and ordinary row-level resampling has its usual asymptotic justification for regular estimators. Independent assignments alone do not establish iid rows under time trends, dependent recruitment, or informative stopping.
DesignFixediBCRDAssignment depends only on the treatment counts (completely randomized / without-replacement urn), inducing negative correlation among the
w_ithrough the fixed-margin constraint. Row-level iid resampling does not replicate this constraint: replicates have a random number of treated subjects. The extra variability isO(1/n), so the bootstrap is conservative by an asymptotically negligible amount.DesignFixedBlockingResampling is within-strata by default (
bootstrap_type = "within_blocks"), preserving stratum sizes; the exact within-stratum treatment/control split is not enforced in replicates, so the block-randomization variance reduction is partially unreplicated. Conservative, minor.bootstrap_type = "resample_blocks"instead resamples whole blocks, preserving within-block composition at the price of fewer resampling atoms.DesignFixedOptimalBlocksSame within-block resampling caveats as
DesignFixedBlocking, plus the blocks themselves are computed from the realized covariate sample: the block structure is a global function of the data that the bootstrap conditions on rather than re-derives. The justification for this conditioning is asymptotic: asngrows the blocking depends on the sample only through the (convergent) empirical distribution ofX, so between-block dependence vanishes. Conservative.DesignFixedClusterAssignment is at the cluster level and outcomes are correlated within clusters, so whole clusters are resampled with replacement. This is the correct exchangeable unit; with few clusters the bootstrap distribution rests on few resampling atoms and becomes unstable. Asymptotics are in the number of clusters, not the number of subjects.
DesignFixedBlockedClusterClusters are resampled within strata, matching both levels of the design's dependence (stratum and cluster). Sound, with the same small-sample caution: few clusters per stratum means few resampling atoms per stratum, and asymptotics are in the number of clusters.
DesignFixedGreedyDOptimal,DesignFixedGreedy,DesignFixedRerandomizationThe observed
wvector is one draw from a tightly constrained (optimized or acceptance-sampled) set of allocations. Resampled replicates carry per-row assignments whose recombinedwvector no longer satisfies the balance constraint, so the bootstrap reflects the variance of unconstrained assignment (cf. Li, Ding & Rubin 2018 for rerandomization). Conservative, moderate-to-large: the stronger the optimization, the greater the over-coverage.DesignFixedMatchingGreedyPairSwitchingThe greedy switching search only ever flips assignments within binary-match pairs, so every pair has exactly one treated subject; the bootstrap resamples intact pairs, preserving the within-pair anticorrelation. Remaining caveats: the pairing is a global function of the sample (conditioned on, justified asymptotically as for the matched designs below), and the greedy choice of which pair member is treated couples the pairs, which resampling does not replicate — the residual effect errs conservative.
DesignFixedBinaryMatchMatched pairs are resampled intact, preserving the within-pair anticorrelation of
wand the pair-level variance reduction. The pairing itself is a global function of the covariate sample (an Abadie & Imbens 2008-type concern): pairs are exchangeable but not exactly independent. Validity is asymptotic — asngrows the pairing depends on the sample only through the empirical distribution ofXand between-pair dependence vanishes — and the bootstrap conditions on the realized match structure.DesignFixedFactorialRow-level resampling does not replicate the balanced allocation across factor combinations. Conservative, minor.
DesignFixedCustomWarning: iid row-level resampling is used because the package has no knowledge of the user-supplied assignment mechanism. If that mechanism balances on covariates, the bootstrap is likely conservative; if it induces clustering or other positive dependence, the bootstrap may not even be valid (anti-conservative). Use the randomization-based inference, which draws from the actual custom mechanism, whenever possible.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap
Methods
Public methods
-
InferenceNonParamBootstrap$get_supported_bootstrap_pval_types() -
InferenceNonParamBootstrap$get_supported_bootstrap_ci_types() -
InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T() -
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval() -
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval() -
InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap() -
InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T() -
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval() -
InferenceNonParamBootstrap$compute_subsampling_confidence_interval() -
InferenceNonParamBootstrap$compute_subsampling_sensitivity() -
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T() -
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval() -
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()
Returns the type values
compute_bootstrap_two_sided_pval() accepts.
Usage
InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()
InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()
Returns the type values
compute_bootstrap_confidence_interval() accepts.
Usage
InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()
InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()
Creates the m-out-of-n bootstrap distribution of the treatment-effect estimate.
Usage
InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T( B = 501, m = NULL, show_progress = TRUE, debug = FALSE, bootstrap_type = NULL, scaling = "sqrt_n", center = "full_estimate" )
Arguments
BNumber of resamples. Default 501.
mNumber of exchangeable resampling units drawn with replacement. If
NULL(default), use the deterministic intermediate-size rulefloor(n_units^0.7), wheren_unitsis the number of exchangeable units used by the design (observations, clusters, pairs, or matched sets). The resolved value must satisfymax(5, p_eff + 2) <= m <= floor(n_units / 2).show_progressA flag indicating whether a progress bar should be displayed.
debugIf
TRUE, return distribution diagnostics in addition to the resampled estimates.bootstrap_typeOptional empirical-resampling scheme. See
approximate_bootstrap_distribution_beta_hat_T().scalingScaling sequence for centered m-out-of-n pivots. The default
"sqrt_n"usessqrt(m)for the m-sample distribution and converts back to the full-sample scale usingsqrt(n_units).centerCentering convention for diagnostics and cache keys.
Returns
A numeric vector of bootstrap estimates, or when
debug = TRUE, a diagnostic list.
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()
Computes a centered m-out-of-n bootstrap two-sided p-value.
Usage
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval( delta = 0, B = 501, m = NULL, type = "centered", show_progress = TRUE, min_number_usable_samples = 5L, bootstrap_type = NULL, scaling = "sqrt_n" )
Arguments
deltaNull treatment effect. Default 0.
BNumber of resamples. Default 501.
mNumber of exchangeable units drawn with replacement. If
NULL(default), usefloor(n_units^0.7)subject to the validation bounds documented forapproximate_m_out_of_n_bootstrap_distribution_beta_hat_T().typeP-value type. Currently only
"centered"is supported.show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite resampled estimates required after filtering. Default 5.
bootstrap_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered m-out-of-n pivots.
Returns
A numeric two-sided p-value, or NA_real_ if the path is
non-estimable.
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()
Computes a basic m-out-of-n bootstrap confidence interval.
Usage
InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval( alpha = 0.05, B = 501, m = NULL, type = "basic", show_progress = TRUE, min_number_usable_samples = 5L, bootstrap_type = NULL, scaling = "sqrt_n" )
Arguments
alphaSignificance level. Default 0.05.
BNumber of resamples. Default 501.
mNumber of exchangeable units drawn with replacement. If
NULL(default), usefloor(n_units^0.7)subject to the validation bounds documented forapproximate_m_out_of_n_bootstrap_distribution_beta_hat_T().typeConfidence-interval type. Currently only
"basic"is supported.show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite resampled estimates required after filtering. Default 5.
bootstrap_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered m-out-of-n pivots.
Returns
A length-2 numeric confidence interval, or
c(NA_real_, NA_real_) if the path is non-estimable.
InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
Selects an m-out-of-n bootstrap size by minimum volatility.
Usage
InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap( B = 251, alpha = 0.05, m_pow_of_n_grid = seq(0.5, 0.9, by = 0.05), m_grid = NULL, objective = "ci_width", target = "ci", volatility_window = 3L, bootstrap_type = NULL, scaling = "sqrt_n", show_progress = TRUE, min_finite_fraction = 0.8 )
Arguments
BNumber of resamples per candidate size. Default 251.
alphaSignificance level for interval-width objectives.
m_pow_of_n_gridCandidate exponent grid used when
m_grid = NULL. Defaults toseq(0.5, 0.9, by = 0.05).m_gridOptional explicit integer candidate sizes.
objectiveSelection objective. Currently
"ci_width".targetTarget summary. Currently
"ci".volatility_windowRolling window size used to measure local volatility across candidate sizes.
bootstrap_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered m-out-of-n pivots.
show_progressA flag indicating whether a progress bar should be displayed.
min_finite_fractionMinimum finite-resample fraction required for a candidate size to be eligible.
Returns
An EDIMOutOfNBootstrapMSelection list with the selected
m, mapped exponent, candidate table, status, and reason.
InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()
Creates the Politis/Romano/Wolf subsampling distribution of the treatment-effect estimate.
Usage
InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T( B = 501, b = NULL, show_progress = TRUE, debug = FALSE, subsampling_type = NULL, scaling = "sqrt_n", center = "full_estimate" )
Arguments
BNumber of subsamples. Default 501.
bNumber of exchangeable units drawn without replacement. If
NULL(default), use the deterministic intermediate-size rulefloor(n_units^0.7), wheren_unitsis the number of exchangeable units used by the design (observations, clusters, pairs, or matched sets). The resolved value must satisfymax(5, p_eff + 2) <= b <= floor(n_units / 2).show_progressA flag indicating whether a progress bar should be displayed.
debugIf
TRUE, return distribution diagnostics in addition to the subsampled estimates.subsampling_typeOptional empirical-resampling scheme. See
approximate_bootstrap_distribution_beta_hat_T().scalingScaling sequence for centered subsampling pivots. The default
"sqrt_n"usessqrt(b)for the subsample distribution and converts back to the full-sample scale usingsqrt(n_units).centerCentering convention for diagnostics and cache keys.
Returns
A numeric vector of subsampled estimates, or when
debug = TRUE, a diagnostic list.
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()
Computes a centered PRW subsampling two-sided p-value.
Usage
InferenceNonParamBootstrap$compute_subsampling_two_sided_pval( delta = 0, B = 501, b = NULL, type = "centered", show_progress = TRUE, min_number_usable_samples = 5L, subsampling_type = NULL, scaling = "sqrt_n" )
Arguments
deltaNull treatment effect. Default 0.
BNumber of subsamples. Default 501.
bNumber of exchangeable units drawn without replacement. If
NULL(default), usefloor(n_units^0.7)subject to the validation bounds documented forapproximate_subsampling_distribution_beta_hat_T().typeP-value type. Currently only
"centered"is supported.show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite subsampled estimates required after filtering. Default 5.
subsampling_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered subsampling pivots.
Returns
A numeric two-sided p-value, or NA_real_ if the path is
non-estimable.
InferenceNonParamBootstrap$compute_subsampling_confidence_interval()
Computes a basic PRW subsampling confidence interval.
Usage
InferenceNonParamBootstrap$compute_subsampling_confidence_interval( alpha = 0.05, B = 501, b = NULL, type = "basic", show_progress = TRUE, min_number_usable_samples = 5L, subsampling_type = NULL, scaling = "sqrt_n" )
Arguments
alphaSignificance level. Default 0.05.
BNumber of subsamples. Default 501.
bNumber of exchangeable units drawn without replacement. If
NULL(default), usefloor(n_units^0.7)subject to the validation bounds documented forapproximate_subsampling_distribution_beta_hat_T().typeConfidence-interval type. Currently only
"basic"is supported.show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite subsampled estimates required after filtering. Default 5.
subsampling_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered subsampling pivots.
Returns
A length-2 numeric confidence interval, or
c(NA_real_, NA_real_) if the path is non-estimable.
InferenceNonParamBootstrap$select_optimal_b_subsampling()
Selects a PRW subsampling size by minimum volatility.
Usage
InferenceNonParamBootstrap$select_optimal_b_subsampling( B = 251, alpha = 0.05, b_pow_of_n_grid = seq(0.5, 0.9, by = 0.05), b_grid = NULL, objective = "ci_width", target = "ci", volatility_window = 3L, subsampling_type = NULL, scaling = "sqrt_n", show_progress = TRUE, min_finite_fraction = 0.8 )
Arguments
BNumber of subsamples per candidate size. Default 251.
alphaSignificance level for interval-width objectives.
b_pow_of_n_gridCandidate exponent grid used when
b_grid = NULL. Defaults toseq(0.5, 0.9, by = 0.05).b_gridOptional explicit integer candidate sizes.
objectiveSelection objective. Currently
"ci_width".targetTarget summary. Currently
"ci".volatility_windowRolling window size used to measure local volatility across candidate sizes.
subsampling_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered subsampling pivots.
show_progressA flag indicating whether a progress bar should be displayed.
min_finite_fractionMinimum finite-subsample fraction required for a candidate size to be eligible.
Returns
An EDISubsamplingBSelection list with the selected
b, mapped exponent, candidate table, status, and reason.
InferenceNonParamBootstrap$compute_subsampling_sensitivity()
Computes PRW subsampling sensitivity over candidate sizes.
Usage
InferenceNonParamBootstrap$compute_subsampling_sensitivity( B = 251, alpha = 0.05, b_pow_of_n_grid = seq(0.5, 0.9, by = 0.05), b_grid = NULL, objective = "ci_width", target = "ci", volatility_window = 3L, subsampling_type = NULL, scaling = "sqrt_n", show_progress = TRUE, min_finite_fraction = 0 )
Arguments
BNumber of subsamples per candidate size. Default 251.
alphaSignificance level for interval-width objectives.
b_pow_of_n_gridCandidate exponent grid used when
b_grid = NULL. Defaults toseq(0.5, 0.9, by = 0.05).b_gridOptional explicit integer candidate sizes.
objectiveSelection objective. Currently
"ci_width".targetTarget summary. Currently
"ci".volatility_windowRolling window size used to measure local volatility across candidate sizes.
subsampling_typeOptional empirical-resampling scheme.
scalingScaling sequence for centered subsampling pivots.
show_progressA flag indicating whether a progress bar should be displayed.
min_finite_fractionMinimum finite-subsample fraction required for a candidate size to be eligible. Defaults to 0 for sensitivity scans.
Returns
An EDISubsamplingSensitivity list containing the candidate
grid table without selecting a final b.
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()
Creates the bootstrap distribution of the estimate for the treatment effect. The resampling unit is design-specific (rows, within-strata rows, matched pairs plus reservoir, or clusters); see the class-level section Design-specific validity caveats for the nonparametric bootstrap for the conservativeness and asymptotics of each concrete design.
Usage
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T( B = 501, show_progress = TRUE, debug = FALSE, bootstrap_type = NULL )
Arguments
BNumber of bootstrap samples. The default is 501.
show_progressA flag indicating whether a progress bar should be displayed.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.bootstrap_typeOptional bootstrap-resampling scheme. Legal public values are:
NULLUse the design's default row-resampling bootstrap. For ordinary non-blocking designs this is the usual subject-level resample-with-replacement bootstrap. For certain blocking designs,
NULLmaps to the same behavior as"within_blocks"."within_blocks"Only legal for blocking-style designs that support block-aware bootstrap resampling:
DesignFixedBlocking,DesignFixedOptimalBlocks,DesignSeqOneByOneSPBR, andDesignFixedBlockedCluster. Resamples observational units within each observed block/stratum. For blocked cluster designs this means resampling clusters within strata."resample_blocks"Only legal for the same blocking-style designs as
"within_blocks". Resamples entire observed blocks/strata with replacement rather than resampling units within each block.
Any non-
NULLvalue is rejected for designs outside that blocking family.
Returns
When debug = FALSE (default), a numeric vector of length B
containing the bootstrap estimates. When debug = TRUE, a list with: values,
errors (list of character vectors, one per iteration), warnings (list of
character vectors, one per iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()
Computes a bootstrap-based two-sided p-value for the treatment effect. Validity is asymptotic and design-dependent; for most covariate-adaptive designs the p-value errs conservative. See the class-level section Design-specific validity caveats for the nonparametric bootstrap.
Usage
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval( delta = 0, B = 501, type = NULL, na.rm = FALSE, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
deltaNull hypothesis value. Default 0.
BNumber of bootstrap samples. Default 501.
typeBootstrap p-value type. Supported values are
"percentile"(default),"symmetric","studentized","bootstrap-t", and"bca"."percentile": shifts the bootstrap distribution to be centred atdeltaand counts the two-tail proportion (Hall 1992)."symmetric": uses|T^* - \bar{T}^*| \ge |t_{\rm obs} - \delta|for a symmetric one-sample test; recommended by Hall & Wilson (1991) when the null distribution may be skewed. This pooled-tail test is offered only as a p-value here, not as a confidence-intervaltypeincompute_bootstrap_confidence_interval: pooling both tails via|\cdot|improves testing power (Hall & Wilson's original use case), but inverting it unstudentized would add no value as an interval. The unstudentized pivot is not asymptotically pivotal, so the resulting interval would have the same first-orderO(n^\{-1/2\})coverage error as"percentile"/"basic", while forcing symmetric bounds around a possibly skewed bootstrap distribution — strictly worse than"percentile"/"basic"for shape-adaptivity, and strictly worse than"symmetric-percentile-t"for accuracy, since studentizing (not the absolute-value pooling) is what buys theO(n^\{-1\})improvement. The CI-worthy symmetric variant is therefore"symmetric-percentile-t"(studentized pivot), not a plain"symmetric"CI type."studentized"/"bootstrap-t": pivots by the per-replicate standard error, giving O(n^{-1}) error versus O(n^{-1/2}) for the percentile method (Hall 1992; Davidson & MacKinnon 1999)."bca": bias-corrected and accelerated p-value via closed-form CI inversion using the jackknife acceleration and bias-correction constants; second-order accurate (Efron 1987; Efron & Tibshirani 1993).na.rmRemove non-finite bootstrap replicates. Default FALSE.
show_progressA flag indicating whether a progress bar should be displayed.
min_number_usable_samplesMinimum number of finite bootstrap samples required after filtering. Default 5. Must be less than or equal to
B.
Returns
A bootstrap two-sided p-value.
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()
Computes a bootstrap-based confidence interval. Coverage is asymptotic and design-dependent; for most covariate-adaptive designs the interval errs conservative (over-coverage). See the class-level section Design-specific validity caveats for the nonparametric bootstrap.
Usage
InferenceNonParamBootstrap$compute_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
alphaThe confidence level 1 -
alpha. Default 0.05.BNumber of bootstrap samples. Default 501.
typeBootstrap CI type. Supported values are
"percentile","basic","studentized","bootstrap-t","symmetric-percentile-t","bca","prepivoted","double-bootstrap","calibrated", and"smoothed". There is no plain"symmetric"CI type (contrast with the"symmetric"p-value type incompute_bootstrap_two_sided_pval): inverting the unstudentized Hall & Wilson pooled-tail statistic would add no value as an interval, since it is not asymptotically pivotal and so has the same first-orderO(n^\{-1/2\})coverage error as"percentile"/"basic", while forcing symmetric bounds around a possibly skewed bootstrap distribution — strictly worse than"percentile"/"basic"for shape-adaptivity, and strictly worse than"symmetric-percentile-t"for accuracy, since studentizing (not the absolute-value pooling) is what buys theO(n^\{-1\})improvement."symmetric-percentile-t"is the CI-worthy symmetric variant.na.rmRemove non-finite bootstrap replicates. Default TRUE. Non-finite replicates are always removed internally.
show_progressShow progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required after filtering. Default 5. Must be less than or equal to
B.
Returns
A bootstrap confidence interval.
InferenceNonParamBootstrap$clone()
The objects of this class are cloneable with this method.
Usage
InferenceNonParamBootstrap$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Bickel, P. J., Gotze, F., and van Zwet, W. R. (1997). Resampling fewer than n observations: gains, losses, and remedies for losses. Statistica Sinica.
Bickel, P. J. and Sakov, A. (2008). On the choice of m in the m out of n bootstrap. The Annals of Statistics.
Politis, D. N., Romano, J. P., and Wolf, M. (1999). Subsampling. Springer.
Adjacent Category Logit Regression Inference for Ordinal Responses
Description
Fits an adjacent-category logit regression for ordinal responses (via
fast_adjacent_category_logit_cpp — see that page for the full
model, an alternative ordinal parameterization to the cumulative-logit
proportional-odds model) using the treatment indicator and, optionally, all
recorded covariates as predictors. This is a full-likelihood class
(likelihood_tier = "full") supporting score, gradient, and
likelihood-ratio tests, plus parametric likelihood-ratio bootstrap
calibration, in addition to Wald and resampling-based inference.
Bayesian-bootstrap inference is temporarily unavailable because the current
non-uniform weighted hook fits a cumulative-logit surrogate rather than the
adjacent-category likelihood. It will remain disabled until the native
weighted adjacent-category backend described in the package implementation
plan lands.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalAdjCatLogitRegr
Methods
Public methods
-
InferenceOrdinalAdjCatLogitRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalAdjCatLogitRegr$supports_rand_pval_for_incidence() -
InferenceOrdinalAdjCatLogitRegr$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalAdjCatLogitRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalAdjCatLogitRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalAdjCatLogitRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalAdjCatLogitRegr$supports_rand_pval_for_incidence()
InferenceOrdinalAdjCatLogitRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalAdjCatLogitRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalAdjCatLogitRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalAdjCatLogitRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalAdjCatLogitRegr$new(seq_des)
inf$compute_estimate()
Cauchit Regression Inference for Ordinal Responses
Description
Cauchit-link cumulative-odds ordinal regression: P(Y \le k \mid w, x) =
F_{\mathrm{Cauchy}}(\alpha_k - \beta_T w - \beta_X^\top x), where
F_{\mathrm{Cauchy}} is the standard Cauchy CDF, \alpha_k are
category-specific cutpoints, and \beta_T is the treatment log-odds
coefficient on the cauchit scale (proportional-odds-style shift common to all
categories). Fit by maximum likelihood. The heavy-tailed Cauchy link is
markedly less sensitive to outlying/extreme response categories than the
logit or probit link, at the cost of a less familiar effect-size
interpretation. likelihood_tier = "full": exposes likelihood-ratio,
score, gradient, and parametric-likelihood-bootstrap inference in addition to
the Wald/asymptotic and Bayesian-bootstrap paths.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalCauchitRegr
Methods
Public methods
-
InferenceOrdinalCauchitRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalCauchitRegr$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalCauchitRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalCauchitRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalCauchitRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalCauchitRegr$supports_rand_pval_for_incidence()
InferenceOrdinalCauchitRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalCauchitRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalCauchitRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalCauchitRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley. Ch. 3-4 (cumulative link models).
See Also
https://en.wikipedia.org/wiki/Ordinal_regression, https://www.statsmodels.org/stable/discretemod.html for an analogous Python cumulative-link API.
Cumulative Cloglog Inference for Ordinal Responses
Description
Complementary log-log cumulative-odds ordinal regression:
P(Y \le k \mid w, x) = 1 - \exp\{-\exp(\alpha_k - \beta_T w -
\beta_X^\top x)\}, where \alpha_k are category-specific cutpoints and
\beta_T is the treatment coefficient on the cloglog scale. Fit by
maximum likelihood. The cloglog link is asymmetric (unlike logit/probit) and
is the natural ordinal generalization of a proportional-hazards/grouped
survival-time model, so it is preferred when the underlying process is
plausibly a discretized time-to-event or extreme-value mechanism.
likelihood_tier = "full": exposes likelihood-ratio, score, gradient,
and parametric-likelihood-bootstrap inference in addition to Wald/asymptotic
and Bayesian-bootstrap paths.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalCloglogRegr
Methods
Public methods
-
InferenceOrdinalCloglogRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalCloglogRegr$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalCloglogRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalCloglogRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalCloglogRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalCloglogRegr$supports_rand_pval_for_incidence()
InferenceOrdinalCloglogRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalCloglogRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalCloglogRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalCloglogRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley. Ch. 3-4 (cumulative link models); McCullagh, P. (1980). "Regression Models for Ordinal Data." JRSS-B, 42(2), 109-142.
See Also
https://en.wikipedia.org/wiki/Ordinal_regression
Continuation Ratio Regression Inference for Ordinal Responses
Description
Fits a conditional (stratified) continuation-ratio logit model for ordinal
responses: for cut j = 1, \dots, K-1, among subjects who have
reached at least category j,
\log\frac{\Pr(Y_i > j \mid Y_i
\ge j)}{\Pr(Y_i = j \mid Y_i \ge j)} = \alpha_j + \beta_T W_i + X_i^\top
\gamma,
a discrete-time-hazard-model analog for ordinal data, with a
treatment coefficient \beta_T constrained equal across all cuts.
\exp(\hat\beta_T) is the common "continue vs. stop here" odds ratio:
a positive \beta_T means treatment pushes subjects toward higher
categories of Y, matching the sign convention of every other ordinal
estimator in the package.
Fitting proceeds by expand_continuation_ratio_data_cpp's
stacked-binary expansion followed by conditional logistic regression on
the expanded data. likelihood_tier = "full": likelihood-ratio,
score, gradient, and Wald tests are all available when the model
converges, plus parametric-likelihood-bootstrap calibration of the
likelihood-ratio test. Validity requires the continuation-ratio
proportionality assumption (a common \beta_T across all K-1
cuts).
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalContRatioRegr
Methods
Public methods
-
InferenceOrdinalContRatioRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalContRatioRegr$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalContRatioRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalContRatioRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalContRatioRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalContRatioRegr$supports_rand_pval_for_incidence()
InferenceOrdinalContRatioRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalContRatioRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalContRatioRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalContRatioRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley, for the continuation-ratio model family.
See Also
InferenceOrdinalStereotypeLogitRegr
and
InferenceOrdinalKKCondAdjCatLogitRegr
for related ordinal-logit expansions. See also:
Ordinal
regression (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalContRatioRegr$new(seq_des)
inf$compute_estimate()
G-Computation Mean-Difference Inference for Ordinal Responses
Description
Fits a proportional-odds working model for an ordinal outcome using
treatment and, optionally, all recorded covariates
(fast_ordinal_regression_with_var_cpp), then estimates the
marginal difference in expected ordinal category score by
G-computation — see gcomp_ordinal_proportional_odds_post_fit_cpp
for the exact standardization formula (mean1 - mean0). Standard
errors are obtained by the delta method: a central finite-difference
gradient of the mean-difference functional with respect to the fitted
[\alpha, \beta] parameters, propagated through the model's fitted
variance-covariance matrix, \widehat{\mathrm{Var}}(\widehat{\mathrm{md}}) =
\nabla^\top \widehat{\mathrm{Var}}(\hat\theta) \nabla. If that delta-method
standard error is unavailable or non-finite, the Wald-style methods
($compute_asymp_confidence_interval(), $compute_asymp_two_sided_pval(),
$compute_wald_confidence_interval(), $compute_wald_two_sided_pval())
all silently fall back to a nonparametric bootstrap interval/p-value instead
(with a warning), rather than returning NA.
Super class
Inference -> InferenceOrdinalGCompMeanDiff
Methods
Public methods
-
InferenceOrdinalGCompMeanDiff$compute_estimate_with_bootstrap_weights() -
InferenceOrdinalGCompMeanDiff$compute_asymp_confidence_interval() -
InferenceOrdinalGCompMeanDiff$compute_asymp_two_sided_pval() -
InferenceOrdinalGCompMeanDiff$compute_wald_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalGCompMeanDiff$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize the ordinal g-computation (G-Comp) inference object for a completed design with an ordinal, uncensored response.
Usage
InferenceOrdinalGCompMeanDiff$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
DesignSeqOneByOneobject with an ordinal response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceOrdinalGCompMeanDiff$compute_estimate()
Computes the G-computation standardized mean-difference treatment-effect estimate (see class documentation for the full proportional-odds-based standardization).
Usage
InferenceOrdinalGCompMeanDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceOrdinalGCompMeanDiff$compute_estimate_with_bootstrap_weights()
Recomputes the G-computation mean-difference estimate under
subject/block bootstrap weights (via
fast_ordinal_regression_weighted_cpp plus
gcomp_ordinal_proportional_odds_post_fit_cpp), used by
the Bayesian bootstrap and related weighted-resampling machinery; see
InferenceNonParamBootstrap.
Runs side-effect free: the ordinary (unweighted) cached fit, warm-start
state, and rank-reduced column selection are saved before the weighted
refit and restored afterward (on.exit), so a weighted bootstrap
replicate cannot corrupt the class's own point estimate or subsequent
fits.
Usage
InferenceOrdinalGCompMeanDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsRow weights for the bootstrap sample.
estimate_onlyIf TRUE, skip variance calculations.
InferenceOrdinalGCompMeanDiff$compute_asymp_confidence_interval()
Computes a 1-\alpha confidence interval for the
G-Comp mean difference using the delta-method standard error (see
class documentation), or falls back (with a warning) to a
nonparametric bootstrap interval if that standard error is
unavailable. Identical to $compute_wald_confidence_interval().
Usage
InferenceOrdinalGCompMeanDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level (default 0.05).
InferenceOrdinalGCompMeanDiff$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\mathrm{md} = \code{delta} using the delta-method standard error (see
class documentation), or falls back (with a warning) to a
nonparametric bootstrap p-value if that standard error is
unavailable. Identical to $compute_wald_two_sided_pval().
Usage
InferenceOrdinalGCompMeanDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect (default 0).
InferenceOrdinalGCompMeanDiff$compute_wald_confidence_interval()
Identical to $compute_asymp_confidence_interval()
(both compute the same delta-method-based Wald interval, with the
same bootstrap fallback); provided as an explicit alias for callers
that want to name the Wald method directly rather than via the
generic "asymptotic" dispatch.
Usage
InferenceOrdinalGCompMeanDiff$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level (default 0.05).
InferenceOrdinalGCompMeanDiff$compute_wald_two_sided_pval()
Identical to $compute_asymp_two_sided_pval() (both
compute the same delta-method-based Wald p-value, with the same
bootstrap fallback); provided as an explicit alias for callers that
want to name the Wald method directly rather than via the generic
"asymptotic" dispatch.
Usage
InferenceOrdinalGCompMeanDiff$compute_wald_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect (default 0).
InferenceOrdinalGCompMeanDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalGCompMeanDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalGCompMeanDiff$new(seq_des)
inf$compute_estimate()
Jonckheere-Terpstra (JT) Test for Ordinal Responses
Description
Two-arm Jonckheere-Terpstra (JT) rank test for an ordinal response — for two
groups, this reduces to the Mann-Whitney U statistic. The point
estimate is the stochastic superiority probability, centered at 0
under the null: \hat\beta_T = \widehat{\Pr}(Y_T > Y_C) +
\tfrac{1}{2}\widehat{\Pr}(Y_T = Y_C) - \tfrac{1}{2}, computed from category
counts as U/(n_T n_C) - 1/2. Asymptotic inference
($compute_asymp_confidence_interval(), $compute_asymp_two_sided_pval())
uses the classical null variance of the Mann-Whitney U statistic,
\mathrm{Var}(U) = n_T n_C (n_T+n_C+1)/12 (no tie correction), matching
clinfun::jonckheere.test()'s normal approximation. This class also
provides an exact, permutation-distribution-based two-sided p-value
via $compute_exact_two_sided_pval_for_treatment_effect()
(exact_jonckheere_terpstra_pval_cpp), which does not rely on the
normal approximation.
Super class
Inference -> InferenceOrdinalJonckheereTerpstraTest
Methods
Public methods
-
InferenceOrdinalJonckheereTerpstraTest$compute_estimate_with_bootstrap_weights() -
InferenceOrdinalJonckheereTerpstraTest$compute_exact_two_sided_pval_for_treatment_effect() -
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_confidence_interval() -
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalJonckheereTerpstraTest$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize the JT test object for a completed design with an ordinal, uncensored response.
Usage
InferenceOrdinalJonckheereTerpstraTest$new( des_obj, model_formula = NULL, verbose = FALSE )
Arguments
des_objA completed
DesignSeqOneByOneobject.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress.
InferenceOrdinalJonckheereTerpstraTest$compute_estimate()
Returns the estimated treatment effect: the stochastic
superiority measure \widehat{\Pr}(Y_T > Y_C) +
\tfrac12\widehat{\Pr}(Y_T=Y_C) - \tfrac12, computed from the
Mann-Whitney U statistic (see class documentation).
Usage
InferenceOrdinalJonckheereTerpstraTest$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceOrdinalJonckheereTerpstraTest$compute_estimate_with_bootstrap_weights()
Recomputes the JT superiority estimate under subject/block
bootstrap weights: the weighted version of the same stochastic
superiority quantity, \sum_{i,j} w_i w_j\left(\mathbb{1}[y_{T,i} >
y_{C,j}] + \tfrac12\mathbb{1}[y_{T,i}=y_{C,j}]\right) \big/ \sum_{i,j}
w_i w_j - \tfrac12, used by the Bayesian bootstrap and related
weighted-resampling machinery. Always leaves the standard error
unavailable (NA) regardless of estimate_only — this
weighted path never computes the null-variance approximation.
Usage
InferenceOrdinalJonckheereTerpstraTest$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject/block level.
estimate_onlyPresent for interface parity; this method never computes variance components regardless of its value.
InferenceOrdinalJonckheereTerpstraTest$compute_exact_two_sided_pval_for_treatment_effect()
Returns the exact, permutation-distribution-based
two-sided p-value (exact_jonckheere_terpstra_pval_cpp) — unlike
$compute_asymp_two_sided_pval(), this does not rely on the
normal approximation to the Mann-Whitney U null distribution.
Usage
InferenceOrdinalJonckheereTerpstraTest$compute_exact_two_sided_pval_for_treatment_effect( )
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_confidence_interval()
Computes the asymptotic normal confidence interval, using the
same Mann-Whitney U null-variance approximation
(n_T n_C(n_T+n_C+1)/12, no tie correction) as
clinfun::jonckheere.test(); see class documentation.
Usage
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaThe significance level (default 0.05).
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_two_sided_pval()
Computes the asymptotic normal two-sided p-value, using the
same Z-approximation as clinfun::jonckheere.test(); see
class documentation and $compute_exact_two_sided_pval_for_treatment_effect()
for the exact (non-approximate) alternative.
Usage
InferenceOrdinalJonckheereTerpstraTest$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null treatment effect (default 0).
InferenceOrdinalJonckheereTerpstraTest$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalJonckheereTerpstraTest$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Jonckheere, A. R. (1954). "A Distribution-Free k-Sample Test Against Ordered Alternatives." Biometrika, 41(1-2), 133-145, doi:10.1093/biomet/41.1-2.133; Terpstra, T. J. (1952). "The Asymptotic Normality and Consistency of Kendall's Test Against Trend, When Ties Are Present in One Ranking." Indagationes Mathematicae, 14, 327-333.
Examples
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneBernoulli$new(n = nrow(x_dat), response_type = "ordinal",
verbose = FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(as.integer(c(1, 2, 2, 3, 3, 4, 4, 5)))
infer <- InferenceOrdinalJonckheereTerpstraTest$
new(seq_des, verbose = FALSE)
infer
Ordinal KK CLMM (Proportional Odds / logit link)
Description
Cumulative-link random-intercept mixed model for ordinal responses under a
KK matching-on-the-fly design, using the logit link (proportional odds):
\mathrm{logit}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top
\gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is
subject i's matched-pair group id (reservoir subjects get singleton
groups). \exp(\hat\beta_T) is the (conditional-on-b_g) treatment
odds ratio. See
InferenceAbstractKKOrdinalCLMM
for the shared model-fitting, caching, and likelihood-tier contract common
to all four link-function siblings
(...Probit,
...Cauchit,
...Cloglog); this class
supplies only the link-function choice (private$clmm_link() ==
"logit").
Super classes
Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMM
Methods
Public methods
+ inherited public methods from InferenceAbstractKKOrdinalCLMM
InferenceAbstractKKOrdinalCLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_estimate()InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKOrdinalCLMM$compute_jackknife_bias_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_std_error()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_sensitivity()InferenceAbstractKKOrdinalCLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$get_mod()InferenceAbstractKKOrdinalCLMM$get_summary()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_testing_types()InferenceAbstractKKOrdinalCLMM$select_optimal_b_subsampling()InferenceAbstractKKOrdinalCLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKOrdinalCLMM$set_testing_type()InferenceAbstractKKOrdinalCLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKCLMM$new()
Initialize the logit-link ordinal KK CLMM subclass; see the
shared ordinal mixed-model contract in
InferenceAbstractKKOrdinalCLMM.
Usage
InferenceOrdinalKKCLMM$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment.
use_rcppUse internal Rcpp implementation (default
TRUE).verbosePrint messages?
smart_cold_start_defaultUse smart cold start values?
InferenceOrdinalKKCLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKCLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMM$new(seq_des)
inf$compute_estimate()
Ordinal KK CLMM (Cauchit link)
Description
Cumulative-link random-intercept mixed model for ordinal responses under a
KK matching-on-the-fly design, using the cauchit (inverse-Cauchy-CDF) link:
\tan(\pi (P(Y_i \le k) - 1/2)) = \alpha_k - (\beta_T W_i + X_i^\top
\gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is
subject i's matched-pair group id. The cauchit link's heavy-tailed
latent distribution makes it more robust than logit/probit to a small
number of subjects near the response's extreme categories. See
InferenceAbstractKKOrdinalCLMM
for the shared model-fitting, caching, and likelihood-tier contract common
to all four link-function siblings; this class supplies only the
link-function choice (private$clmm_link() == "cauchit").
Super classes
Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMMCauchit
Methods
Public methods
+ inherited public methods from InferenceAbstractKKOrdinalCLMM
InferenceAbstractKKOrdinalCLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_estimate()InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKOrdinalCLMM$compute_jackknife_bias_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_std_error()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_sensitivity()InferenceAbstractKKOrdinalCLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$get_mod()InferenceAbstractKKOrdinalCLMM$get_summary()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_testing_types()InferenceAbstractKKOrdinalCLMM$select_optimal_b_subsampling()InferenceAbstractKKOrdinalCLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKOrdinalCLMM$set_testing_type()InferenceAbstractKKOrdinalCLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKCLMMCauchit$new()
Initialize the cauchit-link ordinal KK CLMM subclass; see the
shared ordinal mixed-model contract in
InferenceAbstractKKOrdinalCLMM.
Usage
InferenceOrdinalKKCLMMCauchit$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment.
use_rcppUse internal Rcpp implementation (default
TRUE).verbosePrint messages?
smart_cold_start_defaultUse smart cold start values?
InferenceOrdinalKKCLMMCauchit$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKCLMMCauchit$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMMCauchit$new(seq_des)
inf$compute_estimate()
Ordinal KK CLMM (Complementary log-log link)
Description
Cumulative-link random-intercept mixed model for ordinal responses under a
KK matching-on-the-fly design, using the complementary log-log link:
\log(-\log(1 - P(Y_i \le k))) = \alpha_k - (\beta_T W_i + X_i^\top
\gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where g(i) is
subject i's matched-pair group id. Unlike the symmetric logit/probit/
cauchit links, the cloglog link is asymmetric, making it appropriate when
the ordinal categories arise from an underlying continuous-time
proportional-hazards process discretized into intervals. See
InferenceAbstractKKOrdinalCLMM
for the shared model-fitting, caching, and likelihood-tier contract common
to all four link-function siblings; this class supplies only the
link-function choice (private$clmm_link() == "cloglog").
Super classes
Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMMCloglog
Methods
Public methods
+ inherited public methods from InferenceAbstractKKOrdinalCLMM
InferenceAbstractKKOrdinalCLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_estimate()InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKOrdinalCLMM$compute_jackknife_bias_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_std_error()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_sensitivity()InferenceAbstractKKOrdinalCLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$get_mod()InferenceAbstractKKOrdinalCLMM$get_summary()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_testing_types()InferenceAbstractKKOrdinalCLMM$select_optimal_b_subsampling()InferenceAbstractKKOrdinalCLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKOrdinalCLMM$set_testing_type()InferenceAbstractKKOrdinalCLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKCLMMCloglog$new()
Initialize the cloglog-link ordinal KK CLMM subclass; see the
shared ordinal mixed-model contract in
InferenceAbstractKKOrdinalCLMM.
Usage
InferenceOrdinalKKCLMMCloglog$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment.
use_rcppUse internal Rcpp implementation (default
TRUE).verbosePrint messages?
smart_cold_start_defaultUse smart cold start values?
InferenceOrdinalKKCLMMCloglog$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKCLMMCloglog$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMMCloglog$new(seq_des)
inf$compute_estimate()
Ordinal KK CLMM (Probit link)
Description
Cumulative-link random-intercept mixed model for ordinal responses under a
KK matching-on-the-fly design, using the probit link:
\Phi^{-1}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma) -
b_{g(i)}, b_g \sim N(0, \sigma_b^2), where \Phi is the standard
normal CDF and g(i) is subject i's matched-pair group id. Unlike
the logit-link sibling, \hat\beta_T here is not an odds-ratio scale
parameter; it is the treatment's effect on the latent standard-normal index
underlying the ordinal categories. See
InferenceAbstractKKOrdinalCLMM
for the shared model-fitting, caching, and likelihood-tier contract common
to all four link-function siblings; this class supplies only the
link-function choice (private$clmm_link() == "probit").
Super classes
Inference -> InferenceAbstractKKOrdinalCLMM -> InferenceOrdinalKKCLMMProbit
Methods
Public methods
+ inherited public methods from InferenceAbstractKKOrdinalCLMM
InferenceAbstractKKOrdinalCLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKOrdinalCLMM$compute_asymp_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_asymp_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_estimate()InferenceAbstractKKOrdinalCLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKOrdinalCLMM$compute_jackknife_bias_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_estimate()InferenceAbstractKKOrdinalCLMM$compute_jackknife_std_error()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_rand_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_rand_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_subsampling_sensitivity()InferenceAbstractKKOrdinalCLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKOrdinalCLMM$compute_wald_confidence_interval()InferenceAbstractKKOrdinalCLMM$compute_wald_two_sided_pval()InferenceAbstractKKOrdinalCLMM$get_mod()InferenceAbstractKKOrdinalCLMM$get_summary()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKOrdinalCLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKOrdinalCLMM$get_supported_testing_types()InferenceAbstractKKOrdinalCLMM$select_optimal_b_subsampling()InferenceAbstractKKOrdinalCLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKOrdinalCLMM$set_testing_type()InferenceAbstractKKOrdinalCLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKCLMMProbit$new()
Initialize the probit-link ordinal KK CLMM subclass; see the
shared ordinal mixed-model contract in
InferenceAbstractKKOrdinalCLMM.
Usage
InferenceOrdinalKKCLMMProbit$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject.model_formulaOptional formula for covariate adjustment.
use_rcppUse internal Rcpp implementation (default
TRUE).verbosePrint messages?
smart_cold_start_defaultUse smart cold start values?
InferenceOrdinalKKCLMMProbit$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKCLMMProbit$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCLMMProbit$new(seq_des)
inf$compute_estimate()
Adjacent Category Logit Inference for KK Matching-on-the-fly Designs
Description
Fits a conditional (stratified) adjacent-category logit model for ordinal responses under a KK matching-on-the-fly design:
\log\frac{\Pr(Y_i = j+1 \mid Y_i \in \{j, j+1\})}{\Pr(Y_i = j \mid Y_i
\in \{j, j+1\})} = \alpha_j + \beta_T W_i + X_i^\top \gamma,
for adjacent
category comparisons j = 1, \dots, K-1, with cut-specific intercepts
\alpha_j and a treatment coefficient \beta_T constrained equal
across all cuts (the parallel/proportional adjacent-category assumption).
\exp(\hat\beta_T) is the common adjacent-category odds ratio. Fitting
proceeds by expand_adjacent_category_data_cpp's stacked-binary
expansion (each subject contributes a 0/1 row per adjacent cut they border,
stratified by matched pair) followed by conditional logistic regression on
the expanded data — the matched-pair identity becomes the conditioning
stratum, so the pair's shared nuisance intercept is conditioned out exactly
as in a single binary conditional-logit KK model, and reservoir (unmatched)
subjects each form their own singleton stratum. likelihood_tier =
"partial" (a conditional/partial likelihood, matched-set effects are
profiled out rather than estimated); supports_likelihood_tests() is
hard FALSE — only Wald inference is exposed, not likelihood-ratio,
score, or gradient tests. Validity requires the adjacent-category
proportionality assumption (a common \beta_T across all K-1
cuts) in addition to the usual conditional-logit exchangeability-within-strata
assumption induced by the KK design.
Super class
Inference -> InferenceOrdinalKKCondAdjCatLogitRegr
Methods
Public methods
-
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate_with_bootstrap_weights() -
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_confidence_interval() -
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKCondAdjCatLogitRegr$new()
Initialize inference for the conditional adjacent-category
logit model \log(\Pr(Y_i = j+1 \mid Y_i \in \{j,j+1\}) / \Pr(Y_i =
j \mid Y_i \in \{j,j+1\})) = \alpha_j + \beta_T W_i + X_i^\top \gamma
and prepare KK matched-pair structure for the stratified conditional-logit
fit. Does not fit the model; the fit is deferred to the first call to
compute_estimate() or a method that requires it.
Usage
InferenceOrdinalKKCondAdjCatLogitRegr$new( des_obj, verbose = FALSE, harden = TRUE, model_formula = NULL, smart_cold_start_default = NULL )
Arguments
des_objA completed KK
DesignSeqOneByOneobject with an ordinal response.verboseFlag for progress messages.
hardenWhether to apply robustness measures.
model_formulaOptional formula for covariate adjustment.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate()
Fits the conditional adjacent-category logit model via
stacked-binary expansion (expand_adjacent_category_data_cpp)
plus conditional logistic regression, and returns the shared
log-odds-ratio estimate \hat\beta_T.
Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf
TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate_with_bootstrap_weights()
Recomputes the treatment estimate under subject/block-level
bootstrap weights (Bayesian-bootstrap or nonparametric-bootstrap draw
weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()). When
weights are effectively constant, this collapses to the unweighted
compute_estimate() call. Otherwise, rather than refitting the
full expanded conditional-logit model under weights, it calls
weighted_ordinal_bootstrap_surrogate_fit() — a fast weighted
ordinal-logistic surrogate fit on the raw (unexpanded) design matrix —
as an approximation to the weighted adjacent-category likelihood; this
trades exact reweighted refitting for speed across many bootstrap
replicates. No standard error is computed (s_beta_hat_T is
always NA); the surrogate returns NA if the fit fails.
Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyIf
TRUE, compute only the weighted point estimate.
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_confidence_interval()
Wald confidence interval for the shared adjacent-category
log-odds-ratio \beta_T, using the conditional-logit model's
standard error; see InferenceAsymp
for the shared Wald contract. Fits the model first if not already
cached.
Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_two_sided_pval()
Return the adjacent-category conditional-logit asymptotic
p-value for the treatment coefficient, using the shared Wald semantics
documented in InferenceAsymp.
Usage
InferenceOrdinalKKCondAdjCatLogitRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull hypothesis treatment effect.
InferenceOrdinalKKCondAdjCatLogitRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKCondAdjCatLogitRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Agresti, A. (2010). Analysis of Ordinal Categorical Data (2nd ed.). Wiley, for the adjacent-category logit model family; Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for.
See Also
InferenceOrdinalAdjCatLogitRegr
for the non-KK analog. See also:
Ordinal regression
(Wikipedia).
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKCondAdjCatLogitRegr$new(seq_des)
inf$compute_estimate()
GEE Inference for KK Designs with Ordinal Response
Description
Fits a proportional-odds local-odds-ratio Generalized Estimating
Equations model, via multgee::ordLORgee, for ordinal responses under a
KK matching-on-the-fly design, using the treatment indicator and,
optionally, all recorded covariates as predictors. Each GEE cluster is
either a matched pair (2 members) or a reservoir singleton (1 member) — GEE
is used here purely to fit one marginal cumulative-logit model jointly
across matched-pair and reservoir subjects while accounting for the
within-pair correlation the matching induces, not as a
longitudinal/repeated-measures tool. Unlike the other Inference*KKGEE
classes in this family (continuous/count/incidence/proportion, which use an
internal Rcpp solver or geepack::geeglm with an exchangeable working
correlation), this ordinal class always requires the multgee package
and has no use_rcpp option. The raw multgee treatment coefficient
is negated when reported so that, consistently with EDI's other ordinal
estimators, a positive estimate means movement toward higher response
categories. Inference is quasi-likelihood/
estimating-equation based (likelihood_tier = "quasi"): standard
errors are GEE sandwich (robust) standard errors, not model-likelihood-based.
Bayesian-bootstrap inference is temporarily unavailable because
multgee::ordLORgee does not accept the non-uniform observation
weights needed to refit the same clustered estimator. It will remain
disabled until the weighted ordinal-GEE implementation planned for v1.1.0
is complete.
Super class
Inference -> InferenceOrdinalKKGEE
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKGEE$new()
Initialize KK ordinal GEE inference, validate the ordinal
matched/reservoir design, and prepare the multgee::ordLORgee
proportional-odds local-odds-ratio GEE fitting machinery used by
InferenceOrdinalKKGEE. Requires
the multgee package; errors at construction if it is not installed.
Usage
InferenceOrdinalKKGEE$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with an ordinal response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceOrdinalKKGEE$compute_estimate_with_bootstrap_weights()
Recomputes the KK ordinal treatment estimate under
subject/block bootstrap weights, used by the Bayesian bootstrap and
related weighted-resampling machinery. If the supplied weights are all
(numerically) equal, this short-circuits to the unweighted
$compute_estimate(estimate_only = TRUE) (the multgee
proportional-odds GEE fit) rather than refitting. Otherwise, since
multgee::ordLORgee does not support observation weights, this
falls back to a different, approximating model: a plain
(non-GEE, no matched-pair clustering) weighted proportional-odds
ordinal logistic regression via
fast_ordinal_regression_weighted_cpp, treating the
coefficient on the first predictor column as the treatment effect.
This always leaves the standard error and degrees of freedom
unavailable (s_beta_hat_T = NA, df = Inf) regardless of
estimate_only, since it is a point-estimate-only fallback path.
Usage
InferenceOrdinalKKGEE$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyIf
TRUE, compute only the weighted point estimate. Has no effect on the weighted (non-uniform-weight) fallback path, which never computes a standard error regardless.
InferenceOrdinalKKGEE$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKGEE$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Touloumis, A. (2015). "R Package multgee: A Generalized Estimating Equations Solver for Multinomial Responses." Journal of Statistical Software, 64(8), 1-14, doi:10.18637/jss.v064.i08, for the local-odds-ratio GEE solver used here; Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the underlying GEE estimating-equation framework.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKGEE$new(seq_des)
inf$compute_estimate()
GLMM Inference for KK Designs with Ordinal Response
Description
Fits a cumulative-logit random-intercept mixed model (proportional odds) for
ordinal responses under a KK matching-on-the-fly design:
\mathrm{logit}(P(Y_i \le k \mid w_i, x_i, b_{g(i)})) = \alpha_k - (\beta_T
w_i + x_i^\top \gamma) - b_{g(i)}, for cutpoints \alpha_1 < \cdots <
\alpha_{K-1}, treatment indicator w_i, covariates x_i, and a
matched-pair random intercept b_g \sim N(0, \sigma_b^2) that is
integrated out of the marginal likelihood (either by adaptive Gauss-Hermite
quadrature when use_rcpp = TRUE, the default; see
fast_ordinal_glmm_cpp for the quadrature order and optimizer
details, or by glmmTMB's Laplace approximation when use_rcpp =
FALSE). g(i) is subject i's matched-pair group id; reservoir
(unmatched) subjects each get their own singleton group, contributing no
within-group correlation but still entering the joint likelihood. The
treatment coefficient \beta_T is a log-odds-ratio: \exp(\beta_T)
is the (conditional-on-b_g) odds ratio of being at or above any given
response category. likelihood_tier = "full": likelihood-ratio, score,
Wald, and gradient tests are available when use_rcpp = TRUE and the
fit converges; use_rcpp = FALSE disables likelihood-test support
(private$supports_likelihood_tests() returns FALSE) because
glmmTMB's Laplace-approximate likelihood is not wired into this
package's score/gradient/LR machinery. Validity requires the random-intercept
structure to correctly capture the design's matching dependence, proportional
odds (the treatment/covariate effect is constant across cutpoints), and
correct specification of the fixed-effects formula.
This differs from the GEE sibling
InferenceOrdinalKKGEE (documented
above) in estimand and inference basis: the GLMM's \beta_T is a
subject-specific (conditional) log-odds-ratio with model-likelihood-based
inference, while the GEE's is a population-averaged (marginal) log-odds-ratio
with sandwich-based inference; the two need not numerically agree even on the
same data, and the correct choice depends on whether a
subject-specific/conditional or population-averaged/marginal treatment
effect is of interest.
Super class
Inference -> InferenceOrdinalKKGLMM
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalKKGLMM$new()
Initialize inference for the cumulative-logit random-intercept
mixed model \mathrm{logit}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i +
X_i^\top \gamma) - b_{g(i)}, b_g \sim N(0, \sigma_b^2), where
g(i) is subject i's matched-pair group id (reservoir subjects
get singleton groups). Does not fit the model; the fit is deferred to
the first call to compute_estimate() or a method that requires it.
Usage
InferenceOrdinalKKGLMM$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with an ordinal response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.use_rcppLogical. If
TRUE(default), use internal Rcpp.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceOrdinalKKGLMM$compute_estimate()
Fits the cumulative-logit random-intercept mixed model by
(adaptive-Gauss-Hermite- or Laplace-)approximate maximum likelihood and
returns \hat\beta_T, the estimated treatment log-odds-ratio,
conditional on the matched-pair random intercept. Caches the fitted
model object, full parameter vector, and (when estimate_only =
FALSE) the standard error and degrees of freedom for reuse by
compute_asymp_confidence_interval(),
compute_asymp_two_sided_pval(), and likelihood-test methods; a
fit that fails the kernel's projected-gradient convergence check,
produces non-finite parameters, reaches the upper random-effect variance
boundary, exceeds private$max_abs_reasonable_coef, or lacks a
finite positive treatment-coefficient variance is cached as nonestimable
rather than returned. A valid near-zero random-effect variance boundary
is accepted using conditional fixed-effect information. The native
optimizer retains multistart L-BFGS for basin selection and, only when
its finite selected point fails the projected-score tolerance, applies
a damped-Newton polish using the numerical Hessian. The polished point
is retained only if it remains finite and does not increase the
negative log-likelihood; at a valid lower variance boundary, the
KKT-satisfied variance coordinate is excluded from that Newton system.
Usage
InferenceOrdinalKKGLMM$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf
TRUE, skip standard-error/variance-component computation and cache only the point estimate; used by randomization and bootstrap resampling paths where only\hat\beta_Tis needed per replicate.
InferenceOrdinalKKGLMM$compute_asymp_confidence_interval()
Wald confidence interval for \beta_T using the fitted
model's standard error and degrees of freedom; see
InferenceAsymp for the shared
\hat\beta_T \pm t_{\alpha/2, df} \cdot \widehat{se}(\hat\beta_T)
(or z-based when df = Inf) contract. Fits the model first if not
already cached.
Usage
InferenceOrdinalKKGLMM$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferenceOrdinalKKGLMM$compute_asymp_two_sided_pval()
Two-sided Wald test of H_0: \beta_T = \code{delta}
against H_1: \beta_T \ne \code{delta}, using the fitted model's
standard error and degrees of freedom; see
InferenceAsymp for the shared
t/z test contract. Fits the model first if not already
cached.
Usage
InferenceOrdinalKKGLMM$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaTreatment log-odds-ratio value under the null hypothesis.
InferenceOrdinalKKGLMM$compute_estimate_with_bootstrap_weights()
Refits the mixed model with subject/block-level weights
applied to each row's contribution to the marginal likelihood
(Bayesian-bootstrap or nonparametric-bootstrap draw weights, expanded
from subject/block level to individual rows via
private$expand_subject_or_block_weights_to_row_weights()) and
returns the reweighted estimate \hat\beta_T^{(w)}. Uses
fast_ordinal_regression_weighted_cpp — an ordinary
(non-mixed-effects) weighted cumulative-logit fit, not a reweighted
GLMM refit — as a fast approximation to the weighted marginal
likelihood; this trades exact random-effects refitting for speed across
many bootstrap replicates. When weights are effectively constant, this
collapses to the unweighted compute_estimate() call (returns
df = Inf to signal a degenerate/skipped bootstrap replicate
rather than refitting). Rows with non-finite or non-positive weight, or
non-finite response, are dropped from the weighted fit; if no rows
remain, the estimate is NA.
Usage
InferenceOrdinalKKGLMM$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyIf
TRUE, compute only the weighted point estimate.
InferenceOrdinalKKGLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalKKGLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Hedeker, D., and Gibbons, R. D. (1994). "A Random-Effects Ordinal Regression Model for Multilevel Analysis." Biometrics, 50(4), 933-944, doi:10.2307/2533433, for the random-effects cumulative-logit model; Pinheiro, J. C., and Bates, D. M. (1995). "Approximations to the Log-Likelihood Function in the Nonlinear Mixed-Effects Model." Journal of Computational and Graphical Statistics, 4(1), 12-35, doi:10.1080/10618600.1995.10474663, for the adaptive Gauss-Hermite quadrature approximation used to integrate out the random intercept.
See Also
Comparable Python API: statsmodels MixedLM (continuous analog; no ordinal-GLMM in statsmodels). See also: Ordinal regression and Mixed model (Wikipedia).
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalKKGLMM$new(seq_des)
inf$compute_estimate()
Ordered Probit Regression Inference for Ordinal Responses
Description
Fits a cumulative-probit ("ordered probit") model for ordinal responses:
\Phi^{-1}(P(Y_i \le k)) = \alpha_k - (\beta_T W_i + X_i^\top \gamma),
for cutpoints \alpha_1 < \cdots < \alpha_{K-1}, where \Phi is
the standard normal CDF, W_i is the treatment indicator, and
X_i are optional recorded covariates, by maximum likelihood
(fast_ordinal_probit_regression_cpp/
fast_ordinal_probit_regression_with_var_cpp). As with binary
probit regression, \hat\beta_T is not an odds-ratio-scale parameter:
it is the treatment's effect on the latent standard-normal index
underlying the ordinal categories. likelihood_tier = "full":
likelihood-ratio, score, gradient, and Wald tests are all available when
the model converges, plus parametric-likelihood-bootstrap calibration of
the likelihood-ratio test. Validity requires the proportional/parallel
cutpoints assumption (a single \beta_T shared across all cutpoints)
in addition to the usual latent-normal-index assumption.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalOrderedProbitRegr
Methods
Public methods
-
InferenceOrdinalOrderedProbitRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalOrderedProbitRegr$supports_rand_pval_for_incidence() -
InferenceOrdinalOrderedProbitRegr$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalOrderedProbitRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalOrderedProbitRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalOrderedProbitRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalOrderedProbitRegr$supports_rand_pval_for_incidence()
InferenceOrdinalOrderedProbitRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalOrderedProbitRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalOrderedProbitRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalOrderedProbitRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
McCullagh, P. (1980). "Regression Models for Ordinal Data." Journal of the Royal Statistical Society, Series B, 42(2), 109-142, doi:10.1111/j.2517-6161.1980.tb01109.x, for the cumulative-link ordinal model family this class's probit link instantiates.
See Also
InferenceOrdinalCauchitRegr,
InferenceOrdinalCloglogRegr
for other cumulative-link function choices on the same ordinal model
family. See also:
Ordinal
regression and Probit
model (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalOrderedProbitRegr$new(seq_des)
inf$compute_estimate()
Paired Sign Test Inference for KK Designs with Ordinal Response
Description
Fits the classical paired sign test for ordinal responses under a KK
matching-on-the-fly design. For each matched pair i with treated
member response Y_{i,T} and control member response Y_{i,C},
only the sign of the within-pair difference Y_{i,T} - Y_{i,C} is
used; tied pairs (Y_{i,T} = Y_{i,C}) are dropped from the effective
sample. The estimand is \theta = P(Y_T > Y_C \mid \text{pair
untied}), and the reported treatment effect is \hat\beta_T = \hat p -
0.5, where \hat p is the sample proportion of untied pairs favoring
treatment; \beta_T = 0 corresponds to \theta = 0.5 (no
directional preference). The standard error is the usual binomial-proportion
formula \sqrt{\hat p (1 - \hat p) / n_{\text{eff}}}, where
n_{\text{eff}} is the number of untied pairs.
Reservoir (unmatched) subjects are not included — this is a purely
within-pair test, unlike the IVWC-style classes elsewhere in the KK family
that combine matched-pair and reservoir information.
likelihood_tier = "none" (supports_likelihood_tests() is hard
FALSE): only Wald inference on the proportion scale is exposed.
Bootstrap and jackknife are deliberately unsupported and throw explicit
errors (see approximate_bootstrap_distribution_beta_hat_T() and
approximate_jackknife_distribution_beta_hat_T()), since subject-level
resampling or deletion would violate the matched-pair design's dependence
structure; randomization inference (compute_rand_two_sided_pval())
remains available since it permutes treatment assignment within the design's
own randomization mechanism rather than resampling subjects. Requires a KK
matching-on-the-fly design (DesignSeqOneByOneKK14/KK21) or
DesignFixedBinaryMatch; a design with no discordant (untied) pairs is
cached as nonestimable for the standard error (point estimate 0) or
fully nonestimable, per harden.
Super class
Inference -> InferenceOrdinalPairedSignTest
Methods
Public methods
-
InferenceOrdinalPairedSignTest$compute_estimate_with_bootstrap_weights() -
InferenceOrdinalPairedSignTest$compute_asymp_confidence_interval() -
InferenceOrdinalPairedSignTest$compute_asymp_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalPairedSignTest$new()
Uses the shared randomization-test two-sided p-value
contract; see InferenceRand.
Pinned from plain InferenceRand (not InferenceRandCI)
per the established ordinal-class precedent – Zhang dispatch is
incidence-only.
Initialize inference for the paired sign test on within-pair
response differences Y_{i,T} - Y_{i,C}; see
InferenceOrdinalPairedSignTest
for the model form. Requires a KK matching-on-the-fly design
(DesignSeqOneByOneKK14/KK21) or
DesignFixedBinaryMatch. Does not compute the sign-test statistic;
that is deferred to the first call to compute_estimate() or a
method that requires it.
Usage
InferenceOrdinalPairedSignTest$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed KK matching-on-the-fly design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
rNumber of randomization draws.
deltaNull treatment effect.
transform_responsesOptional response transformation.
na.rmWhether to drop non-finite draws.
show_progressWhether to show a progress bar.
permutationsOptional pre-computed permutations.
zero_one_logit_clampClamp for logit transforms.
InferenceOrdinalPairedSignTest$compute_estimate()
Computes the pair-sign counts (pos/neg, ties
dropped) from the design's matched-pair structure and returns
\hat\beta_T = \hat p - 0.5, where \hat p is the proportion
of untied pairs favoring treatment. If every pair is tied, the
estimate is 0 (no directional preference) and the fit is
cached as standard-error-nonestimable (or fully nonestimable when
harden = FALSE).
Usage
InferenceOrdinalPairedSignTest$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip the standard-error computation and cache only the point estimate.
InferenceOrdinalPairedSignTest$compute_estimate_with_bootstrap_weights()
Recomputes \hat\beta_T under subject/block-level
bootstrap weights (Bayesian-bootstrap draw weights, expanded to row
level via
private$expand_subject_or_block_weights_to_row_weights()): for
each matched pair, a weighted vote is cast toward whichever member has
the higher response, using the mean bootstrap weight of the pair's two
rows; \hat\beta_T^{(w)} is the weighted proportion of
treatment-favoring pairs minus 0.5. No standard error is
computed (s_beta_hat_T is always NA). Pairs with no
discordant (untied) weighted votes are cached as nonestimable.
Usage
InferenceOrdinalPairedSignTest$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyIf
TRUE, compute only the weighted point estimate.
InferenceOrdinalPairedSignTest$compute_asymp_confidence_interval()
Wald confidence interval for \beta_T = \theta - 0.5
(equivalently, for \theta = P(Y_T > Y_C \mid \text{pair
untied})), using the binomial-proportion standard error; see
InferenceAsymp for the shared Wald
contract. Fits (computes the pair-sign counts) first if not already
cached.
Usage
InferenceOrdinalPairedSignTest$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferenceOrdinalPairedSignTest$compute_asymp_two_sided_pval()
Two-sided Wald test of H_0: \theta = 0.5 (equal
chance of favoring treatment vs. control among untied pairs) against
H_1: \theta \ne 0.5, using the binomial-proportion standard
error; see InferenceAsymp for the
shared Wald contract. Only delta = 0 is supported (the sign
test's null is fixed at no directional preference; a non-zero
delta throws).
Usage
InferenceOrdinalPairedSignTest$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null value for
\beta_T; must be0.
InferenceOrdinalPairedSignTest$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalPairedSignTest$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Dixon, W. J., and Mood, A. M. (1946). "The Statistical Sign Test." Journal of the American Statistical Association, 41(236), 557-566, doi:10.2307/2280577, for the classical paired sign test; Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for.
See Also
Sign test (Wikipedia).
Examples
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneKK21$new(n = nrow(x_dat), response_type = "ordinal",
verbose = FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(as.integer(c(1, 2, 2, 3, 3, 4, 4, 5)))
infer <- InferenceOrdinalPairedSignTest$
new(seq_des, verbose = FALSE)
infer
Partial Proportional-Odds Regression Inference for Ordinal Responses
Description
Fits a partial proportional-odds cumulative-logit model
for an ordinal response: a subset of covariates named in nonparallel
are allowed a separate coefficient at each cumulative threshold (relaxing the
proportional-odds/parallel-lines assumption for exactly those covariates),
while every other covariate — including the treatment indicator,
which is always fit as a parallel (proportional) term regardless of
nonparallel — keeps one shared coefficient across all thresholds. The
reported treatment effect is therefore always a single proportional
(threshold-invariant) log-odds shift, even when other covariates' effects
are allowed to vary by threshold. When nonparallel is empty, fitting
uses this package's fast Rcpp full-proportional-odds solver
(fast_ordinal_regression_with_var_cpp); otherwise it falls back,
in order, to VGAM::vglm(family = VGAM::cumulative(parallel = ...)),
ordinal::clm(nominal = ...), and (only when nonparallel is
empty and the earlier fast/VGAM/clm attempts failed) MASS::polr. Each
fallback requires its corresponding package to be installed; unavailable
packages are silently skipped in favor of the next fallback.
Super class
Inference -> InferenceOrdinalPartialProportionalOddsRegr
Methods
Public methods
-
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate() -
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate_with_bootstrap_weights() -
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_confidence_interval() -
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_two_sided_pval() -
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_confidence_interval() -
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_two_sided_pval() -
InferenceOrdinalPartialProportionalOddsRegr$benchmark_asymp_two_sided_pval_breakdown()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalPartialProportionalOddsRegr$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize partial proportional-odds ordinal regression inference for a completed design with an ordinal, uncensored response.
Usage
InferenceOrdinalPartialProportionalOddsRegr$new( des_obj, verbose = FALSE, harden = TRUE, model_formula = NULL, nonparallel = character(0), smart_cold_start_default = NULL )
Arguments
des_objA completed
DesignSeqOneByOneobject with an ordinal response.verboseWhether to print progress messages.
hardenWhether to apply robustness measures.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.nonparallelNames of covariates (not including
"treatment", which is always fit as a parallel/proportional term) allowed a separate coefficient at each cumulative threshold, relaxing the proportional-odds assumption for those covariates specifically.smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate()
Retrieves the estimated (always-parallel) treatment log-odds shift from the partial proportional-odds fit (see class documentation for the fitting backend cascade).
Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate( estimate_only = FALSE )
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
The estimated treatment effect.
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate_with_bootstrap_weights()
Recomputes the partial-proportional-odds treatment estimate
under subject/block bootstrap weights, used by the Bayesian bootstrap
and related weighted-resampling machinery. If the weights are
effectively constant, short-circuits to the unweighted
$compute_estimate(estimate_only = TRUE). Otherwise refits with
weights via the same backend cascade as the unweighted fit
(VGAM/ordinal/MASS::polr, each weighted), and if
all of those fail, falls back further to a plain weighted
binary-logistic surrogate fit
(weighted_ordinal_bootstrap_surrogate_fit(..., method =
"logistic")) that does not model the ordinal structure at all. Never
computes a standard error on any weighted path (s_beta_hat_T is
always NA), regardless of estimate_only.
Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyPresent for interface parity; this method never computes variance components regardless of its value.
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_confidence_interval()
Computes a Wald-style confidence interval for the
treatment log-odds shift, using the model-based standard error
from whichever backend (fast Rcpp solver, VGAM,
ordinal, or MASS::polr) successfully fit the
unweighted model (see class documentation). If that standard
error is unavailable (NA, non-finite, or 0) — e.g. because
the fit succeeded via a fallback path that doesn't report one —
the interval is explicitly marked non-estimable
(c(NA, NA)) when private$harden is TRUE, or
raises an error otherwise, rather than silently returning a
misleading result. Identical to $compute_wald_confidence_interval().
Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaSignificance level for the interval.
Returns
A confidence interval for the treatment effect.
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_two_sided_pval()
Computes a Wald-style two-sided p-value testing
H_0: \beta_T = \code{delta}, using the same model-based
standard error as $compute_asymp_confidence_interval(); if
unavailable, marked non-estimable (NA) or an error is
raised, per private$harden — see that method's
documentation. Identical to $compute_wald_two_sided_pval().
Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_asymp_two_sided_pval( delta = 0 )
Arguments
deltaNull treatment effect to test.
Returns
A two-sided p-value.
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_confidence_interval()
Identical to $compute_asymp_confidence_interval();
provided as an explicit alias for callers that want to name the
Wald method directly rather than via the generic "asymptotic"
dispatch.
Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_confidence_interval( alpha = 0.05 )
Arguments
alphaSignificance level for the interval.
Returns
A confidence interval for the treatment effect.
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_two_sided_pval()
Identical to $compute_asymp_two_sided_pval();
provided as an explicit alias for callers that want to name the
Wald method directly rather than via the generic "asymptotic"
dispatch.
Usage
InferenceOrdinalPartialProportionalOddsRegr$compute_wald_two_sided_pval( delta = 0 )
Arguments
deltaNull treatment effect to test.
Returns
A two-sided p-value.
InferenceOrdinalPartialProportionalOddsRegr$benchmark_asymp_two_sided_pval_breakdown()
Diagnostic helper for performance investigation: runs the
same computation as $compute_asymp_two_sided_pval() (fit the
partial proportional-odds model requiring a standard error, cache the
estimate/SE/df, compute the two-sided Wald p-value) but separately
times each of the three stages — model fit, cache materialization, and
final p-value arithmetic — via proc.time(). If the fit fails or
has no usable standard error, returns immediately with only
fit_time populated and every other timing/result field
NA.
Usage
InferenceOrdinalPartialProportionalOddsRegr$benchmark_asymp_two_sided_pval_breakdown( delta = 0 )
Arguments
deltaNull treatment effect to test.
Returns
A named list: fit_time, cache_time,
pval_math_time, total_time (all in seconds), pval,
beta_hat_T, and s_beta_hat_T.
InferenceOrdinalPartialProportionalOddsRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalPartialProportionalOddsRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Peterson, B., and Harrell, F. E. (1990). "Partial Proportional
Odds Models for Ordinal Response Variables." Journal of the Royal
Statistical Society, Series C (Applied Statistics), 39(2), 205-217,
doi:10.2307/2347760, for the partial (non-parallel-covariate)
proportional-odds model fit here. McCullagh, P. (1980). "Regression
Models for Ordinal Data." Journal of the Royal Statistical
Society, Series B, 42(2), 109-142,
doi:10.1111/j.2517-6161.1980.tb01109.x, for the full proportional-odds
model this generalizes (see
InferenceOrdinalPropOddsRegr).
Proportional Odds Regression Inference for Ordinal Responses
Description
Fits a proportional-odds (cumulative-logit) regression, via
fast_ordinal_regression_with_var_cpp (see that page for the
full model), for ordinal responses using the treatment indicator and,
optionally, all recorded covariates as predictors. This is a full-likelihood
class (likelihood_tier = "full") supporting score, gradient, and
likelihood-ratio tests, plus parametric likelihood-ratio bootstrap
calibration, in addition to Wald and resampling-based inference.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalPropOddsRegr
Methods
Public methods
-
InferenceOrdinalPropOddsRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalPropOddsRegr$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalPropOddsRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalPropOddsRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalPropOddsRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalPropOddsRegr$supports_rand_pval_for_incidence()
InferenceOrdinalPropOddsRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalPropOddsRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalPropOddsRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalPropOddsRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
McCullagh, P. (1980). "Regression Models for Ordinal Data." Journal of the Royal Statistical Society, Series B, 42(2), 109-142, doi:10.1111/j.2517-6161.1980.tb01109.x, for the proportional-odds cumulative-logit model fit here.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'ordinal')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(sample(1:4, 10, replace = TRUE))
inf = InferenceOrdinalPropOddsRegr$new(seq_des)
inf$compute_estimate()
inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)
Ridit Analysis for Ordinal Responses
Description
Performs Ridit analysis (Relative to an Identified Distribution unit) for
comparing two groups on an ordinal scale. Every subject's category k is
converted to a ridit score — its relative rank position within the
reference distribution's empirical CDF: r_k = F_{\mathrm{ref}}(k{-}1)
+ \tfrac12 f_{\mathrm{ref}}(k), where F_{\mathrm{ref}} and
f_{\mathrm{ref}} are the reference group's empirical cumulative and
point probabilities. The reference distribution — controlled by the
reference constructor argument — may be the control arm
(default), the treatment arm, or the pooled sample. The
treatment effect is the mean ridit score among treated subjects minus
0.5 (the value it would take under the null of no group difference,
since a group's own ridit scores against itself as reference always average
to 0.5); this mean ridit score also has a direct interpretation as (an
estimate of) the probability that a randomly selected treated subject's
outcome exceeds a randomly selected reference-distribution subject's outcome
(a Mann-Whitney-type stochastic superiority probability), similar in spirit
to InferenceOrdinalJonckheereTerpstraTest's
superiority measure but referenced against a chosen distribution rather than
always symmetric between the two arms. Standard errors and p-values come
from fast_ridit_analysis_cpp's asymptotic formula, not a resampling
approximation.
Super class
Inference -> InferenceOrdinalRidit
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalRidit$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a Ridit analysis inference object for a completed design with an ordinal, uncensored response.
Usage
InferenceOrdinalRidit$new( des_obj, model_formula = NULL, reference = "control", verbose = FALSE, max_resample_attempts = 50L )
Arguments
des_objA DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.referenceThe group to use as the "Identified Distribution" (reference). Must be one of "control", "treatment", or "pooled". Default is "control".
verboseA flag indicating whether messages should be displayed.
max_resample_attemptsMaximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening. If all attempts fail the replicate is recorded as
NA, silently reducing the effectiveB. Must be a positive integer. Default50L.
InferenceOrdinalRidit$compute_estimate()
Returns the estimated treatment effect: the mean ridit score among treated subjects minus 0.5 (see class documentation for the full ridit-score definition and reference-distribution choice).
Usage
InferenceOrdinalRidit$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
The numeric estimate.
InferenceOrdinalRidit$compute_estimate_with_bootstrap_weights()
Recomputes the ridit treatment estimate under subject/block
bootstrap weights: the reference distribution's category proportions
and every subject's ridit score are recomputed using the weights, then
the weighted mean ridit score among treated subjects (minus 0.5) is
returned. Used by the Bayesian bootstrap and related
weighted-resampling machinery. Always leaves the standard error and
degrees of freedom unavailable (NA) regardless of
estimate_only — this weighted path never computes the
asymptotic variance.
Usage
InferenceOrdinalRidit$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyPresent for interface parity; this method never computes variance components regardless of its value.
InferenceOrdinalRidit$get_mean_ridit_treatment()
Returns the mean ridit score among treated subjects (not
centered — this is the raw mean, unlike $compute_estimate()
which subtracts 0.5).
Usage
InferenceOrdinalRidit$get_mean_ridit_treatment()
Returns
The numeric Mean Ridit.
InferenceOrdinalRidit$get_ridit_scores()
Returns each subject's individual ridit score (see class documentation for the ridit-score formula), in subject order.
Usage
InferenceOrdinalRidit$get_ridit_scores()
Returns
A numeric vector of scores.
InferenceOrdinalRidit$compute_asymp_confidence_interval()
Computes the asymptotic confidence interval for the
treatment effect (mean ridit - 0.5), using
fast_ridit_analysis_cpp's asymptotic standard error.
Usage
InferenceOrdinalRidit$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level.
Returns
A numeric vector of length 2.
InferenceOrdinalRidit$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\text{mean ridit} - 0.5 = \code{delta} (i.e. delta = 0 tests
the null of no group difference, mean ridit = 0.5), using
fast_ridit_analysis_cpp's asymptotic standard error.
Usage
InferenceOrdinalRidit$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null value (centered at 0, so delta=0 means Ridit=0.5).
Returns
The p-value.
InferenceOrdinalRidit$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalRidit$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Bross, I. D. J. (1958). "How to Use Ridit Analysis." Biometrics, 14(1), 18-38, doi:10.2307/2527727, for the ridit transformation and its interpretation used here.
Examples
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneBernoulli$new(n = nrow(x_dat), response_type = "ordinal",
verbose = FALSE)
for (i in seq_len(nrow(x_dat))) {
seq_des$add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$add_all_subject_responses(as.integer(c(1, 2, 2, 3, 3, 4, 4, 5)))
infer <- InferenceOrdinalRidit$
new(seq_des, verbose = FALSE)
infer
Stereotype Logit Regression Inference for Ordinal Responses
Description
Fits Anderson's (1984) stereotype logit model for ordinal responses (see
fast_stereotype_logit_cpp for the full reduced-rank
multinomial-softmax formula and reparameterization): a single linear
predictor \eta_i = \beta_T W_i + X_i^\top \gamma is scaled by a
category-specific score \phi_k \in [0,1] (jointly estimated,
monotone in k) in a softmax over all K categories, rather than
assuming a single proportional/parallel effect across cuts as
InferenceOrdinalContRatioRegr/
InferenceOrdinalKKCondAdjCatLogitRegr
do. This makes the stereotype model a genuinely more flexible
(multinomial-logit-like, reduced-rank) alternative to the standard
proportional-odds/adjacent-category/continuation-ratio ordinal families,
at the cost of a less directly interpretable treatment coefficient
(\beta_T enters multiplicatively through the \phi_k scores
rather than as a single additive log-odds-ratio). likelihood_tier =
"full": likelihood-ratio, score, gradient, and Wald tests are all
available when the model converges, plus parametric-likelihood-bootstrap
calibration of the likelihood-ratio test.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
compute_lik_ratio_two_sided_pval() can be
anti-conservative at small n. The stereotype model's category-score
parameters \phi_k are a Davies (1977) non-regular case, not identified
when \beta_T is at/near zero – exactly the neighborhood every null
hypothesis test sits in – so the LR statistic's null distribution need not
be the chi-square(1) this method assumes. The size of the problem depends
on the constrained refit: before 2026-09-21 the delta-constrained refit
started cold and could stall in a much worse local optimum (e.g.
neg-log-likelihood 131.3 vs the full fit's 118.9 at \delta equal to
the MLE itself), which inflated the LR statistic and produced ~18-23%
Type-I error at n = 100; it also collapsed the bootstrap and
inverted-LR confidence intervals to zero width. The refit now also starts
from the full-fit parameters and keeps the better fit. Measured after that
change (400 simulated null datasets, 5 response categories, a continuous
covariate): about 6.5% Type-I error and 0.93 confidence-interval coverage
at n = 100, with no zero-width intervals, and about 13% Type-I error
at n = 50 with 3 categories (169 converged fits). Part of the
small-sample excess comes from the likelihood being multimodal: in a few
percent of n = 50 fits the reported point estimate sits at a local
optimum whose likelihood is below that of another optimum, which makes the
LR statistic at the estimate positive rather than zero. For small samples,
prefer the better-calibrated alternatives on this class:
compute_lik_ratio_bootstrap_two_sided_pval() (parametric-bootstrap
calibration via the class's own null simulation; ~6-7% Type-I error) and
the cheaper compute_lik_ratio_bartlett_two_sided_pval()
(Monte-Carlo Bartlett correction; ~4% Type-I error); those two figures
predate the refit change and were not re-measured. Both are
roughly 20-40x slower per call than the raw chi-square test (a fresh
bootstrap/Monte-Carlo refit at every delta candidate), which is why they
are excluded from this package's own routine comprehensive test suite
(see comprehensive_slow_paths.R).
compute_lik_ratio_two_sided_pval()
itself is left unchanged (not silently recalibrated) to avoid an
undocumented behavior change to an existing method's contract.
Bayesian-bootstrap inference is temporarily unavailable because the current
non-uniform weighted hook fits a cumulative-logit surrogate rather than the
stereotype likelihood. It will remain disabled until the native weighted
stereotype-logit backend described in the package implementation plan lands.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceOrdinalStereotypeLogitRegr
Methods
Public methods
-
InferenceOrdinalStereotypeLogitRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceOrdinalStereotypeLogitRegr$supports_rand_pval_for_incidence() -
InferenceOrdinalStereotypeLogitRegr$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceOrdinalStereotypeLogitRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceOrdinalStereotypeLogitRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalStereotypeLogitRegr$supports_rand_pval_for_incidence()
Usage
InferenceOrdinalStereotypeLogitRegr$supports_rand_pval_for_incidence()
InferenceOrdinalStereotypeLogitRegr$compute_rand_two_sided_pval()
Usage
InferenceOrdinalStereotypeLogitRegr$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceOrdinalStereotypeLogitRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceOrdinalStereotypeLogitRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Anderson, J. A. (1984). "Regression and Ordered Categorical Variables." Journal of the Royal Statistical Society, Series B, 46(1), 1-30, doi:10.1111/j.2517-6161.1984.tb01276.x, for the stereotype logit model.
See Also
InferenceOrdinalContRatioRegr
for a proportional (non-reduced-rank) ordinal alternative. See also:
Ordinal
regression (Wikipedia).
Parametric-Bootstrap-Capable Likelihood Inference
Description
Intermediate abstract base for the subset of likelihood-backed inference families that are plausible targets for parametric null-bootstrap likelihood-ratio calibration.
This class sits between InferenceAsympLik and
InferenceAsympLikStdModCache in the hierarchy. Families with highly
bespoke partial-likelihood, quadrature, frailty, copula, or custom
combined-likelihood geometry remain direct children of
InferenceAsympLik and do not pass through here.
The only operational user-facing parametric-bootstrap LR methods on this
surface are
compute_lik_ratio_bootstrap_two_sided_pval(...) and
compute_lik_ratio_bootstrap_confidence_interval(...). Diagnostic
accessors such as get_last_param_bootstrap_diagnostics() are
supplementary and not alternative execution entry points.
Parametric-bootstrap LR calibration is available only for concrete classes
that inherit from InferenceParamBootstrap and whose private method
supports_lik_ratio_param_bootstrap() returns TRUE. Families
that are intentionally unsupported are kept off this branch entirely.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI -> InferenceBayesianBootstrap -> InferenceJackknife -> InferenceAsymp -> InferenceMLEorKMSummaryTable -> InferenceAsympLik -> InferenceParamBootstrap
Methods
Public methods
-
InferenceParamBootstrap$get_last_param_bootstrap_diagnostics() -
InferenceParamBootstrap$compute_lik_ratio_bootstrap_two_sided_pval() -
InferenceParamBootstrap$compute_lik_ratio_bootstrap_confidence_interval() -
InferenceParamBootstrap$get_last_param_bootstrap_estimate_diagnostics() -
InferenceParamBootstrap$compute_param_bootstrap_confidence_interval()
+ inherited public methods from InferenceAsympLik
InferenceAsympLik$compute_asymp_confidence_interval()InferenceAsympLik$compute_asymp_two_sided_pval()InferenceAsympLik$compute_gradient_confidence_interval()InferenceAsympLik$compute_gradient_two_sided_pval()InferenceAsympLik$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAsympLik$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAsympLik$compute_lik_ratio_bartlett_confidence_interval()InferenceAsympLik$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAsympLik$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAsympLik$compute_lik_ratio_bartlett_two_sided_pval()InferenceAsympLik$compute_lik_ratio_confidence_interval()InferenceAsympLik$compute_lik_ratio_two_sided_pval()InferenceAsympLik$compute_score_confidence_interval()InferenceAsympLik$compute_score_two_sided_pval()InferenceAsympLik$get_information_preference()InferenceAsympLik$get_information_source_used()InferenceAsympLik$get_supported_information_preferences()InferenceAsympLik$get_supported_testing_types()InferenceAsympLik$get_testing_type()InferenceAsympLik$set_information_preference()InferenceAsympLik$set_testing_type()
+ inherited public methods from InferenceMLEorKMSummaryTable
+ inherited public methods from InferenceAsymp
+ inherited public methods from InferenceJackknife
InferenceJackknife$approximate_jackknife_distribution_beta_hat_T()InferenceJackknife$compute_jackknife_bias_estimate()InferenceJackknife$compute_jackknife_estimate()InferenceJackknife$compute_jackknife_std_error()InferenceJackknife$compute_jackknife_wald_confidence_interval()InferenceJackknife$compute_jackknife_wald_two_sided_pval()
+ inherited public methods from InferenceBayesianBootstrap
InferenceBayesianBootstrap$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceBayesianBootstrap$compute_bayesian_bootstrap_confidence_interval()InferenceBayesianBootstrap$compute_bayesian_bootstrap_two_sided_pval()InferenceBayesianBootstrap$compute_estimate_with_bootstrap_weights()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_ci_types()InferenceBayesianBootstrap$get_supported_bayesian_bootstrap_pval_types()
+ inherited public methods from InferenceRandBootstrapCI
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceParamBootstrap$get_last_param_bootstrap_diagnostics()
Returns diagnostics from the most recent parametric-bootstrap LR run.
Usage
InferenceParamBootstrap$get_last_param_bootstrap_diagnostics()
Returns
A list of diagnostics, or NULL if no parametric-bootstrap LR run has been executed.
InferenceParamBootstrap$compute_lik_ratio_bootstrap_two_sided_pval()
Bootstrap-calibrated likelihood-ratio two-sided p-value.
Fits the null model at delta, simulates B datasets from
that fitted null, refits unrestricted and null models on each, and
returns the empirical tail probability of the observed LR statistic.
This is the primary user-facing entry point for bootstrap-calibrated
likelihood-ratio p-values.
This method is available only for classes whose private
supports_lik_ratio_param_bootstrap() method returns TRUE.
For unsupported classes it errors immediately rather than silently
falling back to another procedure.
The standard user-facing arguments are delta = 0,
B = 199, and show_progress = FALSE. The remaining
arguments control replicate-quality thresholds and retry behavior.
Runtime cost is roughly one unrestricted fit plus B simulated
unrestricted/null refit pairs, so this is typically much more
expensive than the asymptotic LR p-value.
Usage
InferenceParamBootstrap$compute_lik_ratio_bootstrap_two_sided_pval( delta = 0, B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L )
Arguments
deltaNull treatment effect. Default 0.
BNumber of bootstrap replicates. Default 199.
show_progressLogical; show a progress bar. Default
FALSE.min_number_usable_samplesMinimum number of usable bootstrap replicates required to return a finite p-value. Default
5L.max_attempts_per_replicateMaximum number of simulation/refit retries per bootstrap replicate. Default
2L.
Returns
A scalar p-value, or NA_real_ if the computation fails.
InferenceParamBootstrap$compute_lik_ratio_bootstrap_confidence_interval()
Bootstrap-calibrated likelihood-ratio confidence interval.
Inverts compute_lik_ratio_bootstrap_two_sided_pval via a
bracket-and-bisect search seeded with the Wald interval. Each p-value
evaluation costs B bootstrap refits, so this method is
substantially more expensive than the p-value alone.
This is the primary user-facing entry point for bootstrap-calibrated
likelihood-ratio confidence intervals.
This method is available only for classes whose private
supports_lik_ratio_param_bootstrap_confidence_interval() method
returns TRUE.
The standard user-facing arguments are B = 199 and
show_progress = FALSE. Runtime cost is high because each
confidence-interval bound requires repeated bootstrap p-value
evaluations.
Usage
InferenceParamBootstrap$compute_lik_ratio_bootstrap_confidence_interval( alpha = 0.05, B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L, root_tolerance = NULL, max_root_iterations = 8L )
Arguments
alphaSignificance level. Default 0.05.
BBootstrap replicates per p-value evaluation. Default 199.
show_progressLogical; show a progress bar. Default
FALSE.min_number_usable_samplesMinimum number of usable bootstrap replicates required within each p-value evaluation. Default
5L.max_attempts_per_replicateMaximum number of simulation/refit retries per bootstrap replicate. Default
2L.root_toleranceEffect-scale tolerance for the inversion root. If
NULL, a tolerance proportional to the Wald standard error is used. The bootstrap p-value being inverted is Monte Carlo-discrete, so solving to machine precision is not meaningful.max_root_iterationsMaximum number of bisection iterations per bound during interval inversion. Use
0Lto return the first finite outer bracket. Default8L.
Returns
Named two-element numeric vector with the confidence-interval bounds.
InferenceParamBootstrap$get_last_param_bootstrap_estimate_diagnostics()
Returns diagnostics from the most recent
compute_param_bootstrap_estimate() run.
Usage
InferenceParamBootstrap$get_last_param_bootstrap_estimate_diagnostics()
Returns
A list of diagnostics, or NULL if no parametric-bootstrap
estimate bias correction has been run.
InferenceParamBootstrap$compute_param_bootstrap_estimate()
Parametric-bootstrap bias-corrected point estimate.
Simulates B datasets from the model at the unrestricted
fit (not a null-restricted fit, unlike the LR-bootstrap methods on
this class), refits each unrestricted, and returns
2 * theta_hat - mean(theta_hat_star) – the standard
single-level parametric-bootstrap bias correction (Efron & Tibshirani),
the Monte-Carlo analog of an analytic Cox-Snell (1968) first-order bias
correction. Reuses the same simulate_under_lik_null() contract
every supports_lik_ratio_param_bootstrap() == TRUE family already
implements, anchoring the simulation at the unrestricted fit instead of
a null-restricted one.
This method is available only for classes whose private
supports_param_bootstrap_estimate() method returns TRUE
(by default, this delegates to supports_lik_ratio_param_bootstrap()).
Usage
InferenceParamBootstrap$compute_param_bootstrap_estimate( B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L )
Arguments
BNumber of bootstrap replicates. Default 199.
show_progressLogical; show a progress bar. Default
FALSE.min_number_usable_samplesMinimum number of usable bootstrap replicates required to return a finite estimate. Default
5L.max_attempts_per_replicateMaximum number of simulation/refit retries per bootstrap replicate. Default
2L.
Returns
A scalar bias-corrected estimate, or NA_real_ if the computation fails.
InferenceParamBootstrap$compute_param_bootstrap_confidence_interval()
Parametric-bootstrap "basic" (reflected) confidence interval.
Uses the same kind of simulate-at-full_fit replicates as
compute_param_bootstrap_estimate() – run once, not
re-simulated per candidate null – and reflects their empirical
quantiles around the raw estimate:
CI = [2*theta_hat - q_(1-alpha/2)(theta_star), 2*theta_hat - q_(alpha/2)(theta_star)].
This is the standard "basic" bootstrap interval (Efron & Tibshirani):
the bias-aware companion to compute_param_bootstrap_estimate()
– unlike a naive percentile interval, it reflects around the same
point the estimate's own bias correction targets.
Unlike compute_lik_ratio_bootstrap_confidence_interval(), which
must re-simulate B fresh replicates at every candidate null
value visited during bisection (because that method's replicates are
anchored at a null-restricted fit that changes with the candidate
value), this method's replicates are anchored at the single
unrestricted fit and therefore do not need to be regenerated per
alpha or inverted against any particular null – one batch of
B replicates supports the whole interval.
This method is available only for classes whose private
supports_param_bootstrap_estimate() method returns TRUE.
Usage
InferenceParamBootstrap$compute_param_bootstrap_confidence_interval( alpha = 0.05, B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L )
Arguments
alphaSignificance level. Default 0.05.
BNumber of bootstrap replicates. Default 199.
show_progressLogical; show a progress bar. Default
FALSE.min_number_usable_samplesMinimum number of usable bootstrap replicates required to return a finite interval. Default
5L.max_attempts_per_replicateMaximum number of simulation/refit retries per bootstrap replicate. Default
2L.
Returns
Named two-element numeric vector with the confidence-interval bounds.
InferenceParamBootstrap$compute_param_bootstrap_pval()
Parametric-bootstrap two-sided p-value for H0: theta = delta,
obtained by inverting the same "basic" reflected-quantile construction
as compute_param_bootstrap_confidence_interval() rather than by
simulating fresh replicates under each candidate delta (contrast
compute_lik_ratio_bootstrap_two_sided_pval(), which does
resimulate per delta because its replicates are anchored at a
null-restricted fit that changes with delta).
Reflects delta through the raw estimate, t = 2*theta_hat - delta,
and reports twice the smaller empirical tail of the replicate
distribution beyond t (with the same +1 continuity correction
used by compute_lik_ratio_bootstrap_two_sided_pval()). Because
the replicate batch does not depend on delta, this is a direct
lookup against one batch of B replicates – valid for any
delta with no additional simulation.
As with the confidence interval, this test implicitly assumes the
shape of theta_hat's sampling distribution near delta resembles
its shape at the unrestricted fit – an approximation that degrades
the further delta is from the observed estimate, and is why
this p-value should be treated as a cheap default rather than a
replacement for compute_lik_ratio_bootstrap_two_sided_pval()
when accuracy near a specific null matters.
Usage
InferenceParamBootstrap$compute_param_bootstrap_pval( delta, B = 199, show_progress = FALSE, min_number_usable_samples = 5L, max_attempts_per_replicate = 2L )
Arguments
deltaNull value of the coefficient of interest.
BNumber of bootstrap replicates. Default 199.
show_progressLogical; show a progress bar. Default
FALSE.min_number_usable_samplesMinimum number of usable bootstrap replicates required to return a finite p-value. Default
5L.max_attempts_per_replicateMaximum number of simulation/refit retries per bootstrap replicate. Default
2L.
Returns
A scalar p-value, or NA_real_ if the computation fails.
Shared replicate-running core for anything built on B simulated null
datasets via simulate_under_lik_null() (currently:
compute_lik_ratio_bootstrap_two_sided_pval() and, via
InferenceExtBartlettApprox, get_bartlett_factor_approx()). Handles
seeding, multi-core parallelism, the reusable-worker-state optimization,
and deterministic-mode thread budgeting uniformly, so every caller gets
the same performance characteristics for free. Returns a list with
$results (one per-replicate result object per B, each carrying at least
$lr), $used_worker_path, and $used_deterministic_mode.
Callers are responsible for setting private$active_resampling_operation themselves (not done here, to avoid a nested caller clobbering an already-active outer flag via a premature on.exit reset). Simulate a bootstrap dataset under the fitted null likelihood and return a minimal spec list for refitting.
Must be overridden by families that set supports_lik_ratio_param_bootstrap() to TRUE. The returned list must contain at least: - full_fit: unrestricted fit on the simulated data - fit_null: function(delta, start) returning a constrained fit - neg_loglik: function(fit) returning the neg-log-likelihood
InferenceParamBootstrap$clone()
The objects of this class are cloneable with this method.
Usage
InferenceParamBootstrap$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Beta Regression Inference for Proportion Responses
Description
Fits Ferrari and Cribari-Neto's (2004) beta regression for proportion
responses Y_i \in (0, 1): \mathrm{logit}(E[Y_i \mid w_i, x_i]) =
\beta_0 + \beta_T w_i + x_i^\top \gamma, Y_i \mid w_i, x_i \sim
\mathrm{Beta}(\mu_i \phi, (1-\mu_i)\phi) for fitted mean \mu_i and a
single (constant, not covariate-dependent) precision parameter \phi,
by maximum likelihood (fast_beta_regression_cpp/
fast_beta_regression_weighted_cpp). \hat\beta_T is a
log-odds-ratio on the conditional-mean scale: \exp(\hat\beta_T) is
the odds ratio for the expected proportion. Unlike
InferencePropFractionalLogit's
quasi-likelihood (which specifies only the conditional mean), beta
regression also specifies the conditional variance/shape via \phi —
a correctly specified beta model yields a fully efficient likelihood-based
fit and genuine likelihood-ratio/score/gradient tests, at the cost of
requiring the beta-distribution shape assumption to actually hold.
likelihood_tier = "full": likelihood-ratio, score, gradient, and
Wald tests are all available when the model converges, plus
parametric-likelihood-bootstrap calibration of the likelihood-ratio test.
Y_i values of exactly 0 or 1 are not supported by the beta density
and are handled by sanitize_beta_response()'s boundary adjustment
before fitting.
Estimand. Composes
MarginalEstimand
(set_estimand()/get_estimand()/get_supported_estimands()).
Under the default estimand = "conditional", \hat\beta_T is the
log-odds-ratio above. Under estimand = "marginal_mean_diff", the
reported quantity is instead the g-computation marginal mean difference
\frac{1}{n}\sum_i \{\mathrm{plogis}(\hat\beta_0 + \hat\beta_T +
X_i^\top \hat\gamma) - \mathrm{plogis}(\hat\beta_0 + X_i^\top
\hat\gamma)\} (the precision parameter \phi does not enter the
mean, so it plays no role in this functional). Only
"marginal_mean_diff" is supported — a ratio of two mean
proportions, both bounded in [0,1], is not the standard estimand
for a beta-regression treatment effect the way a rate ratio is for count
data. Because there is no latent submodel for this family (unlike e.g.
InferencePropZeroOneInflatedBetaRegr's
zero/one-inflation mixture), the marginal mean function is exactly the
model's own fitted mean; no separate standardization step beyond the
g-computation average is needed. Standard errors under the marginal
estimand use the delta method against the mean-submodel coefficient
covariance (degrees of freedom Inf); testing_type is
restricted to "wald" whenever the estimand is non-conditional. The
underlying model fit is identical regardless of estimand — switching
estimand is a pure post-fit transform, never a refit.
Super class
Inference -> InferencePropBetaRegr
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropBetaRegr$new()
Initialize inference for the beta regression model
\mathrm{logit}(E[Y_i \mid w_i, x_i]) = \beta_0 + \beta_T w_i +
x_i^\top \gamma, Y_i \sim \mathrm{Beta}(\mu_i \phi, (1-\mu_i)
\phi); see
InferencePropBetaRegr for the
model form. Does not fit the model; the fit is deferred to the first
call to compute_estimate() or a method that requires it.
Usage
InferencePropBetaRegr$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a proportion response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
optimization_algCharacter scalar specifying the optimization algorithm. Default is dispatched via policy.
InferencePropBetaRegr$compute_estimate()
Fits the beta regression model by maximum likelihood
(jointly estimating the mean coefficients and the precision parameter
\phi). Under the default estimand = "conditional",
returns the log-odds-ratio estimate \hat\beta_T on the
conditional-mean scale. Under estimand = "marginal_mean_diff"
(set via set_estimand()), returns the g-computation
marginal mean difference instead — see the class-level
@details for the formula. The underlying model fit is
identical either way (a pure post-fit transform of the same
cached fit, no refit).
Usage
InferencePropBetaRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferencePropBetaRegr$compute_asymp_confidence_interval()
Wald confidence interval, dispatched by
testing_type for the conditional estimand (score/gradient/
likelihood-ratio/Bartlett available; see
InferenceAsympLik); under a
marginal estimand testing_type is always "wald" (the
only value set_estimand() permits there), so this always
resolves to the delta-method interval. Calls
self$compute_estimate() first (not private$shared()
directly) so the estimand-aware cache is always current
regardless of call order.
Usage
InferencePropBetaRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferencePropBetaRegr$compute_asymp_two_sided_pval()
Wald two-sided p-value, dispatched by
testing_type exactly as
compute_asymp_confidence_interval(); see that method's
description for the marginal-estimand always-Wald note.
Usage
InferencePropBetaRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment-effect value under the current estimand (conditional log-odds-ratio, or marginal mean difference).
InferencePropBetaRegr$compute_estimate_with_bootstrap_weights()
Refits the beta model with subject/block-level weights
applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()) via
fast_beta_regression_weighted_cpp, and returns the
reweighted estimate \hat\beta_T^{(w)}. Uses the same QR
column-dropping hardening as compute_estimate(); a
hardened-but-still-unreasonable fit is cached as nonestimable.
Usage
InferencePropBetaRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferencePropBetaRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropBetaRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Ferrari, S., and Cribari-Neto, F. (2004). "Beta regression for modelling rates and proportions." Journal of Applied Statistics, 31(7), 799-815, doi:10.1080/0266476042000214501.
See Also
InferencePropFractionalLogit
for a quasi-likelihood proportion model that specifies only the
conditional mean. Comparable Python API: no direct beta-regression
equivalent in statsmodels; see
statsmodels GLM for
the general exponential-family GLM framework. See also:
Beta
distribution (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropBetaRegr$new(seq_des)
inf$compute_estimate()
Fractional Logit Inference for Proportion Responses
Description
Fits Papke and Wooldridge's (1996) fractional logistic (quasi-binomial)
regression for proportion responses Y_i \in [0, 1] (not restricted to
\{0, 1\}):
E[Y_i \mid w_i, x_i] = \mathrm{logit}^{-1}(\beta_0 + \beta_T w_i +
x_i^\top \gamma), fit by maximizing the Bernoulli quasi-log-likelihood
\sum_i \{Y_i \log \mu_i + (1 - Y_i) \log(1 - \mu_i)\} treating
Y_i as if it were binary (a valid estimating equation for the
conditional mean even though Y_i is fractional — the Bernoulli
log-likelihood's score is unbiased for the true mean regardless of the
actual distribution of Y_i on [0,1]). \hat\beta_T is a
log-odds-ratio on the conditional-mean scale: \exp(\hat\beta_T) is
the odds ratio for the expected proportion. Standard errors use the
model-based (non-robust/non-sandwich) Fisher information from this
quasi-likelihood, scaled by an estimated quasi-binomial dispersion
parameter \hat\phi (Papke & Wooldridge's own prescription; fixed
2026-09-06 – the unscaled Bernoulli-based variance systematically
overstates \mathrm{Var}(\hat\beta_T) for a genuinely fractional
response, since Bernoulli is the maximum-variance distribution on [0,1]
for a given mean); only Wald inference is
exposed (private$supports_likelihood_tests() is hard FALSE
here even though likelihood_tier = "full" metadata is set for
component-composition purposes — this class deliberately does not compose
ParametricLikelihoodBootstrap, so no likelihood-ratio/score/gradient
test surface is exposed). Validity requires that the conditional mean is
correctly specified on the logit scale; unlike beta regression, no
assumption is made about the conditional variance or shape of Y_i's
distribution.
Super class
Inference -> InferencePropFractionalLogit
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropFractionalLogit$new()
Initialize inference for the fractional logit model
E[Y_i \mid w_i, x_i] = \mathrm{logit}^{-1}(\beta_0 + \beta_T w_i +
x_i^\top \gamma); see
InferencePropFractionalLogit
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Usage
InferencePropFractionalLogit$new( des_obj, model_formula = NULL, verbose = FALSE, harden = TRUE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a proportion response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
hardenWhether to apply robustness measures.
smart_cold_start_defaultWhether to use smart cold start values.
InferencePropFractionalLogit$compute_estimate()
Fits the fractional logit model by maximizing the Bernoulli
quasi-log-likelihood on the fractional response and returns the
log-odds-ratio estimate \hat\beta_T. When estimate_only =
TRUE and hardening is disabled (harden = FALSE), uses a fast
path via base R's glm.fit(family = quasibinomial()) instead of
the package's own fitting routine; otherwise dispatches through the
shared hardened-fit path.
Usage
InferencePropFractionalLogit$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations; when combined with
harden = FALSE, also switches to thequasibinomial()fast path.
InferencePropFractionalLogit$compute_estimate_with_bootstrap_weights()
Refits the fractional logit model with subject/block-level
weights applied to the fitting quasi-log-likelihood (Bayesian-bootstrap
or nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()), and
returns the reweighted estimate \hat\beta_T^{(w)}. Uses the same
QR column-dropping hardening as compute_estimate()'s hardened
path; a hardened-but-still-unreasonable fit is cached as nonestimable.
Usage
InferencePropFractionalLogit$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferencePropFractionalLogit$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropFractionalLogit$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Papke, L. E., and Wooldridge, J. M. (1996). "Econometric Methods for Fractional Response Variables with an Application to 401(K) Plan Participation Rates." Journal of Applied Econometrics, 11(6), 619-632, doi:10.1002/(SICI)1099-1255(199611)11:6<619::AID-JAE418>3.0.CO;2-1.
See Also
InferencePropBetaRegr for
a proportion model that also specifies the conditional variance/shape.
Comparable Python API:
statsmodels GLM
(family=Binomial() on fractional response data). See also:
Logistic
regression (Wikipedia).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropFractionalLogit$new(seq_des)
inf$compute_estimate()
G-Computation Mean-Difference Inference for Proportion Responses
Description
Fits a fractional-logit working model, \mathrm{logit}\,E[Y_i \mid x_i] =
x_i^\top\hat\beta (via fast_logistic_regression_cpp, treating
the continuous-in-(0,1) proportion as a quasi-binomial mean model —
the same mean-model idea as
InferencePropFractionalLogit),
for a proportion outcome using treatment and, optionally, all recorded
covariates, then estimates the marginal mean difference by G-computation:
standardizing predicted mean proportions under all-treated and all-control
assignments over the empirical covariate distribution (see
gcomp_fractional_logit_point_estimate_cpp for the exact
standardization formula). Inference uses a Huber-White sandwich-robust
covariance for the regression coefficients and the delta method
(analytic gradient of the standardized mean-difference functional with
respect to \hat\beta, by default) to propagate that covariance onto
the mean-difference scale, \widehat{\mathrm{Var}}(\widehat{\mathrm{md}})
= \nabla^\top \widehat{\mathrm{Var}}(\hat\beta) \nabla.
The implementation is optimized for resampling-based inference. It utilizes a fast C++ IRLS solver for the underlying fractional logit regression. During resampling draws, it bypasses the calculation of the sandwich covariance matrix and delta-method standard errors, providing a significant speedup when computing bootstrap or randomization distributions.
Variance fallback cascade. If the primary analytic-gradient/sandwich-covariance
variance is non-finite (e.g. near-boundary fitted probabilities), up to eight
progressively more conservative fallback strategies are tried in order (see
the variance_fallback_methods constructor argument for the full list
and their individual definitions): stabilized (PSD-projected) sandwich
covariance, model-based (Fisher information) covariance, finite-difference
gradients in place of the analytic delta-method gradient, and combinations
with progressively stronger probability clipping. The first strategy in the
ordered list that yields a finite, positive variance is used; an empty
variance_fallback_methods vector always returns NA variance
rather than erroring.
Super class
Inference -> InferencePropGCompMeanDiff
Methods
Public methods
-
InferencePropGCompMeanDiff$compute_estimate_with_bootstrap_weights() -
InferencePropGCompMeanDiff$compute_asymp_confidence_interval() -
InferencePropGCompMeanDiff$compute_wald_confidence_interval() -
InferencePropGCompMeanDiff$compute_bootstrap_two_sided_pval() -
InferencePropGCompMeanDiff$compute_bootstrap_confidence_interval() -
InferencePropGCompMeanDiff$approximate_bootstrap_distribution_beta_hat_T()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropGCompMeanDiff$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize the g-computation inference object.
Usage
InferencePropGCompMeanDiff$new(
des_obj,
model_formula = NULL,
verbose = FALSE,
prob_clip_eps = 1e-06,
prob_clip_strong_eps = 1e-04,
max_resample_attempts = 50L,
smart_cold_start_default = NULL,
harden = TRUE,
variance_fallback_methods = c("robust", "stabilized_robust", "model_based",
"stabilized_robust_fd", "model_based_fd", "stabilized_robust_strong_clip",
"model_based_strong_clip", "model_based_fd_strong_clip")
)
Arguments
des_objA completed
DesignSeqOneByOneobject with a proportion response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseWhether to print progress messages.
prob_clip_epsPrimary probability clamp applied to fitted values during model-based variance computation. Predicted probabilities are clipped to
[prob_clip_eps, 1 - prob_clip_eps]before computing IWLS weights. Must be in[0, 0.5). Default1e-6.prob_clip_strong_epsStronger clamp used as a fallback when the primary variance strategy fails. Predicted probabilities and gradients are clipped to
[prob_clip_strong_eps, 1 - prob_clip_strong_eps]. Must be in[0, 0.5)and should be\geprob_clip_eps. Default1e-4.max_resample_attemptsMaximum number of times a single bootstrap replicate may be redrawn when the drawn sample fails validity screening (e.g. near-perfect separation, too few observations per arm, excessive boundary mass). If all attempts fail the replicate is recorded as
NA, silently reducing the effectiveB. Can be overridden per-call inapproximate_bootstrap_distribution_beta_hat_T. Must be a positive integer. Default50L.smart_cold_start_defaultWhether to use smart cold start values.
hardenWhether to apply robustness measures.
variance_fallback_methodsOrdered character vector of variance strategies to attempt in sequence. Each name corresponds to a (gradient, covariance-matrix) pair; the first strategy that yields a finite, positive variance is used. Allowed values (in their default order) are:
"robust"Analytic delta-method gradient with the sandwich (HC) covariance.
"stabilized_robust"Same gradient; covariance projected to the nearest PSD matrix.
"model_based"Same gradient; Fisher-information (IWLS) covariance. Fitted probabilities are first clipped to
[prob_clip_eps, 1 - prob_clip_eps], then the binomial variance weightw_i = \hat\mu_i(1-\hat\mu_i)is further capped at0.25. The cap is the global maximum ofp(1-p), attained atp = 0.5; it prevents a near-boundary fitted probability that slips through the clip from inflating the information matrix, which is mathematically correct because no Bernoulli variance can exceed 0.25."stabilized_robust_fd"Finite-difference (central-difference) gradient; stabilized sandwich covariance. The step size for coefficient
jish_j = \varepsilon^{1/3}(|\hat\beta_j| + 1), where\varepsilon =.Machine$double.eps. The cube-root of machine epsilon is the theoretically optimal step that balances truncation error (O(h^2)for central differences) against floating-point cancellation (O(\varepsilon / h)), giving a total error ofO(\varepsilon^{2/3})."model_based_fd"Finite-difference gradient (same step rule as
"stabilized_robust_fd"); model-based covariance with the 0.25 weight cap."stabilized_robust_strong_clip"Strong-clipped analytic gradient; stabilized sandwich covariance.
"model_based_strong_clip"Strong-clipped analytic gradient; model-based covariance (strong-clipped) with the 0.25 weight cap.
"model_based_fd_strong_clip"Strong-clipped finite-difference gradient (same step rule); model-based covariance (strong-clipped) with the 0.25 weight cap.
Pass a shorter vector or a single string to restrict which strategies are tried. An empty vector always returns
NAvariance.
InferencePropGCompMeanDiff$compute_estimate()
Computes the g-computation treatment-effect estimate.
Usage
InferencePropGCompMeanDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferencePropGCompMeanDiff$get_standard_error()
Returns the standard error of the g-computation mean-difference
estimate (NA if it is unavailable).
Usage
InferencePropGCompMeanDiff$get_standard_error()
Returns
A single numeric standard error, or NA_real_.
InferencePropGCompMeanDiff$compute_estimate_with_bootstrap_weights()
Recomputes the g-computation mean-difference estimate under the supplied subject- or block-level bootstrap weights and caches it; used by the Bayesian-bootstrap and resampling paths.
Usage
InferencePropGCompMeanDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsNumeric vector of bootstrap weights, one per subject (or per block when the design is blocked).
estimate_onlyIf TRUE, only the point estimate is required (no variance).
Returns
The weighted mean-difference estimate, or NA_real_ if the weighted fit is unusable.
InferencePropGCompMeanDiff$compute_asymp_confidence_interval()
Computes a 1 - alpha confidence interval.
Usage
InferencePropGCompMeanDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha.
InferencePropGCompMeanDiff$compute_asymp_two_sided_pval()
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
InferencePropGCompMeanDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null mean difference. Defaults to 0.
InferencePropGCompMeanDiff$compute_wald_two_sided_pval()
Computes a Wald two-sided p-value for the treatment effect.
Usage
InferencePropGCompMeanDiff$compute_wald_two_sided_pval(delta = 0)
Arguments
deltaThe null mean difference. Defaults to 0.
InferencePropGCompMeanDiff$compute_wald_confidence_interval()
Computes a Wald confidence interval for the treatment effect.
Usage
InferencePropGCompMeanDiff$compute_wald_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level. Default 0.05.
InferencePropGCompMeanDiff$compute_bootstrap_two_sided_pval()
Computes a bootstrap two-sided p-value for the treatment effect.
Usage
InferencePropGCompMeanDiff$compute_bootstrap_two_sided_pval( delta = 0, B = 501, type = "symmetric", na.rm = FALSE, boundary_tol = 0.02, max_boundary_mass = 0.95, sep_tol = 0.02, min_group_n = 5L, show_progress = TRUE, min_number_usable_samples = 5L )
Arguments
deltaThe null mean difference. Defaults to 0.
BNumber of bootstrap samples.
typeBootstrap p-value type. See
InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval.na.rmWhether to remove non-finite bootstrap replicates.
boundary_tolResample screening threshold for boundary mass near 0/1.
max_boundary_massReject a resample when at least this fraction is near the boundary.
sep_tolSeparation tolerance used to reject nearly perfectly separated resamples.
min_group_nMinimum number of observations required in each treatment arm.
show_progressWhether to show a progress bar.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
InferencePropGCompMeanDiff$compute_bootstrap_confidence_interval()
Generic (non-screening-modified) bootstrap two-sided p-value, aliased directly from 'InferenceNonParamBootstrap' so 'compute_bootstrap_two_sided_pval()' can dispatch to it without relying on 'super$', which does not resolve under flat component composition.
Computes a bootstrap confidence interval.
Usage
InferencePropGCompMeanDiff$compute_bootstrap_confidence_interval( alpha = 0.05, B = 501, type = NULL, na.rm = TRUE, show_progress = TRUE, boundary_tol = 0.02, max_boundary_mass = 0.95, sep_tol = 0.02, min_group_n = 5L, min_number_usable_samples = 5L )
Arguments
alphaThe confidence level 1 -
alpha.BNumber of bootstrap samples.
typeBootstrap CI type.
na.rmWhether to remove non-finite bootstrap replicates.
show_progressWhether to show bootstrap progress.
boundary_tolResample screening threshold for boundary mass near 0/1.
max_boundary_massReject a resample when at least this fraction is near the boundary.
sep_tolSeparation tolerance used to reject nearly perfectly separated resamples.
min_group_nMinimum number of observations required in each treatment arm.
min_number_usable_samplesMinimum number of finite bootstrap samples required.
InferencePropGCompMeanDiff$approximate_bootstrap_distribution_beta_hat_T()
Generic (non-screening-modified) bootstrap confidence interval; see 'compute_bootstrap_two_sided_pval_generic'.
Abbreviated bootstrap sampler that reuses a bootstrap worker.
Usage
InferencePropGCompMeanDiff$approximate_bootstrap_distribution_beta_hat_T( B = 501, show_progress = TRUE, max_resample_attempts = NULL, boundary_tol = 0.02, max_boundary_mass = 0.95, sep_tol = 0.02, min_group_n = 5L, debug = FALSE )
Arguments
BThe number of bootstrap samples (default 501).
show_progressWhether to show a progress bar.
max_resample_attemptsMaximum redraw attempts per bootstrap replicate before the replicate is recorded as
NA.NULL(default) uses the value set at construction time.boundary_tolResample screening threshold for boundary mass near 0/1.
max_boundary_massReject a resample when at least this fraction is near the boundary.
sep_tolSeparation tolerance used to reject nearly perfectly separated resamples.
min_group_nMinimum number of observations required in each treatment arm.
debugIf TRUE, return per-replicate diagnostics (values, errors, warnings) instead of just the bootstrap values.
InferencePropGCompMeanDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropGCompMeanDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropGCompMeanDiff$new(seq_des)
inf$compute_estimate()
GEE Inference for KK Designs with Proportion Response
Description
Fits a Generalized Estimating Equations model with a binomial (quasi-likelihood,
fractional-response) family and logit link, \mathrm{logit}\,E[Y_i \mid x_i]
= x_i^\top\beta, for proportion (continuous values in (0, 1)) responses
under a KK matching-on-the-fly design — the same fractional-logit mean-model
idea as InferencePropFractionalLogit,
extended to jointly account for matched-pair and reservoir clustering via GEE.
Each GEE cluster is either a matched pair (2 members) or a reservoir singleton
(1 member), with an exchangeable working correlation structure — see
$compute_estimate()'s method-level documentation for the full fitting
contract (internal Rcpp solver vs. geepack fallback, hardening/retry
behavior). Inference is quasi-likelihood/estimating-equation based
(likelihood_tier = "quasi"): standard errors are GEE sandwich (robust)
standard errors, not model-likelihood-based.
Super class
Inference -> InferencePropKKGEE
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropKKGEE$new()
Initialize KK proportion-response GEE inference, validate
the matched/reservoir design, and prepare the exchangeable-working-correlation
fractional-logit GEE fitting machinery used by
InferencePropKKGEE.
Usage
InferencePropKKGEE$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a proportion response.model_formulaOptional formula for covariate adjustment.
use_rcppWhether to use the internal Rcpp GEE solver (
TRUE, default) with automatic fallback togeepack::geeglmon failure, or always usegeepack::geeglmdirectly (FALSE, requires geepack to be installed).verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferencePropKKGEE$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropKKGEE$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Liang, K.-Y., and Zeger, S. L. (1986). "Longitudinal Data Analysis Using Generalized Linear Models." Biometrika, 73(1), 13-22, doi:10.1093/biomet/73.1.13, for the GEE estimating-equation framework and sandwich variance estimator used here.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropKKGEE$new(seq_des)
inf$compute_estimate()
KK GLMM Inference for Proportion Responses
Description
Fits a combined conditional-logit-plus-random-intercept-GLMM likelihood for
proportion responses under a KK matching-on-the-fly design. Matched pairs
with a discordant pair-difference are handled by a conditional (fixed
pair-effect) logistic term, while concordant/reservoir subjects are handled
by a random-intercept logistic mixed model with intercept
b_g \sim N(0, \sigma_b^2) per matched-set/reservoir group g;
both terms share the same treatment coefficient \beta_T, jointly
maximized by the internal fast_clogit_plus_glmm_cpp routine. This combines the
design-exact conditional-logit treatment of matched pairs (no nuisance
pair-intercept to estimate) with a GLMM's ability to still contribute
information from concordant pairs and reservoir subjects, which a pure
conditional-logit-on-discordant-pairs-only approach would discard.
\exp(\hat\beta_T) is the common treatment odds ratio.
likelihood_tier = "full": likelihood-ratio, score, and Wald tests
are all available when the model converges. See
InferenceAbstractKKCondLogitGLMM
for the shared model-fitting and caching contract used by this class's
incidence-response siblings
(InferenceIncidKKCondLogitGLMMIVWC,
InferenceIncidKKCondLogitGLMMOneLik).
Super classes
Inference -> InferenceAbstractKKCondLogitGLMM -> InferencePropKKGLMM
Methods
Public methods
+ inherited public methods from InferenceAbstractKKCondLogitGLMM
InferenceAbstractKKCondLogitGLMM$approximate_bayesian_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_jackknife_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_rand_bootstrap_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_randomization_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$approximate_subsampling_distribution_beta_hat_T()InferenceAbstractKKCondLogitGLMM$compute_asymp_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_asymp_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_bayesian_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_bayesian_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_estimate()InferenceAbstractKKCondLogitGLMM$compute_estimate_with_bootstrap_weights()InferenceAbstractKKCondLogitGLMM$compute_gradient_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_gradient_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_jackknife_bias_estimate()InferenceAbstractKKCondLogitGLMM$compute_jackknife_estimate()InferenceAbstractKKCondLogitGLMM$compute_jackknife_std_error()InferenceAbstractKKCondLogitGLMM$compute_jackknife_wald_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_jackknife_wald_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_approx_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_approx_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_exact_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_exact_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bartlett_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_lik_ratio_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_m_out_of_n_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_estimate()InferenceAbstractKKCondLogitGLMM$compute_param_bootstrap_pval()InferenceAbstractKKCondLogitGLMM$compute_rand_bootstrap_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_rand_bootstrap_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_rand_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_rand_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_score_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_score_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_subsampling_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_subsampling_sensitivity()InferenceAbstractKKCondLogitGLMM$compute_subsampling_two_sided_pval()InferenceAbstractKKCondLogitGLMM$compute_wald_confidence_interval()InferenceAbstractKKCondLogitGLMM$compute_wald_two_sided_pval()InferenceAbstractKKCondLogitGLMM$get_information_preference()InferenceAbstractKKCondLogitGLMM$get_information_source_used()InferenceAbstractKKCondLogitGLMM$get_last_param_bootstrap_diagnostics()InferenceAbstractKKCondLogitGLMM$get_last_param_bootstrap_estimate_diagnostics()InferenceAbstractKKCondLogitGLMM$get_mod()InferenceAbstractKKCondLogitGLMM$get_summary()InferenceAbstractKKCondLogitGLMM$get_supported_bayesian_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_bayesian_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_information_preferences()InferenceAbstractKKCondLogitGLMM$get_supported_rand_bootstrap_ci_types()InferenceAbstractKKCondLogitGLMM$get_supported_rand_bootstrap_pval_types()InferenceAbstractKKCondLogitGLMM$get_supported_testing_types()InferenceAbstractKKCondLogitGLMM$get_testing_type()InferenceAbstractKKCondLogitGLMM$select_optimal_b_subsampling()InferenceAbstractKKCondLogitGLMM$select_optimal_m_out_of_n_bootstrap()InferenceAbstractKKCondLogitGLMM$set_information_preference()InferenceAbstractKKCondLogitGLMM$set_testing_type()InferenceAbstractKKCondLogitGLMM$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropKKGLMM$new()
Initialize inference for the combined conditional-logit
(discordant matched pairs) plus random-intercept-GLMM (concordant
pairs/reservoir) proportion model; see
InferencePropKKGLMM for the model
form. Does not fit the model; the fit is deferred to the first call to
compute_estimate() or a method that requires it.
Usage
InferencePropKKGLMM$new( des_obj, model_formula = NULL, max_abs_reasonable_coef = 10000, max_abs_log_sigma = 8, verbose = FALSE, smart_cold_start_default = NULL, optimization_alg = NULL )
Arguments
des_objA completed
Designobject with a proportion response.model_formulaOptional formula for covariate adjustment.
max_abs_reasonable_coefCap for reasonable coefficient estimates.
max_abs_log_sigmaCap for reasonable log random effect variance.
verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
optimization_algCharacter. Optimization algorithm (default "lbfgs").
InferencePropKKGLMM$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropKKGLMM$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kapelner, A. and Krieger, A. M. (2014). "Matching on-the-fly: Sequential allocation with higher power and efficiency." Biometrics, 70(2), 378-388, doi:10.1111/biom.12148, for the KK matching-on-the-fly design this class is built for; Breslow, N. E., and Clayton, D. G. (1993). "Approximate Inference in Generalized Linear Mixed Models." Journal of the American Statistical Association, 88(421), 9-25, doi:10.2307/2290687, for the GLMM likelihood framework combined with the conditional-logit term here.
Quantile Regression Compound Estimator for KK Matching-on-the-Fly Designs (Proportion Outcomes)
Description
A variance-weighted compound quantile regression estimator for KK matching-on-the-fly
designs with proportion responses. Inference is performed on the logit (log-odds)
scale: responses y \in (0,1) are transformed via \text{logit}(y) = \log(y/(1-y))
before quantile regression.
The estimator combines:
Quantile regression on logit-scale within-pair differences
\text{logit}(y_T) - \text{logit}(y_C)(matched pairs)Quantile regression of
\text{logit}(y)on treatment and covariates (reservoir)
using the same variance-weighted combination logic as the OLS compound estimator.
The estimated treatment effect is a log-odds-ratio shift at quantile tau.
At beta_T = 1 (one log-odds-ratio unit of treatment effect), the population
treatment effect on the logit scale is exactly 1, so no skip_ci is needed.
Default quantile: tau = 0.5 (median regression).
To target a different quantile — for example the 25th or 75th percentile — pass
tau = 0.25 or tau = 0.75 to the constructor:
inf = InferencePropKKQuantileRegrIVWC$ new(seq_des, tau = 0.75)
Any value strictly between 0 and 1 is accepted.
Standard errors use Powell's "nid" sandwich estimator (non-iid), falling back to "iid" on failure. Asymptotic z-based inference is used throughout.
This class requires the quantreg package, which is listed in Suggests and is not installed automatically with EDI. Install quantreg before using this class.
Legacy class. Not fully tested in comprehensive_tests.R.
Super class
Inference -> InferencePropKKQuantileRegrIVWC
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropKKQuantileRegrIVWC$new()
Initialize proportion-response KK IVWC quantile-regression
inference on the logit response scale; see
InferencePropKKQuantileRegrIVWC.
Usage
InferencePropKKQuantileRegrIVWC$new( des_obj, model_formula = NULL, tau = 0.5, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.tauThe quantile level for regression on the logit scale, strictly between 0 and 1. The default
tau = 0.5estimates the median log-odds-ratio treatment effect. Pass a different value (e.g.tau = 0.25ortau = 0.75) to target a different percentile of the treatment effect distribution.verboseA flag indicating whether messages should be displayed to the user. Default is
FALSE.smart_cold_start_defaultWhether to use smart cold start values.
InferencePropKKQuantileRegrIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropKKQuantileRegrIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropKKQuantileRegrIVWC$new(seq_des)
inf$compute_estimate()
Quantile Regression Combined-Likelihood Compound Estimator for KK Designs (Proportion)
Description
Fits the combined stacked quantile regression (matched-pair differences + reservoir)
using the treatment indicator and all recorded covariates for proportion responses.
Responses y \in (0,1) are transformed via
\mathrm{logit}(y) = \log(y/(1-y)) before regression; the estimated
treatment effect \hat\beta_T is a log-odds-ratio shift at quantile
tau of the logit-transformed response. Minimizes the joint
check-function (pinball) loss \rho_\tau(u) = u(\tau - \mathbb{1}\{u<0\})
over both data sources simultaneously in one quantreg fit, unlike the
IVWC sibling, which fits
matched-pair and reservoir quantile regressions separately and pools them by
inverse-variance weighting. Standard errors use Powell's sandwich estimator.
likelihood_tier = "none": quantile regression minimizes an
asymmetric-loss objective, not a proper likelihood, so no likelihood-ratio or
parametric-bootstrap methods are exposed. Requires the quantreg
package.
Super class
Inference -> InferencePropKKQuantileRegrOneLik
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropKKQuantileRegrOneLik$new()
Initialize proportion-response KK combined-likelihood quantile-regression inference.
Responses are fitted on the logit scale; the shared stacked quantile-regression
fit is documented in
InferenceContinKKQuantileRegrOneLik
(this class's continuous-response sibling, sharing the same
KKQuantileRegrOneLik component).
Usage
InferencePropKKQuantileRegrOneLik$new( des_obj, model_formula = NULL, tau = 0.5, verbose = FALSE )
Arguments
des_objA DesignSeqOneByOne object whose entire n subjects are assigned and response y is recorded within.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.tauThe quantile level on the logit scale, strictly between 0 and 1. Default is 0.5.
verboseWhether to print progress messages.
InferencePropKKQuantileRegrOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropKKQuantileRegrOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Koenker, R. (2005). Quantile Regression. Cambridge University Press. doi:10.1017/CBO9780511754098
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropKKQuantileRegrOneLik$new(seq_des)
inf$compute_estimate()
Quantile Regression Inference for Proportion Responses
Description
Fits a quantile regression for proportion responses (constrained to (0, 1)) using
the treatment indicator and, optionally, all recorded covariates as predictors.
Inference is performed on the logit (log-odds) scale: responses
y \in (0,1) are transformed via \text{logit}(y) = \log(y/(1-y)) before
quantile regression, so the estimated treatment effect is a log-odds-ratio
shift at quantile tau; by default tau = 0.5, so this is a median
log-odds-ratio shift.
Fitting is via rq (method "br", the Barrodale-Roberts
simplex algorithm, for the point estimate; the default Frisch-Newton-adjacent
interior-point path for the variance-computing fit) on logit(y) ~ w + covariates
with no intercept column (the design matrix already carries one). Standard errors use
quantreg's Powell (1991) kernel sandwich "nid" estimator (heteroskedasticity-
and design-robust, valid under non-i.i.d. errors) when available, falling back to the
i.i.d.-errors "iid" estimator if "nid" extraction fails; inference on the
resulting standard error uses the asymptotic normal (Wald) approximation, not a
resampling-based reference distribution, for the asymptotic CI/p-value paths.
compute_asymp_confidence_interval/compute_asymp_two_sided_pval use the
fit's residual degrees of freedom n - p in a t-reference (via
compute_z_or_t_ci_from_s_and_df) rather than a plain normal reference, so the
interval/test remain slightly conservative in small samples relative to a bare Wald z.
This class requires the quantreg package, which is listed under
Suggests and is not installed automatically with EDI.
Install quantreg manually before use. Only uncensored proportion responses are
supported (checked via assertNoCensoring at construction).
Super class
Inference -> InferencePropQuantileRegr
Methods
Public methods
-
InferencePropQuantileRegr$compute_estimate_with_bootstrap_weights() -
InferencePropQuantileRegr$compute_asymp_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropQuantileRegr$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize a quantile-regression inference object for a completed design with a proportion response.
Usage
InferencePropQuantileRegr$new( des_obj, model_formula = NULL, tau = 0.5, verbose = FALSE )
Arguments
des_objA completed
Designobject with a proportion response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.tauThe quantile to estimate (default 0.5).. Default 0.5.
verboseWhether to print progress messages.. Default FALSE.
InferencePropQuantileRegr$compute_estimate()
Computes the fitted treatment coefficient of the tau-quantile
regression of logit(y) on the treatment indicator (plus any adjustment
covariates) — a log-odds-ratio shift at quantile tau of the
proportion response, not a difference in means or in the raw-scale quantile.
Caches beta_hat_T (and, unless estimate_only, the standard error
and residual degrees of freedom) so repeated calls are cheap; returns
NA_real_ if the reduced design matrix is degenerate (fewer usable rows
than columns) or the quantreg fit fails/errors.
Usage
InferencePropQuantileRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferencePropQuantileRegr$compute_estimate_with_bootstrap_weights()
Recomputes compute_estimate's treatment log-odds-ratio-shift
coefficient with each subject's (or block's) contribution to the tau-quantile
fit reweighted by subject_or_block_weights (expanded to per-row weights and
passed as quantreg::rq(..., weights = ...)), for the Bayesian bootstrap
contract; see InferenceBayesianBootstrap.
Writes into the same beta_hat_T/s_beta_hat_T/df cache
fields that compute_estimate reads from — a call to this method overwrites
the cached original-data estimate with the bootstrap-reweighted one, so a subsequent
compute_estimate() call will return the bootstrap replicate's value
from cache rather than recomputing on the original data, until the cache is reset by
whatever higher-level bootstrap driver owns this object's lifecycle. Returns
NA_real_ under the same degenerate-design/fit-failure conditions as
compute_estimate.
Usage
InferencePropQuantileRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferencePropQuantileRegr$compute_asymp_confidence_interval()
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Usage
InferencePropQuantileRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
InferencePropQuantileRegr$compute_asymp_two_sided_pval()
Uses the shared asymptotic two-sided p-value contract; see
InferenceAsymp.
Usage
InferencePropQuantileRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. Default is zero.
InferencePropQuantileRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropQuantileRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Koenker, R. and Bassett, G. (1978). "Regression Quantiles."
Econometrica, 46(1), 33-50, doi:10.2307/1913643, for quantile regression itself.
Powell, J. L. (1991). "Estimation of Monotonic Regression Models under Quantile
Restrictions," in Nonparametric and Semiparametric Methods in Econometrics and
Statistics, Cambridge University Press, for the "nid" sandwich standard error.
See Also
InferenceContinQuantileRegr for the
untransformed (continuous-scale) analogue of this class.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropQuantileRegr$new(seq_des)
inf$compute_estimate()
Zero/One-Inflated Beta Inference for Proportion Responses
Description
Internal class for non-KK zero/one-inflated beta regression models. The
response is modeled as a three-component mixture with point masses at 0 and 1
plus a beta-distributed interior component on (0, 1). The reported
treatment effect is the treatment coefficient from the beta mean submodel, on
the logit scale, conditional on the response falling strictly inside
(0, 1): it is not, and should not be read as, the effect on the
unconditional mean E[Y], which also depends on how treatment shifts the
zero/one inflation probabilities. A design where treatment moves mass between
the point masses and the interior, with no shift in the interior beta mean,
will report a null treatment coefficient here even though E[Y] changed.
The beta mean submodel uses treatment alone in the univariate class and
treatment plus covariates in the multivariate class. The zero/one inflation
submodels use model_formula_zero_one, which defaults to ~ .
so that treatment plus all available covariates enter those auxiliary pieces.
See the class-level description above for what the reported coefficient does
and does not represent.
Marginal estimand. This class composes
MarginalEstimand
(set_estimand()/get_estimand()/get_supported_estimands());
in addition to the default "conditional" estimand described above, it
supports "marginal_mean_diff": the model-implied unconditional mean
recombines all three mixture components,
E[Y \mid x, w] = \pi_1(x, w) \cdot 1 + (1 - \pi_0(x, w) - \pi_1(x, w))
\cdot \mathrm{logit}^{-1}(x^\top \beta + \beta_T w) (the zero mass
contributes nothing), where \pi_0/\pi_1 are the normalized
zero/one-inflation mixture probabilities. The reported treatment effect
under "marginal_mean_diff" is the g-computation average
\hat\tau = n^{-1} \sum_i [\hat E(Y \mid x_i, w=1) - \hat E(Y \mid x_i,
w=0)], on the response's natural [0,1] scale (not the log-odds
scale of the conditional estimand). Standard errors are delta-method,
against the joint covariance of [\beta, \log\phi, \gamma_0, \gamma_1]
already returned by fast_zero_one_inflated_beta_cpp, using a
numerical (central-difference) gradient of \hat\tau — see
marginal_estimand_report.md → TODO-4 for why analytic
differentiation was not used. This is a pure post-fit transform of the same
cached maximum-likelihood fit (no refit), so only Wald-via-delta-method
inference is available under a marginal estimand (no likelihood-ratio/
score/gradient test — see set_estimand()'s testing-type
interaction).
Super class
Inference -> InferencePropZeroOneInflatedBetaRegr
Methods
Public methods
-
InferencePropZeroOneInflatedBetaRegr$compute_asymp_confidence_interval() -
InferencePropZeroOneInflatedBetaRegr$compute_asymp_two_sided_pval() -
InferencePropZeroOneInflatedBetaRegr$compute_estimate_with_bootstrap_weights()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferencePropZeroOneInflatedBetaRegr$new()
Initialize inference for the three-component zero/one-inflated
beta mixture model; see
InferencePropZeroOneInflatedBetaRegr
for the model form and the important caveat that the reported treatment
coefficient is conditional on the interior (0,1) component, not an
unconditional-mean effect. Does not fit the model; the fit is deferred to
the first call to compute_estimate() or a method that requires it.
Usage
InferencePropZeroOneInflatedBetaRegr$new( des_obj, model_formula = NULL, model_formula_zero_one = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a proportion response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.model_formula_zero_oneFormula for the zero/one inflation submodels. Defaults to
~ ., meaning treatment plus all available covariates.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferencePropZeroOneInflatedBetaRegr$compute_estimate()
Fits the zero/one-inflated beta mixture model by maximum
likelihood (jointly the beta mean submodel, the zero/one inflation
submodels, and the beta precision). Under the default
estimand = "conditional", returns \hat\beta_T, the
treatment log-odds-ratio from the beta mean submodel conditional
on the interior (0,1) component — see
InferencePropZeroOneInflatedBetaRegr's
estimand caveat. Under estimand = "marginal_mean_diff" (set via
set_estimand()), returns the g-computation marginal mean
difference instead — see the class-level @details for the
formula. The underlying model fit is identical either way (a pure
post-fit transform of the same cached fit, no refit).
Usage
InferencePropZeroOneInflatedBetaRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip standard-error computation and cache only the point estimate; used by randomization and bootstrap resampling paths.
InferencePropZeroOneInflatedBetaRegr$compute_asymp_confidence_interval()
Wald confidence interval, dispatched by
testing_type for the conditional estimand (score/gradient/
likelihood-ratio/Bartlett available; see
InferenceAsympLik); under a
marginal estimand testing_type is always "wald" (the
only value set_estimand() permits there), so this always
resolves to the delta-method interval. Calls
self$compute_estimate() first (not
private$shared() directly) so the estimand-aware cache is
always current regardless of call order.
Usage
InferencePropZeroOneInflatedBetaRegr$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaTwo-sided miscoverage rate; the returned interval targets
1 - alphacoverage.
InferencePropZeroOneInflatedBetaRegr$compute_asymp_two_sided_pval()
Wald two-sided p-value, dispatched by testing_type
exactly as compute_asymp_confidence_interval(); see that
method's description for the marginal-estimand always-Wald note.
Usage
InferencePropZeroOneInflatedBetaRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment-effect value under the current estimand (conditional log-odds-ratio, or marginal mean difference).
InferencePropZeroOneInflatedBetaRegr$compute_estimate_with_bootstrap_weights()
Refits the zero/one-inflated beta model with subject/block-level
weights applied to the fitting log-likelihood (Bayesian-bootstrap or
nonparametric-bootstrap draw weights, expanded to row level via
private$expand_subject_or_block_weights_to_row_weights()), and
returns the reweighted conditional log-odds-ratio estimate
\hat\beta_T^{(w)}.
Usage
InferencePropZeroOneInflatedBetaRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferencePropZeroOneInflatedBetaRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferencePropZeroOneInflatedBetaRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Ospina, R., and Ferrari, S. L. P. (2010). "Inflated beta distributions." Statistical Papers, 51(1), 111-126, doi:10.1007/s00362-008-0125-4, for the zero/one-inflated beta mixture density; Ferrari, S., and Cribari-Neto, F. (2004). "Beta regression for modelling rates and proportions." Journal of Applied Statistics, 31(7), 799-815, doi:10.1080/0266476042000214501, for the interior beta-regression submodel.
See Also
InferencePropBetaRegr
for the plain (non-inflated) beta regression model.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'proportion')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferencePropZeroOneInflatedBetaRegr$new(seq_des)
inf$compute_estimate()
Randomization-based Inference
Description
Abstract class for randomization-based inference.
Super class
Inference -> InferenceRand
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceRand$approximate_randomization_distribution_beta_hat_T()
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Usage
InferenceRand$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
Returns
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
InferenceRand$supports_rand_pval_for_incidence()
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Usage
InferenceRand$supports_rand_pval_for_incidence()
Returns
A single logical.
InferenceRand$compute_rand_two_sided_pval()
Computes a randomization-based p-value.
Usage
InferenceRand$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors.
deltaNull difference.
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
Returns
Randomization p-value.
InferenceRand$clone()
The objects of this class are cloneable with this method.
Usage
InferenceRand$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Bootstrap Randomization Test Inference
Description
Each of the B null draws is generated by (1) resampling n subject rows
with replacement from the observed data and (2) drawing one fresh assignment
vector w from the actual experimental design run on the resampled covariates.
The test statistic computed on each such draw forms a null distribution against which
the observed statistic (actual data, actual w) is compared.
Details
Abstract class implementing the bootstrap randomization test (BRT), a hybrid of the
nonparametric bootstrap and the randomization test of Fisher's sharp null. This
construction is inspired by Kallus, N. (2018), "Optimal a priori balance in the design
of controlled experiments," Journal of the Royal Statistical Society: Series B,
80(1), 85-112, Section 5 ("Algorithms for inference"), Algorithm 4 – see
@references below for the full attribution and what EDI's implementation adds
beyond it.
Motivating heuristic: the sharp null makes the science table fully known.
Under Fisher's sharp null H_0: y_i(0) = y_i(1) = y_i for all i, the observed
outcome is the outcome regardless of assignment, so each observed row (x_i, y_i)
is a complete description of that subject: it can be paired with any w and the
outcome is still correct. Resampling rows and drawing a fresh w therefore
simulates an entire new experiment — new subjects from \hat{F}_n (the empirical
distribution of subjects), new assignment from the true, known assignment mechanism —
under the assumption that outcomes do not respond to treatment. Both simulation
ingredients are faithful to the real data-generating process under H_0: the
assignment mechanism is exact (drawn from the actual design, including sequential
covariate-dependent designs, since the design depends only on the covariates), and the
subject distribution is approximately right by the usual bootstrap argument
\hat{F}_n \to F. This is a motivating heuristic, not a proof — see "Known
limitations and open theoretical question" below for exactly where it stops short of
establishing asymptotic validity, and why. If H_0 is false, the construction
still generates the null distribution (it forcibly treats y as invariant to
w) while the observed statistic drifts into the tail — that asymmetry is the
intended source of power, independent of the open validity question below.
Why this outgrew its original motivation. This construction was originally
implemented for a narrow reason: it works even when w is deterministic or
near-deterministic, exactly the degenerate case Kallus (2018) needed it for (see
@references). It turns out to be useful well beyond that: (1) it is
design-agnostic for free, since it calls draw_ws_according_to_design() on the
resampled data rather than enumerating a permutation space – for the matching-on-the-fly
designs (DesignSeqOneByOneKK14/KK21/KK21stepwise), that permutation
space depends on the whole sequential arrival/matching history and has no closed form
worth enumerating, so a classic randomization test would need bespoke combinatorics this
construction avoids entirely; (2) it generalizes the inferential target from
finite-population (conditional on exactly these n subjects) to superpopulation,
independent of the degenerate-design motivation; (3) it inherits the rest of the
inference machinery (CI inversion, the estimand/testing-type axes) automatically by
living in the same class hierarchy as everything else, rather than being a bolted-on
special case usable only for the one scenario that motivated it.
How it differs from the pure randomization test. The classic randomization
test conditions on the realized sample (the "science table") and is exact for any
n. The BRT is unconditional: it targets the distribution of the statistic over
both subject sampling and randomization, and the null tested is the compound "sharp
null AND subjects are i.i.d. draws from F". Exactness is traded for a
population-level (superpopulation) interpretation. Equivalently, the BRT is a
parametric bootstrap test whose "parameter" is the pair (F, design): the design
part is known exactly, F is plugged in via \hat{F}_n, and the sharp null is
precisely what makes \hat{F}_n estimable from one arm-agnostic dataset. The pure
randomization test is the special case where F is conditioned away.
Known limitations and open theoretical question: asymptotic validity is not proven, for two distinct, identifiable reasons. Kallus (2018) himself states that asymptotic validity of this bootstrap test is "believed" rather than established, and explicitly calls it an open question. EDI's implementation does not resolve that question; the two specific gaps are:
-
Glivenko-Cantelli is not bootstrap consistency. The motivating heuristic above leans on
\hat{F}_n \to Fto justify substituting a bootstrap resample for a genuinely fresh i.i.d. draw fromF. But\hat{F}_n \to F(uniform convergence of the empirical CDF, a Glivenko-Cantelli statement) is a claim about the empirical distribution itself; it is not the same as consistency of the bootstrap, i.e. convergence of the sampling distribution of the test statistic computed under resampling from\hat{F}_nto the statistic's true sampling distribution underF. That second, load-bearing claim is well known to require statistic-specific regularity conditions (smoothness/Hadamard differentiability of the statistic as a functional of the empirical process, adequate moment conditions, etc. — the classical Bickel-Freedman-type distinction) and is known to fail for some statistics even when\hat{F}_n \to Fholds trivially (e.g. non-smooth statistics, extreme-value/max-type statistics, heavy-tailed distributions without enough moments, non-regular estimators). No such regularity condition is verified, case by case, for the estimators this class is composed into. -
The naive version of this concern is already mitigated by design; a narrower residual question remains open. A naive i.i.d. row bootstrap would create exact duplicate rows that never occur in real data, and for a design whose assignment mechanism depends on the joint covariate configuration of the whole sample — exactly what EDI's matching-on-the-fly designs (
DesignSeqOneByOneKK14/KK21/KK21stepwise) and fixed matched-pair designs (DesignFixedBinaryMatch) do — those artificial exact ties would fabricate zero-distance ("perfect") matches that a genuinely fresh sample fromFwould essentially never produce. This package does not resample rows i.i.d. for matching-capable designs.DesignMatchingAbstract$draw_bootstrap_indices()(design_matching_abstract.R), inherited by every matching-capable design, dispatches to a pair-aware resampler (private$draw_matching_bootstrap_indices()) that resamples reservoir subjects i.i.d. and matched pairs as intact units — preserving each pair's true, historically realized within-pair covariate distance rather than fabricating an artificial zero-distance one. This is the correct fix for the naive-duplicate-row concern, and it is already active for every BRT draw on a matching-capable design (InferenceRandBootstrapinheritsbootstrap_sample_indices()from the same chain, so no separate wiring is needed). What remains open, narrower than the naive concern above: reservoir subjects are still resampled i.i.d. independently of the intact pairs, so a single reservoir subject can still appear more than once in one bootstrap draw; whether two duplicate copies of the same original reservoir subject can subsequently be matched to each other by the re-run sequential matching algorithm (fabricating a same-subject zero-distance pair as a second-order effect, distinct from the naive first-order concern this mitigation addresses) has not been analyzed in this package. This is a narrower, unquantified residual question, not a demonstrated bias.
Status: point 1 above is a documented open theoretical question, not a settled result; point 2's naive form is mitigated by the pair-aware resampler described above, with only the narrower residual question left open. Neither this package nor Kallus (2018) supplies a general proof of asymptotic validity covering point 1; both explicitly flag it as unresolved rather than claiming a proof. Treat the p-values and confidence intervals from this class as resting on a well-motivated but unproven asymptotic argument. A rigorous asymptotic proof (or a demonstrated counterexample) remains future work.
Other implementation notes. (1) For sequential designs the bootstrap sample needs an arrival order; the i.i.d. resampling order plays that role, matching the i.i.d.-arrivals assumption of the KK designs. (2) Studentization is not needed for the motivating heuristic above but, as with any bootstrap test, an asymptotically pivotal statistic improves the level's rate of convergence where the construction is valid.
Users do not instantiate this class directly: every concrete inference class in the
package inherits from it, so its methods (compute_rand_bootstrap_two_sided_pval,
approximate_rand_bootstrap_distribution_beta_hat_T) are available on any
inference object. See InferenceRandBootstrapCI for the companion
confidence interval.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap
Methods
Public methods
-
InferenceRandBootstrap$get_supported_rand_bootstrap_pval_types() -
InferenceRandBootstrap$approximate_rand_bootstrap_distribution_beta_hat_T() -
InferenceRandBootstrap$compute_rand_bootstrap_two_sided_pval()
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceRandBootstrap$get_supported_rand_bootstrap_pval_types()
Returns the type values
compute_rand_bootstrap_two_sided_pval() accepts.
Usage
InferenceRandBootstrap$get_supported_rand_bootstrap_pval_types()
InferenceRandBootstrap$approximate_rand_bootstrap_distribution_beta_hat_T()
Computes the bootstrap randomization null distribution of the test
statistic under Fisher's sharp null (shifted by delta): each draw resamples
subject rows with replacement and draws one fresh assignment vector from the design.
Usage
InferenceRandBootstrap$approximate_rand_bootstrap_distribution_beta_hat_T( B = 501, delta = 0, transform_responses = "none", show_progress = TRUE, debug = FALSE, bootstrap_type = NULL, rand_bootstrap_draws = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
BNumber of bootstrap randomization draws. Default 501.
deltaThe null treatment effect (on the
transform_responsesscale). Default 0.transform_responsesType of response transformation used to impose the sharp null shift for nonzero
delta. Default "none".show_progressA flag indicating whether a progress bar should be displayed.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.bootstrap_typeOptional bootstrap-resampling scheme; see
approximate_bootstrap_distribution_beta_hat_Tfor legal values. DefaultNULL.rand_bootstrap_drawsOptional pre-generated draws (as returned by the private method
generate_rand_bootstrap_draws) enabling common random numbers across calls with differentdeltavalues (used by the CI inversion). DefaultNULL.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging.
Returns
When debug = FALSE (default), a numeric vector of length B
containing the null-distribution draws. When debug = TRUE, a list with:
values, errors, warnings, num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
InferenceRandBootstrap$compute_rand_bootstrap_two_sided_pval()
Computes a bootstrap randomization two-sided p-value for Fisher's sharp
null (shifted by delta): the observed statistic (actual data, actual w)
is compared against the null distribution generated by resampling rows and drawing
fresh assignments from the design.
Usage
InferenceRandBootstrap$compute_rand_bootstrap_two_sided_pval( B = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, bootstrap_type = NULL, rand_bootstrap_draws = NULL, zero_one_logit_clamp = .Machine$double.eps, type = "percentile" )
Arguments
BNumber of bootstrap randomization draws. Default 501.
deltaThe null treatment effect. Default 0.
transform_responsesType of response transformation for the sharp null shift. The default "none" resolves by response type (logit for proportion, log for count and survival, identity otherwise), matching
compute_rand_two_sided_pval.na.rmRemove non-finite null draws. Default
TRUE.show_progressA flag indicating whether a progress bar should be displayed.
bootstrap_typeOptional bootstrap-resampling scheme; see
approximate_bootstrap_distribution_beta_hat_Tfor legal values. DefaultNULL.rand_bootstrap_drawsOptional pre-generated draws for common random numbers across
deltavalues. DefaultNULL.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging.
typeTest statistic type.
"percentile"(default) uses the raw estimator as the BRT test statistic, giving an asymptotic p-value that inherits the unconditional superpopulation validity of the BRT."studentized"divides each null draw's signed deviation from the null by its per-draw standard error:p = 2\min(P(z^0_b \ge z), P(z^0_b \le z))wherez^0_b = (t^0_b - \delta)/\hat{s}^0_b; yields asymmetric CI under inversion."symmetric-percentile-t"uses the absolute pivotp = P(|t^0_b - \delta|/\hat{s}^0_b \ge |t - \delta|/\hat{s}); CI is symmetric. Both SE-based types require the class to exposes_beta_hat_T; fall back to"percentile"if the SE is unavailable."smoothed"adds kernel noise\varepsilon_b \sim N(0, \hat{\sigma}/\sqrt{n})to each resampled draw before imposing the null shift, reducing discreteness in the null distribution. Only meaningful for continuous responses. For count responses the noisy draw is rounded and floored at zero so it stays on the non-negative integer support the Poisson-family likelihoods require; at the default bandwidth this makes the smoothing nearly a no-op for low counts.Theoretical justification. Order-statistic/rank-based estimators (e.g. the Hodges-Lehmann pseudo-median) take only finitely many values, so their bootstrap/ randomization null distribution is a step function; this both coarsens p-values and destabilizes the
uniroot-baseddeltasearch inInferenceRandBootstrapCI'scompute_rand_bootstrap_confidence_interval, which assumes an approximately continuous, monotone p-value curve. Restoring continuity by convolving the resampling distribution with a shrinking-bandwidth kernel is the classical "smoothed bootstrap" device: Silverman (1981, "Density ratios, empirical likelihood and cot death", Applied Statistics 30(2):142-145) for kernel smoothing of a resampled empirical distribution; Silverman & Young (1987, "The bootstrap: To smooth or not to smooth?", Biometrika 74(3):469-479) for applying that smoothing directly to the bootstrap resampling scheme; and Hall, DiCiccio & Romano (1989, "On smoothing and the bootstrap", Annals of Statistics 17(2):692-704) for the bandwidth conditions under which smoothing improves, rather than degrades, coverage accuracy.Implementation caveat (ad hoc, not a certified instance of the above). The bandwidth used here,
\hat{\sigma}/\sqrt{n}(the raw SE-of-the-mean scale of the response), is a pragmatic engineering choice, not one derived from or validated against the bandwidth-selection results in the sources above. Those references smooth the resampling distribution itself with a bandwidth chosen to trade off bias against variance (often shrinking slower thann^{-1/2}, e.g.n^{-1/5}-type KDE rates); here, noise is instead added directly to each already-resampled response at a fixedn^{-1/2}rate. The bandwidth is not exposed as a parameter, has no zero-noise escape hatch (the only way to disable smoothing is to pick a differenttype), and its coverage behavior has not been validated by simulation in this package. Treat it as a discreteness patch that is qualitatively motivated by the literature above, not a certified implementation of it.Performance. Every class with a C++
compute_fast_rand_bootstrap_distrkernel that operates on real-valued responses (Wilcox HL, simple mean difference, OLS, robust regression, CoxPH, Weibull marginal, log-rank, RMST, KM-diff) accepts the smoothing noise directly in its kernel. This only speeds upInferenceRandBootstrapCI'scompute_rand_bootstrap_confidence_interval(type = "smoothed"), since CI inversion pre-materializes one set of fresh assignments up front (common random numbers reused across everydeltaevaluated during root-finding) and the fast kernel can engage on each evaluation; measured onInferenceAllSimpleWilcox,n = 30,B = 99: the forced-slow-fallback CI took 25.5s, the fast-kernel CI took 0.5s (about 50x). A standalonecompute_rand_bootstrap_two_sided_pval(type = "smoothed")call (no CI inversion) is not accelerated by this: it deliberately draws the fresh assignment lazily per replicate (materialize_w = FALSE), sorand_bootstrap_draw_matrices()cannot build the matrices the fast kernels need and the R-level fallback still runs regardless of this fix. The two ordinal classes (InferenceOrdinalRidit,InferenceOrdinalJonckheereTerpstraTest) still use the slower R-level fallback in every case, because adding continuous Gaussian noise to integer category codes is not statistically meaningful — see the response-type caveat above.
Returns
A two-sided p-value.
InferenceRandBootstrap$clone()
The objects of this class are cloneable with this method.
Usage
InferenceRandBootstrap$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kallus, N. (2018), "Optimal a priori balance in the design of controlled
experiments," Journal of the Royal Statistical Society: Series B, 80(1), 85-112,
Section 5 ("Algorithms for inference"), Algorithm 4 – the nearest prior appearance of
this exact construction (resample subjects, then redraw one fresh assignment from the
design on the resampled covariates), introduced there for the degenerate case of his a
priori balancing designs (PSODs), where a design can admit only one or very few
distinct treatment permutations, so the classic Fisher randomization test has no power
(it always returns p-value 1). Kallus himself cites Efron, B. and Tibshirani, R.
(1993), An Introduction to the Bootstrap, Chapman & Hall, for the bootstrap
ingredient, and Good, P. (2005), Permutation, Parametric and Bootstrap Tests of
Hypotheses, Springer, for the test/confidence-interval duality used to invert his
Algorithm 4 into intervals – and states asymptotic validity of the bootstrap test as
an open question, not a proven result (see "Known limitations and open theoretical
question" above, which this package inherits and extends with the matching-design-
specific failure mode). EDI's contribution beyond Algorithm 4 is generalizing the
construction to arbitrary designs (fixed and sequential, including the matching-on-
the-fly family) and response types, and pairing it with a confidence-interval
inversion (InferenceRandBootstrapCI); it does not resolve the open
validity question.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceAllSimpleAverageDiff$new(seq_des)
seq_des_inf$compute_rand_bootstrap_two_sided_pval(B = 101)
Bootstrap Randomization Confidence Intervals
Description
The CI is the set of delta values whose bootstrap randomization p-value exceeds
alpha. All p-value evaluations across candidate delta values share one set
of pre-generated draws (resampled row indices and fresh design assignments) — common
random numbers — so the p-value is a deterministic, near-monotone function of
delta and the bound search is stable. See InferenceRandBootstrap
for the statistical justification of the test being inverted; the resulting interval
inherits its unconditional, superpopulation interpretation and asymptotic validity.
Users do not instantiate this class directly: every concrete inference class in the
package inherits from it, so compute_rand_bootstrap_confidence_interval is
available on any inference object.
Details
Abstract class implementing confidence intervals by inverting the bootstrap
randomization test of InferenceRandBootstrap over the null effect delta.
Super classes
Inference -> InferenceRand -> InferenceRandCI -> InferenceNonParamBootstrap -> InferenceRandBootstrap -> InferenceRandBootstrapCI
Methods
Public methods
-
InferenceRandBootstrapCI$get_supported_rand_bootstrap_ci_types() -
InferenceRandBootstrapCI$compute_rand_bootstrap_confidence_interval()
+ inherited public methods from InferenceRandBootstrap
+ inherited public methods from InferenceNonParamBootstrap
InferenceNonParamBootstrap$approximate_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_m_out_of_n_bootstrap_distribution_beta_hat_T()InferenceNonParamBootstrap$approximate_subsampling_distribution_beta_hat_T()InferenceNonParamBootstrap$compute_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_confidence_interval()InferenceNonParamBootstrap$compute_m_out_of_n_bootstrap_two_sided_pval()InferenceNonParamBootstrap$compute_subsampling_confidence_interval()InferenceNonParamBootstrap$compute_subsampling_sensitivity()InferenceNonParamBootstrap$compute_subsampling_two_sided_pval()InferenceNonParamBootstrap$get_supported_bootstrap_ci_types()InferenceNonParamBootstrap$get_supported_bootstrap_pval_types()InferenceNonParamBootstrap$select_optimal_b_subsampling()InferenceNonParamBootstrap$select_optimal_m_out_of_n_bootstrap()
+ inherited public methods from InferenceRandCI
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceRandBootstrapCI$get_supported_rand_bootstrap_ci_types()
Returns the type values
compute_rand_bootstrap_confidence_interval() accepts.
Usage
InferenceRandBootstrapCI$get_supported_rand_bootstrap_ci_types()
InferenceRandBootstrapCI$compute_rand_bootstrap_confidence_interval()
Computes a confidence interval by inverting the bootstrap randomization
test over the null effect delta. For statistics that are affine in the
additive sharp-null shift (e.g. the simple mean difference and the OLS treatment
coefficient), the inversion is performed in closed form from the breakpoints of the
p-value step function — exact given the draws and requiring no bisection. Otherwise,
the generic bisection search is used. When the p-value does not drop below
alpha / 2 anywhere within the search radius, a conservative bound is returned
at the search boundary rather than NA; each such event emits a
message() and increments the private field
rand_bootstrap_ci_conservative_count.
Usage
InferenceRandBootstrapCI$compute_rand_bootstrap_confidence_interval( alpha = 0.05, B = 501, pval_epsilon = 0.005, show_progress = TRUE, max_expansions = 7L, bootstrap_type = NULL, zero_one_logit_clamp = .Machine$double.eps, type = "percentile" )
Arguments
alphaThe confidence level 1 -
alpha. Default 0.05.BNumber of bootstrap randomization draws. Default 501.
pval_epsilonBisection tolerance (on both the
deltabracket width and the p-value span). Default 0.005.show_progressA flag indicating whether progress should be displayed.
max_expansionsMaximum number of bound-doubling expansions when the seed interval does not bracket the target p-value. Default 7.
bootstrap_typeOptional bootstrap-resampling scheme; see
approximate_bootstrap_distribution_beta_hat_Tfor legal values. DefaultNULL.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging.
typeCI type.
"percentile"(default) inverts the raw BRT p-value via bisection (or closed-form for affine statistics)."studentized"inverts the signed-pivot BRT p-valuep(\delta) = 2\min(P(z^0_b \ge z), P(z^0_b \le z))wherez^0_b = (t^0_b(\delta) - \delta)/\hat{s}^0_b; yields asymmetric CI."symmetric-percentile-t"inverts the absolute-pivot versionp(\delta) = P(|t^0_b - \delta|/\hat{s}^0_b \ge |t - \delta|/\hat{s}); yields a CI symmetric around the observed estimate. Both SE-based types pre-compute\hat{s}^0_bonce at\delta = 0and reuse it across bisection steps. Yield O(n^{-1}) coverage error versus O(n^{-1/2}) for"percentile"when the pivot is asymptotically normal. Require the class to expose a standard error; returnNAbounds in harden mode if unavailable."smoothed"adds per-draw kernel noise\varepsilon_b \sim N(0, \hat{\sigma}/\sqrt{n})to the resampled responses before imposing the null shift, reducing discreteness. Only meaningful for continuous responses; count responses are rounded and floored at zero after the noise so the Poisson-family refits stay on the integer support. SeeInferenceRandBootstrap'scompute_rand_bootstrap_two_sided_pvalfor the theoretical justification (Silverman 1981; Silverman & Young 1987; Hall, DiCiccio & Romano 1989), an explicit caveat that the bandwidth used here (\hat{\sigma}/\sqrt{n}, fixed, unexposed, with no zero-noise escape hatch) is a pragmatic ad hoc choice rather than one derived from or validated against those sources, and measured performance (this CI, unlike a standalone smoothed p-value, is accelerated by the noise-aware fast kernels: about 50x onn = 30,B = 99).
Returns
A bootstrap randomization confidence interval. The interval lives on the response-transformation scale used by the test (identity for continuous, logit for proportion, log for count and survival). Bounds may be conservative (wider than necessary) when the p-value inversion cannot be completed within the search radius.
InferenceRandBootstrapCI$clone()
The objects of this class are cloneable with this method.
Usage
InferenceRandBootstrapCI$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 20, response_type = "continuous")
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
seq_des_inf = InferenceAllSimpleAverageDiff$new(seq_des)
seq_des_inf$compute_rand_bootstrap_confidence_interval(alpha = 0.05, B = 101)
Randomization-based Confidence Intervals
Description
Abstract class for randomization-based confidence interval inference.
Super classes
Inference -> InferenceRand -> InferenceRandCI
Methods
Public methods
+ inherited public methods from InferenceRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceRandCI$compute_rand_two_sided_pval()
Compute a randomization-based two-sided p-value for the treatment effect.
Usage
InferenceRandCI$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, type = NULL, args_for_type = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors.
deltaNull treatment effect value.
transform_responsesResponse transformation to apply during the test. For survival responses the default
"log"multiplies the recorded times of the units treated under each reference allocation bye^\delta, event and censoring times alike, with censoring indicators unchanged – the rank-based AFT residual construction (Tsiatis 1990; Wei, Ying and Lin 1990; Jin, Lin, Wei and Ying 2003); seecompute_rand_confidence_interval()for the assumptions.na.rmWhether to remove non-finite simulated statistics.
show_progressWhether to show progress.
permutationsOptional pre-generated assignment draws.
typeOptional incidence-specific exact randomization type.
args_for_typeOptional arguments keyed by
type.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
Returns
A two-sided p-value.
InferenceRandCI$compute_rand_confidence_interval()
Computes a randomization-based confidence interval by inverting the
randomization test: the interval is the set of \delta for which the two-sided
randomization p-value of the sharp null "treatment effect = \delta" is at
least alpha. For each candidate \delta the control potential
outcomes are first imputed under that sharp null by removing the hypothesised
effect from the units that were actually treated; then, for each reference
allocation w_b, the effect is applied to the units treated under
w_b — on the response type's scale (additive for continuous responses,
multiplicative e^\delta for counts and survival times, a logit shift for
proportions) — the estimator is recomputed, and the observed estimate is
compared with the resulting reference distribution. This is the
impute-then-permute construction of Rosenbaum (2002, ch. 2) and Imbens and
Rubin (2015, ch. 5): y_{sim} = y - \delta w_{obs} + \delta w_b on the
additive scale.
Survival responses. The sharp null is an accelerated-failure-time
(AFT) effect: for units treated under w_b the recorded time is multiplied
by e^\delta — both event times and censoring times — and the
censoring indicator is carried over unchanged. Equivalently, the test is run on
the residual times \log y_i - \delta w_i of every observation, censored
or not. This is the residual construction underlying rank-based inference for
the AFT model: Tsiatis (1990) shows that linear rank statistics computed on
these residuals have mean zero at the true \delta under independent
censoring, even though the residual censoring distribution then depends on
treatment; Wei, Ying and Lin (1990) invert exactly this family of tests to
obtain confidence intervals for AFT regression coefficients; Jin, Lin, Wei and
Ying (2003) give the modern estimation and inference machinery for the same
model. The alternative of rescaling only the event times and re-deriving the
censoring indicator is not identifiable, because a unit's censoring time is
unobserved whenever its event was. Two assumptions therefore apply: (i)
independent censoring (C \perp T \mid w) for the asymptotic validity of
the inverted test, and (ii) for the finite-sample exactness of the
permutation version specifically, that censoring times are on the same
accelerated clock as event times (C_i(1) = e^\delta C_i(0)), which is
plausible for health-driven dropout and does not hold for calendar-time
administrative censoring; under administrative censoring the interval is
asymptotically, not exactly, valid. Because \delta is a log time-ratio,
this construction is coherent only for classes whose estimand is on that scale
(the Weibull AFT, marginal Weibull, Weibull-frailty, and rank-regression
classes); classes whose estimand is a log hazard ratio cannot invert an AFT
shift without a parametric link between the two scales and refuse this method
(the six classes in EDI_LOG_HAZARD_RATIO_INFERENCE_CLASSES — the Cox
family — which also do not advertise the randomization_ci capability, so
InferenceSuite never offers it for them; their randomization p-value and
randomization-bootstrap CI are unaffected).
Usage
InferenceRandCI$compute_rand_confidence_interval( alpha = 0.05, r = 501, pval_epsilon = 0.005, show_progress = TRUE, type = NULL, args_for_type = NULL, ci_search_control = NULL )
Arguments
alphaSignificance level.
rNumber of randomization vectors.
pval_epsilonBisection tolerance.
show_progressShow progress.
typeOptional incidence-specific exact randomization type.
args_for_typeOptional arguments keyed by
type.ci_search_controlOptional control list for randomization-CI search. Supported entries are
fallback,seed,max_radius_se_mult(default 25),max_radius_scale_mult(default 6),max_expansions(default 7),seed_boot_B, Monte Carlo settingsmc_enable,mc_batch_size,mc_min_draws, andmc_conf_level, midpoint-cache settingspval_cache_enableandpval_cache_resolution, and model-fit reuse settingsfit_warm_start_enableandfit_reuse_factorizations. Setmc_enable = FALSEto force full enumeration of all requested randomization draws. The search radius ismax(max_radius_se_mult * se_guess, max_radius_scale_mult * sd(y)). When the randomization p-value does not drop belowalpha/2anywhere within the search radius (e.g. the test has low power or the design has few unique permutations), a conservative CI bound is returned at the search boundary rather thanNA. This guarantees a valid (though possibly wide) interval. Each such event emits amessage()and increments the private fieldrand_ci_conservative_countfor monitoring.high_precision_confirm(defaultTRUE) re-checks each converged bound with one full-enumeration (no early stopping) p-value evaluation and, only if that disagrees with the early-stopped bisection's conclusion, spends a short additional full-precision re-bisection to correct it – a final high-precision confirmation pass that catches sequential-Monte-Carlo early-stopping noise the cheap bisection has no way to notice on its own; seehigh_precision_confirm_and_refine_ci_bound()'s own comment for the full rationale andR/package_metadata/new_feature_plans/for the deeper, not-yet-implemented redesign (anytime-valid confidence sequences) this pass is a pragmatic stopgap for.
Returns
Randomization CI. Bounds may be conservative (wider than necessary) when the
p-value inversion cannot be completed within the search radius; see
ci_search_control for details.
InferenceRandCI$clone()
The objects of this class are cloneable with this method.
Usage
InferenceRandCI$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Tsiatis, A. A. (1990). Estimating regression parameters using linear rank tests for censored data. The Annals of Statistics, 18(1), 354-372, doi:10.1214/aos/1176347504. Wei, L. J., Ying, Z., and Lin, D. Y. (1990). Linear regression analysis of censored survival data based on rank tests. Biometrika, 77(4), 845-851, doi:10.1093/biomet/77.4.845. Jin, Z., Lin, D. Y., Wei, L. J., and Ying, Z. (2003). Rank-based inference for the accelerated failure time model. Biometrika, 90(2), 341-353, doi:10.1093/biomet/90.2.341.
Randomization test/CI on a user-supplied statistic
Description
Runs a randomization test (and, via compute_rand_confidence_
interval(), a randomization confidence interval) for an arbitrary
user-supplied statistic, without writing an Inference subclass.
The statistic function is called once per permutation as
fn(y, w, dead) – plain numeric/integer vectors: the current
(possibly permuted) response, treatment assignment, and event indicator.
It must return one scalar. For maximum speed, supply
custom_randomization_statistic_cpp instead: C++ source defining a
function of (NumericVector y, IntegerVector w) or
(NumericVector y, IntegerVector w, IntegerVector dead) returning a
scalar double, using the same convention as
DesignFixedOptimal's custom_objective.
Only a randomization test and randomization confidence interval are available – there is no package point-estimator, Wald path, or bootstrap machinery on this class, so no other action needs disabling.
Super classes
Inference -> InferenceCustomRand -> InferenceRandCustom
Methods
Public methods
+ inherited public methods from InferenceCustomRand
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceRandCustom$new()
Initialize.
Usage
InferenceRandCustom$new( des_obj, custom_randomization_statistic_function = NULL, custom_randomization_statistic_cpp = NULL, verbose = FALSE )
Arguments
des_objSee class description.
custom_randomization_statistic_functionSee class description.
custom_randomization_statistic_cppSee class description.
verboseSee class description.
InferenceRandCustom$fit()
Calls the user-supplied statistic on the observed data.
Usage
InferenceRandCustom$fit(estimate_only = FALSE)
Arguments
estimate_onlyUnused; present for the 'fit()' contract.
InferenceRandCustom$clone()
The objects of this class are cloneable with this method.
Usage
InferenceRandCustom$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Inference Suite: Discover and Bundle Every Applicable Inference Class for a Design
Description
A lightweight coordinator (not itself an Inference subclass, and not
part of the Inference R6 hierarchy) that pairs a single completed
Design object with the full set of concrete
Inference classes compatible with it. On construction, the suite
consults package-level inference metadata to discover every exported,
non-abstract Inference subclass whose declared response-type,
matched-design (KK), blocking, and censoring requirements are all satisfied
by des_obj, storing the resulting sorted class-name vector in
applicable_design_classes. Because discovery is driven by metadata
lookups rather than by actually attempting to construct each candidate
class, the applicable list automatically stays current as new inference
classes are registered elsewhere in the package, without this class needing
any changes, and without risking side effects (e.g. an optional-package
load failure inside some class's constructor) from a doomed construction
attempt.
Compatibility rules (see
is_inference_class_compatible_with_design_metadata(), also used by
Design's own
applicable_inference_class_names()): a candidate class is
excluded if it is abstract or not exported; if it declares no compatible
response types, or none match des_obj's response type; if its name
contains "KK" (a matched-design-only class) but des_obj does
not support KK matching; if the class's requires_blocking_design()
is TRUE (currently only InferenceIncidExtendedRobins –
InferenceIncidCMH works on both blocking and non-blocking designs
via different standard-error estimators, so it is not excluded)
but des_obj does not support blocking; or if des_obj has any
left-/interval-censored subjects (a finite y_R) but the class's
supports_interval_or_left_censored_data() is FALSE –
both of the latter two mirror Inference$initialize()'s own
construction-time gate exactly, via each class's registered
requires_blocking_design/supports_general_censoring
metadata (see infer_inference_requires_blocking_design()/
infer_inference_supports_general_censoring() in
inference_class_registry.R). Ordinary right-censoring alone never
excludes a class. The KK-name-pattern rule is still hardcoded in this
class rather than looked up from a central registry.
A class that is design-compatible but whose registered
required_packages are not all installed is excluded from
applicable_design_classes and reported separately, by class name,
in unavailable_due_to_missing_packages – so callers can tell "not
applicable to this design" apart from "applicable, but an optional
dependency isn't installed" (see the Discovery section of
fix_inference_hierarchy.md). Package availability is never a
reason a class is treated as design-incompatible.
Construction itself does not compute any estimates, p-values, or
confidence intervals – it only discovers and validates which inference
classes are applicable and does not eagerly construct any of them. This
class's run_all_inference() method does that: it constructs and
fits every applicable class and returns a uniform
comparison across them (see that method's own documentation for the
output schema, the CI/p-value method selection policy, and the
screen/html/plots/pdf/
save_results_as_JSON output options). lock_objects = FALSE
allows ad hoc fields to be attached to an instance after construction.
Every row this class discovers and fits is a test about the
same outcome variable, by construction. response_type is a
required, immutable constructor argument on Design
(read-only thereafter via get_response_type()), and this class
discovers every candidate in applicable_design_classes from one
attached Design object's one response_type – there is no
code path here that spans two response types in a single instance. This
matters beyond bookkeeping: a planned comparison-across-classes feature
(a single combined-evidence p-value summarizing every row, via a
dependence-robust combination test) relies on every combined class
sharing one sharp null of "no treatment effect on this outcome" –
which only holds when every test concerns the same outcome variable
(combining, say, a survival model's p-value with an unrelated
continuous-outcome model's p-value would not be valid, since a real
effect on one with none on the other is entirely plausible). That
precondition is guaranteed here structurally, not by caller discipline.
Same-Y does not mean same estimand, and that matters for
interpretation. Rows discovered here routinely test genuinely different
\theta_i on the same outcome – a mean difference, a log-odds
ratio under one link, a rank-based effect, a quantile shift – each
\theta_i its own parameterization's own null. The combined
H_0: \theta_1 = 0 \cap \dots \cap \theta_k = 0 is one coherent
claim (Fisher's unit-level sharp null: no treatment effect
whatsoever) only under a randomization-based/exact procedure,
where every possible summary of "effect" is simultaneously zero by
construction. Asymptotic/likelihood-based procedures instead test a
weak, population-level null specific to their own
parameterization, and weak nulls for genuinely different summaries are
not generally nested – treatment can shift a distribution's variance or
a high quantile while leaving its mean exactly unchanged, so "mean
difference = 0" does not imply "quantile effect = 0" outside
location-shift-style models. This is comparatively safe across
different link functions for the same latent effect (e.g.
cauchit/probit/cloglog/logit on the same binary/ordinal response, which
typically share one underlying latent-variable null) and comparatively
riskier across different kinds of summary (a mean-difference test
combined with a rank-based test combined with a quantile-regression
test). Under that weak-null reading, a rejection is more honestly read
as "at least one specific summary of this outcome's distribution
differs" rather than "there is one coherent nonzero effect" – reinforcing,
not loosening, the interpretation caveat below.
Combined Evidence interpretation caveat (read before using
combined_evidence$pval): the Cauchy combination test is a
union-intersection test of H_0: \theta_1 = 0 \cap \theta_2 = 0 \cap
\dots \cap \theta_k = 0 against the alternative that at least
one \theta_i \neq 0. A significant combined_evidence$pval
is therefore evidence of an effect in at least one of these
senses, not evidence for a specific estimate or direction – it does
not say which class's estimand is nonzero, nor does a small combined
p-value imply every (or even most) constituent p-values were small. Do
not report combined_evidence$pval as if it estimated a single
effect size, and do not treat it as validating any one class's estimate
over another's; its only valid use is as evidence that *some* legitimate
way of looking for a treatment effect on this outcome found one.
This guarantee assumes every constituent p_i is itself valid (i.e.
uniform under its own null) – the Cauchy combination's dependence-robust
size control says nothing about which such p_i is small,
so a single miscalibrated or misspecified procedure (e.g. an asymptotic
approximation that breaks down for this sample) can dominate the combined
result the same way it would dominate a plain minimum-p-value test, even
when every other procedure shows nothing. A significant
combined_evidence$pval is worth cross-checking against the
per-estimand CI-forest plot (plots$ci_forest) before trusting it –
one outlying interval sitting apart from a cluster of concordant ones is a
miscalibration flag, not confirmed evidence.
This same-Y precondition is guaranteed within one
InferenceSuite instance structurally (one Design, one
response_type), but the architecture cannot stop a caller from
manually combining raw pvals pulled from two separate
InferenceSuite objects' results_tables outside
cct_combine_pvalues() – doing so is outside this function's
validity guarantee and is not a supported use of this metric.
Public fields
applicable_design_classesCharacter vector of applicable inference class names derived during initialization.
unavailable_due_to_missing_packagesA named list, one entry per otherwise-design-compatible class whose registered
required_packagesare not all installed: names are class names, values are the character vector of missing package names. These classes are excluded fromapplicable_design_classesbut are reported here separately from plain design/response-type incompatibility (see the class-level docs' "Discovery" rules), so callers can distinguish "not applicable to this design" from "applicable, but an optional dependency isn't installed." Empty named list if every design-compatible class has all required packages available.
Methods
Public methods
InferenceSuite$new()
Discover every Inference class applicable to des_obj
(see the class-level documentation for the compatibility rules) and validate
any per-class constructor overrides in inference_params, storing
applicable_design_classes for later use. This constructor does
not instantiate any Inference object itself; callers are
expected to construct the specific classes they need (optionally passing the
validated inference_params) from the discovered list.
Usage
InferenceSuite$new(des_obj, model_formula = NULL, inference_params = list())
Arguments
des_objA completed
Designobject (validated viais(des_obj, "Design")when assertions are enabled; seetoggle_asserts).model_formulaAccepted for interface/future-extension purposes but currently not used anywhere in this method's body – supplying a non-
NULLvalue has no effect on discovery, validation, or any stored state. Do not rely on this parameter to affect covariate adjustment; a design's own model formula and design matrix are what individualInferenceclasses actually consult when later constructed from this suite's discovered class list.inference_paramsA named list of lists supplying additional constructor arguments for specific inference classes. Each name must be the name of a concrete
Inferencesubclass that is applicable todes_obj(checked againstapplicable_design_classesonce discovered – an inapplicable class name raises an error); the corresponding list contains keyword arguments (beyonddes_obj) forwarded to that class'sinitialize, and every argument name supplied must match a formal parameter of that class'sinitializemethod (other thandes_objand...) or an error is raised naming the unknown argument(s) and the valid ones. Defaults to an empty list (no extra arguments for any class).
InferenceSuite$run_all_inference()
Construct and fit every class in applicable_design_classes,
and report one uniform comparison row per class – estimate, SE, CI, p-value
(each via the highest-priority available method; see
inference_suite_inspect.md's "Method Selection Policy"), likelihood
tier, estimand (where declared), fit time, captured warnings, and status.
Unlike the constructor, this method fits models and is not free; call it
explicitly when you want the comparison, not automatically.
A single class's construction or fit failure never aborts the report – it
is caught and recorded as that class's status/message (see
"Per-Class Failure Isolation" in the design doc).
Side effects (v1.0.0 slice): screen prints each row as its
class finishes fitting (computation order), not buffered to the end, with a
percent-done/estimated-time-remaining progress bar line underneath each row
(the ETA is the mean per-class elapsed time so far times classes remaining),
followed by a footer listing classes excluded for missing optional packages.
html = TRUE writes a self-contained, timestamped HTML report (the
same table plus the same footer) to output_dir
and opens it via browseURL; it requires the
knitr package. The plots ggplot2 visualizations, and their
embedding into this HTML report, are not yet implemented
(inference_suite_inspect.md TODO-7); pdf output is not yet a
parameter of this method.
Usage
InferenceSuite$run_all_inference(
screen = TRUE,
html = FALSE,
alpha = 0.05,
save_results_as_JSON = FALSE,
plots = screen,
pdf = FALSE,
classes = NULL,
exclude_classes = character(),
max_secs_per_class = NULL,
num_cores = 1L,
formulas = NULL,
methods = NULL,
basic_bootstrap = FALSE,
compute_conf_intervals = FALSE,
output_dir = "~",
combined_evidence_estimands = NULL,
combined_evidence_weighting = c("estimand_grouped", "equal", "custom"),
combined_evidence_weights = NULL
)
Arguments
screenPrint results to the console as each class finishes. At least one of
screen/htmlmust beTRUE.htmlRender, save (
output_dir, timestamped filename), and auto-open a self-contained HTML report of the results.alphaSignificance level: confidence intervals are computed at
1 - alphaandalphais the significance threshold used anywhere the report flags significance. Default0.05.save_results_as_JSONIf
TRUE, serialize the return object (excluding plot objects) to a timestamped JSON file inoutput_dir. Requires the optional jsonlite package; if it is not installed, awarning()is issued and this artifact is skipped rather than erroring. DefaultFALSE.plotsIf
TRUE, build and display (on the current graphics device) one visualization per estimand: an annotated confidence-interval forest plot (p-value left of each interval, interval width right of it, class/method label right-aligned on each row, color-keyed to significance atalpha) stacked over its own “Estimates” box-and-whisker subplot – a free x-axis (same label, same log10/linear choice as the forest, but its own limits) summarizing the point estimates, one point per inference class/formula, collapsed over method/type since those share one estimate. That subplot scales with how many distinct estimates there are: none for a single estimate (redundant with the forest's own dot), dots alone for 2-5, and dots over a box-and-whisker for more than 5. Built with ggplot2 and stacked into a single gtable grob (draw withgrid::grid.draw()); requires the optional ggplot2 package, if it is not installed, awarning()is issued and plotting is skipped rather than erroring. Defaults to the value ofscreen.pdfIf
TRUE, save the visualization to one timestamped multi-page PDF file inoutput_dir(one page per estimand; page height scales with the largest estimand's number of CI rows). Same ggplot2 dependency and missing-package handling asplots. DefaultFALSE.classesOptional character vector restricting which applicable classes to fit – e.g. re-running against only the few classes a user is actually deciding between, without reconstructing the suite. Every name must already be in
applicable_design_classesor this errors, naming the unknown name(s) and the valid ones.NULL(default) fits every applicable class.exclude_classesOptional character vector of applicable classes to skip, applied after
classes. Same validation asclasses. Default none.max_secs_per_classOptional per-class elapsed-time limit in seconds (via
setTimeLimit), after which that class's row getsstatus = "timeout"instead of hanging the whole report. Protects against one pathological class (e.g. a bootstrap/randomization method with many replicates) blocking every other class. Known limitation: R's time limits are checked at R-level interrupt points, so this reliably cuts off slow R-level work but is not guaranteed to interrupt one very slow single native (C/C++/BLAS) call with no intervening R-level check.NULL(default) means no limit.num_coresIf greater than
1, fit classes in parallel across this many forked workers (makeForkCluster) – Unix/Linux only; on other platforms this falls back to sequential (num_cores = 1) with awarning(). Screen output changes under parallel execution: fitting is a single blocking call that only returns once every worker has finished, so there is no meaningful per-class ETA to show while running –screen = TRUEinstead prints a "fitting N classes across K workers" message up front, then every result row together once complete, then a total-elapsed-time summary line (a deliberate design choice, not a degraded default – seeinference_suite_inspect.md's TODO-13). Default1L(sequential, with the normal incremental streaming/progress bar).formulasNULL(default), a single formula (~ .), a single formula string ("~ ."), or a collection of either – includingc(~ 1, ~ .), which base R already returns as a plainlistofformulaobjects (formulas have noc()method of their own), or a character vector (c("~ .", "~ age + sex * smoking")).NULLmeans each class fits once with its own default formula – identical to omitting this argument entirely, sincemodel_formula = NULLat construction already resolves todes_obj$get_design_formula()(default~ .). When non-NULL, only classes whose constructor syntactically accepts amodel_formulaargument are fit once per formula informulas(oneresults_tablerow each, disambiguated inresultsby"<class>[<formula>]"names); classes without amodel_formulaconstructor argument at all still fit exactly once, ignoringformulas. Note this is a syntactic check (does the constructor accept one), not a semantic one (does the fit actually use it) – some classes accept-and-ignoremodel_formula(e.g.InferenceAllSimpleAverageDiff's unadjusted Welch's t-test); seefix_inference_hierarchy.md'sadjusts_for_covariatesregistry-metadata audit, which makes thecov_modelcolumn semantics-aware wherever that audit has landed.methodsNULL(default), a character vector of method sentinel strings, or (TODO-22) a named list,sentinelto a character vector of requestedtypevalues, orNULL, restricting which inference method(s) – and, for the three resampling sentinels marked "typed" below, which resampling/CI- constructiontypeflavor(s) – get fit and reported per class.NULLconsiders every sentinel inEDI_INFERENCE_SUITE_METHOD_SENTINELS, and for each typed sentinel, everytypevalue that class supports (queried at runtime via its ownget_supported_bootstrap_ {pval,ci}_types()/get_supported_bayesian_bootstrap_ {pval,ci}_types()/get_supported_rand_bootstrap_ {pval,ci}_types()accessor – never a hardcoded type table in this package), except class/method/type combinations declared inEDI_COMPREHENSIVE_SLOW_PATHS. Those implemented but prohibitively slow paths are omitted only from this default selection. Supplyingmethodsexplicitly opts into the named sentinel/type combinations even when the registry marks them slow. Thus the default remains broad without allowing known multi-minute paths to dominate a routine report; it is not a single "best available" cascade. List-shaped example:methods = list(bootstrap = c("percentile", "bca"), rand_bootstrap = NULL)fits only"bootstrap"(restricted to the"percentile"/"bca"types that class actually supports) and"rand_bootstrap"(every type that class supports); a sentinel present as a list name with valueNULLstill means "every valid type" for it, exactly like the flat-vector shape – to not fit a sentinel at all, simply don't name it. Requesting atypefor a sentinel with notypeaxis (any sentinel not marked "typed" below) errors. Valid sentinels, each corresponding to one asymptotic/exact/ randomization/resampling inference family a class may (or may not) implement:"wald"Asymptotic Wald inference –
compute_asymp_confidence_interval()/compute_asymp_two_sided_pval()(capability"wald"). The standard closed-form normal-approximation CI/test."exact"Exact inference –
compute_exact_confidence_interval()/compute_exact_two_sided_pval_for_treatment_effect()(capability"exact_test"). Finite-sample-exact methods (e.g. Fisher's exact test, exact binomial)."rand"Randomization inference –
compute_rand_confidence_interval()/compute_rand_two_sided_pval()(capabilities"randomization_ci"/"randomization_test"– distinct capability names for the CI vs. p-value side, since a class can support one without the other). Design-based inference via re-randomizing the observed treatment assignment."rand_bootstrap"(typed)Randomization-bootstrap inference –
compute_rand_bootstrap_confidence_interval()/compute_rand_bootstrap_two_sided_pval()(capabilities"randomization_bootstrap_ci"/"randomization_bootstrap"– distinct capability names for the CI vs. p-value side, since a class can support one without the other). Resamples under the randomization null rather than the usual iid-resampling bootstrap.type(both sides agree on the same four values, unlike"bootstrap"/"bayes_boot"below):"percentile","studentized","symmetric-percentile-t","smoothed"."jackknife"Jackknife-Wald inference –
compute_jackknife_wald_confidence_interval()/compute_jackknife_wald_two_sided_pval()(capability"jackknife"). Leave-one-out variance estimate feeding a Wald-style CI/test."score"Score (Rao) test –
compute_score_confidence_interval()/compute_score_two_sided_pval()(capability"likelihood_tests", one of its three independent sub-procedures)."lik_ratio"Likelihood-ratio test –
compute_lik_ratio_confidence_interval()/compute_lik_ratio_two_sided_pval()(capability"likelihood_tests")."gradient"Gradient test –
compute_gradient_confidence_interval()/compute_gradient_two_sided_pval()(capability"likelihood_tests").
(The plain, "best available" auto-selecting
compute_lik_ratio_bartlett_two_sided_pval()/compute_lik_ratio_bartlett_confidence_interval()dispatcher is deliberately not its own sentinel – it only picks between the two explicit variants below depending on which this class implements, so it is never a distinct inference procedure.)"lik_ratio_bartlett_approx"Bartlett-corrected likelihood-ratio test, Monte-Carlo-approximated correction factor pinned explicitly (for reproducibility) –
compute_lik_ratio_bartlett_approx_confidence_interval()/compute_lik_ratio_bartlett_approx_two_sided_pval()(capability"likelihood_tests"; degrades toNAfor classes without an approximate Bartlett factor)."lik_ratio_bartlett_exact"Bartlett-corrected likelihood-ratio test, closed-form analytic correction factor pinned explicitly –
compute_lik_ratio_bartlett_exact_confidence_interval()/compute_lik_ratio_bartlett_exact_two_sided_pval()(capability"likelihood_tests"; degrades toNAfor classes without an exact Bartlett factor)."param_boot"Bootstrap-calibrated likelihood-ratio test –
compute_lik_ratio_bootstrap_confidence_interval()/compute_lik_ratio_bootstrap_two_sided_pval()(capability"parametric_likelihood_bootstrap")."param_boot_direct"Direct parametric-bootstrap estimate/CI/pval for the treatment coefficient itself –
compute_param_bootstrap_confidence_interval()/compute_param_bootstrap_pval()(capability"parametric_likelihood_bootstrap"; distinct from"param_boot"above, which is a bootstrap-calibrated likelihood-ratio test, not a direct estimate)."bayes_boot"(typed)Bayesian bootstrap inference –
compute_bayesian_bootstrap_confidence_interval()/compute_bayesian_bootstrap_two_sided_pval()(capability"bayesian_bootstrap"). CI-sidetype:"percentile","basic","wald","studentized","bootstrap-t","bca"; pval-sidetypeswaps"basic"for"symmetric"(all others the same)."bootstrap"(typed)Nonparametric bootstrap inference –
compute_bootstrap_confidence_interval()/compute_bootstrap_two_sided_pval()(capability"nonparametric_bootstrap"). CI-sidetype:"percentile","basic","studentized","bootstrap-t","symmetric-percentile-t","bca","prepivoted","double-bootstrap","calibrated","smoothed"; pval-sidetypeis a smaller set –"percentile","symmetric","studentized","bootstrap-t","bca"– neither"basic"nor the other CI-only variants apply on the pval side.
For the three "typed" sentinels above, an exhaustive
typelist is documented here for orientation only – the actual set consulted at runtime always comes from that class's own accessor (see the top of this section), so a class need not support every value listed. ("likelihood_ratio"/"estimating_equation_likelihood_ratio"are deliberately not separate sentinels – both capabilities gate the exact same method pair"lik_ratio"above already covers.) For each class, only sentinels the class has any CI or p-value capability for (among the requestedmethods) get a row; for a typed sentinel, one row pertypethat class actually supports (intersected with any requested type subset) – a class with zero applicable sentinels, or a typed sentinel with zero resulting types, still gets exactly one row withmethod/type = NA_character_(mirrors the pre-methods"no capability" row) rather than being silently dropped. A class contributing more than one applicable-sentinel row is disambiguated inresults/results_tableby"<class>{<method>}"or"<class>{<method>:<type>}"(or with a"[<formula>]"tag too under simultaneousformulasfan-out) names. Unlike the removed cascade,ci_method/pval_methodon a given row now always match that row's ownmethod(or areNAif this class lacks that half of the sentinel's capability, including whentypeis valid on one side but not the other) – there is no fallback to a different sentinel within one row.basic_bootstrapFALSE(default). Convenience flag: whenTRUE, restricts every typed sentinel ("bootstrap"/"bayes_boot"/"rand_bootstrap") to just that class's first (i.e. default)typevalue instead of fitting everytypeit supports – "just run the default bootstrap flavor for nonparametric/Bayesian/randomization resampling" without having to spell outmethods = list(bootstrap = ..., bayes_boot = ..., rand_bootstrap = ...)by hand. Only takes effect for a typed sentinel the caller didn't already restrict via an explicitmethodslist entry – an explicittyperequest there always wins over this flag. No effect on non-typed sentinels ("param_boot"/"param_boot_direct"included – neither has atypeaxis, so they already run their one procedure).compute_conf_intervalsFALSE(default). WhenFALSE, every task's confidence-interval side (ci_a/ci_b/ci_method) is skipped entirely – only the p-value side runs. Several sentinels' CI search (Bartlett-approx,"rand","rand_bootstrap","param_boot") re-invokes the same expensive machinery as its own p-value roughly 15-40 times per bound during root-finding, by far the dominant cost of a fullrun_all_inference()run for those sentinels; skipping it can cut total runtime dramatically.ci_a/ci_b/ci_methodstay present but alwaysNAinresults_table(stable schema either way) and are omitted entirely from the live/print/HTML display tables whenFALSE. SetTRUEto compute CIs as before.output_dirDirectory for the
html/pdf/save_results_as_JSONoutput files. Default"~"(the user's home directory), not the current working directory – these calls are routinely made from inside the package's own source tree (a demo/dev script run fromR/EDI/), and a stray timestamped report left ingetwd()there is exactly the kind of untracked file that can end up bundled into a source tarball (R CMD build) or committed by accident. Pass an explicit path (e.g. the current directory, or a temp directory) to write elsewhere.combined_evidence_estimandsNULL(default: include every declaredestimand), or a character vector ofestimandvalues to restrict the Combined Evidence p-value/weights to. Validated argument-time against theestimandvalues actually declared amongclasses/exclude_classes-filtered candidates.combined_evidence_weightingOne of
"estimand_grouped"(default –w_i = 1 / (G * m_i)),"equal"(flatw_i = 1/k), or"custom"(caller suppliescombined_evidence_weights). Seeinference_suite_inspect.md's TODO-15.combined_evidence_weightsNamed numeric vector (
inference_classname -> weight), required when and only whencombined_evidence_weighting = "custom". Names must be a subset of the classes being fit; an unnamed usable class defaults to weight0(excluded). Need not pre-sum to 1.
Returns
Invisibly, an object of class c("EDIInferenceSuiteResults", "list")
with elements results (one named sub-list per class, in computation
order), results_table (the same rows as a flat data.frame,
sorted/grouped by estimand – NA_character_ last – with a
secondary sort by inference_class; includes the weight
column driven by combined_evidence_weighting/
combined_evidence_estimands), combined_evidence
(list(pval, stat, method = "cauchy_combination", n_classes_used,
n_estimand_groups, estimands_used, weighting, weights_used,
classes_used) – the Cauchy-combination-test p-value/statistic
across all usable rows under the resolved weighting policy;
weights_used/classes_used are keyed/valued by each
row's results name, not results_table$inference_class
directly, since that column can repeat under formulas;
pval = stat = NA_real_ if fewer than 2 rows are usable),
design, alpha, unavailable_due_to_missing_packages,
plots (list(ci_forest); ci_forest is a named list
of one gtable grob per estimand – the CI forest
stacked over its “Estimates” box-and-whisker subplot; draw
with grid::grid.draw() or pass to ggplot2::ggsave() –
possibly empty – rather than a single plot, since the visualization
is split one-per-estimand, per user request, 2026-08-19; the former
separate estimates plot became that subplot, 2026-08-21),
files (list(html, pdf, json), each a
path or NULL; pdf is one multi-page PDF with one page
per estimand), timestamp, total_secs, and
edi_version.
InferenceSuite$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSuite$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Madigan, D., Ryan, P. B., and Schuemie, M. (2013), "Does design matter? Systematic evaluation of the impact of analytical choices on effect estimates in observational studies," Therapeutic Advances in Drug Safety, 4(2), 53-62, PMID 25083251 – the motivating finding behind this class's "every legitimate way to look for an effect, compared honestly" default output.
Liu, Y. and Xie, J. (2020), "Cauchy combination test: a powerful test
with analytic p-value calculation under arbitrary dependency
structures," Journal of the American Statistical Association,
115(529), 393-402 – run_all_inference()'s Combined Evidence
Metric.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 20, response_type = "continuous")
for (i in 1:20) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x = rnorm(1)))
}
seq_des$add_all_subject_responses(rnorm(20))
suite = InferenceSuite$new(seq_des)
suite$applicable_design_classes
# Fit and compare every applicable class:
results = suite$run_all_inference(screen = TRUE)
results$results_table
Cox Proportional Hazards Regression Inference for Survival Responses
Description
Fits a Cox proportional hazards model, \lambda(t \mid x_i) = \lambda_0(t)
\exp(x_i^\top\beta), for survival responses using the treatment indicator
and, optionally, all recorded covariates as predictors, by maximizing the
Breslow-tie-corrected partial likelihood. For exact/right-censored
data, fitting uses this package's internal Newton-Raphson C++ solver
(fast_coxph_regression_prebuilt_cpp) by default (use_rcpp =
TRUE), falling back to survival::coxph.fit()/survival::coxph()
if that fails to converge or use_rcpp = FALSE. For left- or
interval-censored data (a genuinely different likelihood, with no
closed-form partial-likelihood score/information), fitting instead dispatches
to icenReg::ic_sp(model = "ph"), a semiparametric NPMLE Cox fit whose
standard errors come from icenReg's own internal bootstrap (not a
closed-form covariance) — only Wald inference (testing_type = "wald")
is supported on that path; score/gradient/likelihood-ratio/Bartlett testing
types raise an informative error for such data. This is a partial-likelihood
class (likelihood_tier = "partial") supporting score, gradient, and
likelihood-ratio tests, plus parametric likelihood-ratio bootstrap
calibration (both only for the exact/right-censored path), in addition to
Wald and resampling-based inference. Fitted coefficients exceeding a fixed
magnitude threshold (20, on the log-hazard-ratio scale) are treated as
non-estimable (a numerical-divergence guard) rather than returned.
Super class
Inference -> InferenceSurvivalCoxPHRegr
Methods
Public methods
-
InferenceSurvivalCoxPHRegr$compute_asymp_confidence_interval() -
InferenceSurvivalCoxPHRegr$compute_estimate_with_bootstrap_weights()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalCoxPHRegr$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand. Deliberately pulled from
InferenceRand, not InferenceRandCI – despite
InferenceRandCI's override having a richer signature
(type/args_for_type), its body calls super$...(),
which resolves against this class's *actual* R6 superclass
(Inference, which has no such method) once the method body is
extracted and merged flatly by the component system – not against
InferenceRand the way it would inside InferenceRandCI's
own real inheritance chain. InferenceRandCI's only other content
is an incidence-response special case (Zhang exact test) that never
applies to survival data anyway, so the two are behaviorally identical
for this class – confirmed by tracing the body, not assumed. Matches
the pattern already used by the sibling migrated classes (LogRank/
GehanWilcox/KMDiff/RestrictedMeanDiff).
Initialize a Cox PH inference object for a completed design with a survival response. Unlike most survival inference classes in this package, this one accepts left- and interval-censored data (via an icenReg-backed fallback fit; see class documentation), not only exact/right-censored.
Usage
InferenceSurvivalCoxPHRegr$new( des_obj, model_formula = NULL, use_rcpp = TRUE, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a survival response.model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.use_rcppLogical. If
TRUE(default), enable internal Rcpp score/information helpers for likelihood inference.verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceSurvivalCoxPHRegr$compute_estimate()
Computes the Cox PH treatment coefficient \hat\beta_T
(log hazard ratio) — see class documentation for the fitting backend
used (partial-likelihood C++/survival solver for exact/
right-censored data; icenReg NPMLE for left-/interval-censored
data). NA if the fit fails or the fitted coefficients are
numerically extreme.
Usage
InferenceSurvivalCoxPHRegr$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceSurvivalCoxPHRegr$compute_asymp_confidence_interval()
Computes an asymptotic confidence interval using the configured likelihood-backed test.
Usage
InferenceSurvivalCoxPHRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level 1 -
alpha. Default 0.05.
InferenceSurvivalCoxPHRegr$compute_asymp_two_sided_pval()
Computes an asymptotic two-sided p-value using the configured likelihood-backed test.
Usage
InferenceSurvivalCoxPHRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect to test against. Default 0.
InferenceSurvivalCoxPHRegr$compute_estimate_with_bootstrap_weights()
Recomputes the Cox PH treatment estimate under subject/block
bootstrap weights (via weighted_cox_bootstrap_surrogate_fit(),
which assumes ordinary right-censoring semantics), used by the
Bayesian bootstrap and related weighted-resampling machinery. If the
weights are effectively constant, short-circuits to the unweighted
$compute_estimate(estimate_only = TRUE). Not supported
for left- or interval-censored data — raises an error immediately,
since the surrogate weighted fit has no extension for that likelihood.
Always leaves the standard error unavailable (NA) regardless of
estimate_only — this weighted path never computes a variance.
Usage
InferenceSurvivalCoxPHRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyPresent for interface parity; this method never computes variance components regardless of its value.
InferenceSurvivalCoxPHRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalCoxPHRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220, for the proportional hazards model and partial likelihood; Breslow, N. E. (1974). "Covariance Analysis of Censored Survival Data." Biometrics, 30(1), 89-99, doi:10.2307/2529620, for the tied-event partial-likelihood approximation used for exact/right-censored data.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalCoxPHRegr$new(seq_des)
inf$compute_estimate()
Dependent-Censoring Transformation Inference for Survival Responses
Description
Fits a joint bivariate log-normal transformation model for a latent event
time T^E_i and a latent censoring time T^C_i that are allowed
to be dependent (a violation of the usual independent-censoring
assumption): \log T^E_i = X_i^\top \beta_{\mathrm{event}} +
\sigma_{\mathrm{event}} \epsilon^E_i, \log T^C_i = X_i^\top
\beta_{\mathrm{cens}} + \sigma_{\mathrm{cens}} \epsilon^C_i, with
(\epsilon^E_i, \epsilon^C_i) jointly standard bivariate normal with
correlation \rho (estimated via an atanh-reparameterized,
clamped nuisance parameter). X_i includes the treatment indicator
W_i as its first column, so \hat\beta_T (the first entry of
\hat\beta_{\mathrm{event}}) is a log-time-ratio for the
event submodel, on the same AFT interpretation scale as
InferenceSurvivalWeibullRegr
but log-normal rather than Weibull, and jointly modeling the censoring
mechanism rather than assuming it independent. This is the correct tool
when censoring is suspected to depend on the same latent factors driving
the event time (e.g. sicker patients are both more likely to be censored
— dropout — and more likely to fail early), a scenario under which
ordinary Kaplan-Meier/Cox/AFT methods (which assume independent
censoring) are biased. likelihood_tier = "full": likelihood-ratio,
score, gradient, and Wald tests are available when the model converges,
plus parametric-likelihood-bootstrap calibration of the likelihood-ratio
test. Substantial method-support limitations, all deliberate:
randomization inference and jackknife bias correction/standard errors are
hard-unsupported (each randomization draw would require a full
dependent-censoring likelihood refit, too unstable/slow for the
comprehensive test suite; jackknife bias correction is unstable for this
likelihood on small censored samples) — every jackknife/randomization
method returns NA and marks the result nonestimable rather than
computing a value. Nonparametric-bootstrap confidence intervals are
computed but additionally validated/sanity-checked (excessively wide or
zero-excluding-by-construction intervals are treated as unstable and
replaced with NA), and Bayesian-bootstrap weighted re-estimation
uses a fast Cox-model surrogate fit (weighted_cox_bootstrap_surrogate_fit())
as an approximation rather than a full weighted joint-likelihood refit.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Value
A single logical.
Super class
Inference -> InferenceSurvivalDepCensTransformRegr
Methods
Public methods
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalDepCensTransformRegr$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalDepCensTransformRegr$supports_rand_pval_for_incidence()
InferenceSurvivalDepCensTransformRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalDepCensTransformRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
See Also
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik
for a different (Clayton-copula, KK-design) approach to dependence
between two survival-type quantities.
Survival
analysis (Wikipedia, general orientation; no direct Wikipedia page for
dependent-censoring copula/transformation models specifically).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalDepCensTransformRegr$new(seq_des)
inf$compute_estimate()
Clayton Copula / Standard Weibull Compound Inference for KK Designs
Description
This class implements a compound estimator for KK matching-on-the-fly designs with survival responses using a Clayton copula with Weibull AFT margins for matched pairs and a standard Weibull AFT model for the reservoir. The two treatment-effect estimates (on the log-time ratio scale) are combined by inverse-variance weighting.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Frailty distribution. The Clayton copula for a matched pair,
S(t1,t2) = (S1(t1)^-theta + S2(t2)^-theta - 1)^(-1/theta) with Weibull
margins S_i, is exactly the closed-form bivariate survival function
obtained by multiplying two conditionally-independent Weibull hazards by a
shared gamma frailty term Z ~ Gamma(1/theta, 1/theta) and
integrating Z out analytically (Clayton 1978; Oakes 1989); theta
is the frailty variance / dependence parameter (see ClaytonWeibullLikelihood
in fast_survival_models_optim.cpp, which builds the likelihood from the
per-subject Weibull cumulative hazards H1, H2). This is the classic
textbook Weibull-gamma shared-frailty model, fit here in its closed form (no
numerical integration required) rather than as an AFT Gaussian-random-intercept
model.
This is a different (and equally standard) frailty assumption from the
log-normal-frailty Weibull AFT GLMM implemented by
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC /
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik,
which instead places a Gaussian random intercept on the log-time (AFT) scale and
integrates it out by Gauss-Hermite quadrature. Prefer this Clayton-copula class
for the classic gamma-frailty / proportional-hazards dependence structure;
prefer the Weibull-frailty class for a Gaussian-random-intercept / GLMM-style
dependence structure.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC
Methods
Public methods
-
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$supports_rand_pval_for_incidence() -
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$supports_rand_pval_for_incidence( )
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$compute_rand_two_sided_pval()
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Clayton DG (1978). "A Model for Association in Bivariate Life Tables and Its Application in Epidemiological Studies of Familial Tendency in Chronic Disease Incidence." Biometrika, 65(1), 141-151. doi:10.2307/2335289
Oakes D (1989). "Bivariate Survival Models Induced by Frailties." Journal of the American Statistical Association, 84(406), 487-493. doi:10.2307/2289934
Legacy class. Not fully tested in comprehensive_tests.R.
See Also
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC
for the corresponding log-normal-frailty IVWC estimator.
One-Likelihood Clayton-Copula Weibull AFT Inference for KK Survival Designs
Description
Estimates a treatment log-time-ratio \beta_T for right-censored
survival outcomes collected under a KK matching-on-the-fly design
(DesignSeqOneByOneKK14 or
subclass) by maximizing a single combined likelihood: matched-pair
survival times are modeled with a Weibull accelerated-failure-time (AFT)
margin joined by a Clayton copula (dependence parameter \theta)
to account for within-pair correlation induced by shared matching
covariates, while unmatched reservoir subjects are modeled by the same
Weibull AFT margin marginally (no dependence term). All subjects share
one treatment coefficient, estimated jointly.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Estimand. \beta_T, the treatment coefficient of a Weibull
AFT model \log T = \beta_0 + \beta_T W + X\beta + \sigma\epsilon
with \epsilon extreme-value-distributed; \exp(\hat\beta_T)
is the treatment-vs-control survival-time ratio (an acceleration
factor). This is a distinct scale from the log hazard ratio reported by
Cox-based KK survival classes.
Model. .fit_clayton_weibull_aft() jointly optimizes the
AFT regression coefficients, the Weibull shape (\log\sigma), and
the Clayton copula dependence parameter (\log\theta) by direct
maximum likelihood over the combined matched-pair-copula /
reservoir-marginal log-likelihood; right-censoring enters as the usual
survival contribution (density for observed failures, survival function
for censored times). likelihood_tier = "full", so a parametric
likelihood bootstrap (simulate_under_lik_null, which draws new
pair times from the fitted Clayton copula and new singleton times from
the marginal Weibull) is available alongside Wald inference.
Assumptions. Weibull AFT margin correctly specified; Clayton copula correctly captures within-pair dependence (a positive-dependence, single-parameter Archimedean copula); independent censoring given covariates; a KK matching-on-the-fly design supplying the matched/ reservoir partition.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik
Methods
Public methods
-
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$supports_rand_pval_for_incidence() -
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$supports_rand_pval_for_incidence( )
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$compute_rand_two_sided_pval()
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Clayton, D. G. (1978). "A model for association in bivariate life tables
and its application in epidemiological studies of familial tendency in
chronic disease incidence." Biometrika, 65(1), 141-151.
doi:10.1093/biomet/65.1.141. (Clayton1978 in REFERENCES.md.)
Oakes, D. (1989). "Bivariate survival models induced by frailties."
Journal of the American Statistical Association, 84(406),
487-493. doi:10.1080/01621459.1989.10478795. (Oakes1989 in
REFERENCES.md.)
See Also
Analogous Python API for AFT/copula survival models: lifelines WeibullAFTFitter, copulas. Copula (probability theory) (orientation).
Weibull Frailty IVWC Inference for KK Designs
Description
Log-normal (Gaussian random-intercept) frailty Weibull AFT estimator; see
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC
for the frailty-distribution details and contrast with the gamma-frailty
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC
(Clayton copula) alternative.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalGLMMWeibullFrailtyNormalIVWC
Methods
Public methods
-
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$supports_rand_pval_for_incidence() -
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$supports_rand_pval_for_incidence( )
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$compute_rand_two_sided_pval()
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Weibull Frailty Combined-Likelihood Inference for KK Designs
Description
Log-normal (Gaussian random-intercept) frailty Weibull AFT estimator; see
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik
for the frailty-distribution details and contrast with the gamma-frailty
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik
(Clayton copula) alternative.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalGLMMWeibullFrailtyNormalOneLik
Methods
Public methods
-
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$supports_rand_pval_for_incidence() -
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$supports_rand_pval_for_incidence( )
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$compute_rand_two_sided_pval()
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Gehan-Wilcoxon (Peto-Prentice) Inference for Survival Data with Censoring
Description
Non-parametric inference for survival outcomes supporting censored data, using
the Peto-Prentice modification of the Gehan-Wilcoxon test. The treatment effect
estimate is the mean difference in Peto-Prentice weighted martingale residuals
between the treatment and control groups. Specifically, for each subject the
weighted residual is M_i^w = \hat{S}(t_i^-) \cdot M_i, where
M_i = \delta_i - \hat\Lambda_0(t_i) is the martingale residual and
\hat{S}(t_i^-) is the overall Kaplan-Meier survival estimate just before
time t_i. These weights downweight late events, analogously to the
Wilcoxon rank-sum test for uncensored data (which also weights early observations
more heavily via their larger rank denominator).
The p-value uses survival::survdiff(rho = 1) (Peto-Prentice / Fleming-Harrington
p=1, q=0), which is distinct from the log-rank test (rho = 0) used in
InferenceSurvivalKMDiff.
Super class
Inference -> InferenceSurvivalGehanWilcox
Methods
Public methods
-
InferenceSurvivalGehanWilcox$compute_estimate_with_bootstrap_weights() -
InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval() -
InferenceSurvivalGehanWilcox$compute_rand_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalGehanWilcox$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize Gehan-Wilcoxon survival inference and prepare the
rank-based treatment statistic used by
InferenceSurvivalGehanWilcox.
Usage
InferenceSurvivalGehanWilcox$new( des_obj, model_formula = NULL, verbose = FALSE )
Arguments
des_objThe design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseIf TRUE, print additional information.
InferenceSurvivalGehanWilcox$compute_estimate()
Returns the mean difference in Peto-Prentice weighted martingale residuals
between the treatment and control groups. Positive values indicate that treatment
subjects experienced fewer early events than expected. For left- or
interval-censored data, dispatches instead to interval::ictest(...,
scores = "wmw") (the Wilcoxon-Mann-Whitney interval-censored
generalization of the Peto-Prentice test) and returns its estimate.
Usage
InferenceSurvivalGehanWilcox$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
A numeric scalar (the Peto-Prentice weighted score treatment effect estimate).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival") seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10]) seq_des$add_all_subject_responses( ys = c(4.71, NA, 4.78, 6.11, NA, 8.43), y_Ls = c(NA, 1.23, NA, NA, 5.95, NA), y_Rs = c(NA, Inf, NA, NA, Inf, NA) ) seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des) seq_des_inf$compute_estimate()
InferenceSurvivalGehanWilcox$compute_estimate_with_bootstrap_weights()
Recomputes the class-specific treatment estimate under bootstrap weights; see
InferenceBayesianBootstrap.
Usage
InferenceSurvivalGehanWilcox$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval()
Computes a (1 - alpha)-level confidence interval based on the asymptotic normality of the Peto-Prentice weighted martingale residual mean difference. Falls back to bootstrap if the SE is unavailable.
Usage
InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level. Default is 0.05.
Returns
A numeric vector of length 2: (lower, upper) confidence bounds.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival") seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10]) seq_des$add_all_subject_responses( ys = c(4.71, NA, 4.78, 6.11, NA, 8.43), y_Ls = c(NA, 1.23, NA, NA, 5.95, NA), y_Rs = c(NA, Inf, NA, NA, Inf, NA) ) seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des) seq_des_inf$compute_asymp_confidence_interval()
InferenceSurvivalGehanWilcox$compute_asymp_two_sided_pval()
Computes the Peto-Prentice (Gehan-Wilcoxon) two-sided p-value via
survival::survdiff(rho = 1), which puts greater weight on early events
relative to the standard log-rank test (rho = 0).
For delta != 0, not yet implemented.
Usage
InferenceSurvivalGehanWilcox$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect to test against. Default is 0.
Returns
A p-value in [0, 1].
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival") seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10]) seq_des$add_all_subject_responses( ys = c(4.71, NA, 4.78, 6.11, NA, 8.43), y_Ls = c(NA, 1.23, NA, NA, 5.95, NA), y_Rs = c(NA, Inf, NA, NA, Inf, NA) ) seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des) seq_des_inf$compute_asymp_two_sided_pval()
InferenceSurvivalGehanWilcox$compute_rand_confidence_interval()
Randomization confidence intervals are not supported for this class because the Peto-Prentice weighted score scale is not commensurate with the time-ratio null used by the randomization CI bisection algorithm.
Usage
InferenceSurvivalGehanWilcox$compute_rand_confidence_interval( alpha = 0.05, r = 501, pval_epsilon = 0.005, show_progress = TRUE, ci_search_control = NULL )
Arguments
alphaUnused.
rUnused.
pval_epsilonUnused.
show_progressUnused.
ci_search_controlUnused.
InferenceSurvivalGehanWilcox$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalGehanWilcox$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Gehan, E. A. (1965). "A generalized Wilcoxon test for
comparing arbitrarily singly-censored samples." Biometrika,
52(1-2), 203-223, doi:10.1093/biomet/52.1-2.203, for the original
generalized (Gehan) Wilcoxon test for censored data. Peto, R., and Peto,
J. (1972). "Asymptotically Efficient Rank Invariant Test Procedures."
Journal of the Royal Statistical Society, Series A, 135(2),
185-207, doi:10.2307/2344317, for the survival-weighted (Peto-Prentice)
modification this class implements via \rho=1.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalGehanWilcox$new(seq_des)
inf$compute_estimate()
## ------------------------------------------------
## Method `InferenceSurvivalGehanWilcox$compute_estimate()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)
seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_estimate()
## ------------------------------------------------
## Method `InferenceSurvivalGehanWilcox$compute_asymp_confidence_interval()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)
seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_asymp_confidence_interval()
## ------------------------------------------------
## Method `InferenceSurvivalGehanWilcox$compute_asymp_two_sided_pval()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2:10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2:10])
seq_des$add_all_subject_responses(
ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)
seq_des_inf = InferenceSurvivalGehanWilcox$new(seq_des)
seq_des_inf$compute_asymp_two_sided_pval()
LWA-style Marginal Cox IVWC Compound Inference for KK Designs
Description
Fits a compound (IVWC) estimator for KK matching-on-the-fly designs with
survival responses: matched pairs are analyzed with a marginal Cox model
\lambda(t \mid w) = \lambda_0(t)\exp(\beta_T w) whose robust variance
uses the Lee-Wei-Amato (1992) cluster-robust sandwich (treating each matched
pair as an independent cluster of correlated failure times), while reservoir
subjects are analyzed with a standard (independent-subjects) Cox partial
likelihood; the two log-hazard-ratio estimates are then combined by
inverse-variance weighting. likelihood_tier = "partial" (Cox partial
likelihood), but likelihood-ratio/score/gradient tests are not exposed on
this IVWC compound (only on the
OneLik sibling, which
fits one combined partial likelihood across both sources instead of pooling
two separate fits).
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalKKLWACoxPHIVWC
Methods
Public methods
-
InferenceSurvivalKKLWACoxPHIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalKKLWACoxPHIVWC$supports_rand_pval_for_incidence() -
InferenceSurvivalKKLWACoxPHIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalKKLWACoxPHIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalKKLWACoxPHIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKLWACoxPHIVWC$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalKKLWACoxPHIVWC$supports_rand_pval_for_incidence()
InferenceSurvivalKKLWACoxPHIVWC$compute_rand_two_sided_pval()
Usage
InferenceSurvivalKKLWACoxPHIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKLWACoxPHIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalKKLWACoxPHIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Lee, E. W., Wei, L. J., and Amato, D. A. (1992). "Cox-Type Regression Analysis for Large Numbers of Small Groups of Correlated Failure Time Observations." In Survival Analysis: State of the Art, 237-247. Springer. doi:10.1007/978-94-015-7983-4_14
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKKLWACoxPHIVWC$new(seq_des)
inf$compute_estimate()
LWA-style Marginal Cox Combined-Likelihood Inference for KK Designs
Description
Fits a single combined Cox partial likelihood
\lambda(t \mid w, x) = \lambda_0(t)\exp(\beta_T w + \beta_X^\top x)
jointly over matched-pair and reservoir subjects for KK matching-on-the-fly
designs with survival responses (a marginal, not stratified, Cox model:
matched pairs do not get pair-specific baseline hazards). This is the
one-likelihood combined-fit analog of
InferenceSurvivalKKLWACoxPHIVWC,
which instead fits and pools two separate estimators. likelihood_tier
= "partial": exposes likelihood-ratio and parametric-likelihood-bootstrap
inference in addition to Wald/asymptotic and Bayesian-bootstrap paths.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalKKLWACoxPHOneLik
Methods
Public methods
-
InferenceSurvivalKKLWACoxPHOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalKKLWACoxPHOneLik$supports_rand_pval_for_incidence() -
InferenceSurvivalKKLWACoxPHOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalKKLWACoxPHOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalKKLWACoxPHOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKLWACoxPHOneLik$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalKKLWACoxPHOneLik$supports_rand_pval_for_incidence()
InferenceSurvivalKKLWACoxPHOneLik$compute_rand_two_sided_pval()
Usage
InferenceSurvivalKKLWACoxPHOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKLWACoxPHOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalKKLWACoxPHOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKKLWACoxPHOneLik$new(seq_des)
inf$compute_estimate()
Rank Regression Inference for Survival Responses under KK Designs
Description
Fits a multivariate Gehan-Wilcoxon rank regression for survival outcomes under a KK matching-on-the-fly design. The model adjusts for the treatment indicator and, optionally, all recorded covariates.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalKKRankRegrIVWC
Methods
Public methods
-
InferenceSurvivalKKRankRegrIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalKKRankRegrIVWC$supports_rand_pval_for_incidence() -
InferenceSurvivalKKRankRegrIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalKKRankRegrIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalKKRankRegrIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKRankRegrIVWC$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalKKRankRegrIVWC$supports_rand_pval_for_incidence()
InferenceSurvivalKKRankRegrIVWC$compute_rand_two_sided_pval()
Usage
InferenceSurvivalKKRankRegrIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKRankRegrIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalKKRankRegrIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneKK14$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1), x2 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKKRankRegrIVWC$new(seq_des)
inf$compute_estimate()
Stratified Cox / Standard Cox Compound Inference for KK Designs
Description
This class implements a compound estimator for KK matching-on-the-fly designs with survival responses. For matched pairs, it uses stratified Cox proportional hazards regression (each pair is a stratum). For reservoir subjects, it uses standard Cox regression. The two estimates (both log-hazard ratios) are combined via a variance-weighted linear combination.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Details
Under harden = TRUE, multivariate fits preserve the treatment column and
progressively retry reduced covariate sets after QR-based rank reduction and
correlation-based pruning. Extreme finite coefficients / standard errors are
rejected and treated as non-estimable.
The matched-pair sub-estimate treats each pair as its own stratum (a pair-specific baseline hazard, exactly canceling shared-frailty effects within the pair via Cox's partial likelihood) and is a special case of the Lee-Wei-Amato (1992) large-numbers-of-small-groups stratified Cox approach; the reservoir sub-estimate is a standard unstratified Cox partial-likelihood fit (Cox 1972). The two log-hazard-ratio estimates are combined by inverse-variance weighting, the same rule used throughout the KK IVWC family.
Legacy class. Not fully tested in comprehensive_tests.R.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalKKStratCoxPHIVWC
Methods
Public methods
-
InferenceSurvivalKKStratCoxPHIVWC$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalKKStratCoxPHIVWC$supports_rand_pval_for_incidence() -
InferenceSurvivalKKStratCoxPHIVWC$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalKKStratCoxPHIVWC$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalKKStratCoxPHIVWC$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKStratCoxPHIVWC$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalKKStratCoxPHIVWC$supports_rand_pval_for_incidence()
InferenceSurvivalKKStratCoxPHIVWC$compute_rand_two_sided_pval()
Usage
InferenceSurvivalKKStratCoxPHIVWC$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKStratCoxPHIVWC$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalKKStratCoxPHIVWC$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220.
Lee, E. W., Wei, L. J., and Amato, D. A. (1992). "Cox-Type Regression Analysis for Large Numbers of Small Groups of Correlated Failure Time Observations." In Survival Analysis: State of the Art, 237-247. Springer. doi:10.1007/978-94-015-7983-4_14
Stratified Cox Combined-Likelihood Compound Inference for KK Designs
Description
Fits a single combined partial likelihood for KK matching-on-the-fly designs
with survival responses: matched pairs contribute a stratified Cox term (one
stratum per pair, canceling shared-frailty effects within the pair) and
reservoir subjects contribute a standard unstratified Cox term, summed into
one joint partial log-likelihood and optimized jointly for a single shared
treatment coefficient. This differs from the two-stage IVWC sibling
InferenceSurvivalKKStratCoxPHIVWC,
which fits the matched and reservoir sub-models separately and combines the
two log-hazard-ratio estimates by inverse-variance weighting; this class
instead estimates one coefficient from the combined likelihood directly,
which additionally supports likelihood-ratio tests and parametric
likelihood bootstrap (likelihood_tier = "partial").
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Computes a randomization-based p-value.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Randomization p-value.
Super class
Inference -> InferenceSurvivalKKStratCoxPHOneLik
Methods
Public methods
-
InferenceSurvivalKKStratCoxPHOneLik$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalKKStratCoxPHOneLik$supports_rand_pval_for_incidence() -
InferenceSurvivalKKStratCoxPHOneLik$compute_rand_two_sided_pval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalKKStratCoxPHOneLik$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalKKStratCoxPHOneLik$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKStratCoxPHOneLik$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalKKStratCoxPHOneLik$supports_rand_pval_for_incidence()
InferenceSurvivalKKStratCoxPHOneLik$compute_rand_two_sided_pval()
Usage
InferenceSurvivalKKStratCoxPHOneLik$compute_rand_two_sided_pval( r = 501, delta = 0, transform_responses = "none", na.rm = TRUE, show_progress = TRUE, permutations = NULL, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
rNumber of randomization vectors.
deltaThe null difference. Default 0.
deltaNull difference.
transform_responsesType of transformation. Default "none".
transform_responsesTransformation.
na.rmRemove NAs.
show_progressShow progress bar. Default TRUE.
show_progressShow progress.
permutationsPre-computed permutations. Default NULL.
permutationsPre-computed permutations.
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalKKStratCoxPHOneLik$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalKKStratCoxPHOneLik$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Cox, D. R. (1972). "Regression Models and Life-Tables." Journal of the Royal Statistical Society, Series B, 34(2), 187-220.
Lee, E. W., Wei, L. J., and Amato, D. A. (1992). "Cox-Type Regression Analysis for Large Numbers of Small Groups of Correlated Failure Time Observations." In Survival Analysis: State of the Art, 237-247. Springer. doi:10.1007/978-94-015-7983-4_14
Marginal (Cluster-Robust) Weibull Inference for KK Matched-Pair Survival Designs
Description
Initialize the marginal (cluster-robust) Weibull inference object.
Returns the pooled treatment effect estimate (log-time-ratio scale).
Recomputes the treatment estimate under Bayesian-bootstrap weights.
Computes the asymptotic (cluster-robust) confidence interval.
Computes the asymptotic (cluster-robust) two-sided p-value.
Duplicates this subclass while preserving fit caches; see
Inference.
Usage
SurvivalKKWeibullMarginalSource
Details
Fits a single pooled Weibull Accelerated Failure Time (AFT) model across all subjects (treatment plus, optionally, all recorded covariates), ignoring the matched-pair structure in the mean model. Standard errors are computed via a cluster-robust (sandwich) covariance estimator: matched pairs from a KK matching-on-the-fly or binary-match design form size-2 clusters, and unmatched reservoir subjects each form their own singleton cluster.
This is the "marginal" competitor to InferenceSurvivalGLMMWeibullFrailtyNormalOneLik:
rather than modeling the within-pair correlation explicitly via a frailty
term, it fits an ordinary (working-independence) Weibull AFT model and
corrects the treatment-effect standard error post hoc for the within-pair
dependence. The model is fit via the package's fast C++ Weibull AFT backend
(fast_weibull_regression_general_cpp) and the cluster-robust sandwich is
assembled from per-subject dfbeta contributions collapsed within clusters,
which is numerically equivalent to
survival::survreg(..., cluster = ..., robust = TRUE) (retained as a
fallback if the C++ fit fails to converge).
Examples
des = DesignSeqOneByOneKK14$new(n = 20, response_type = 'survival')
for (i in 1:20) {
des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
des$add_all_subject_responses(rexp(20))
inf = InferenceSurvivalKKWeibullMarginal$new(des)
inf$compute_estimate()
Kaplan-Meier Median-Difference Inference for Survival Responses
Description
Fits a non-parametric treatment-effect estimator for censored survival
responses: the difference in Kaplan-Meier median survival times
between the treated and control arms, \hat m_T - \hat m_C. Standard
errors are obtained by back-calculating from each arm's separate
Brookmeyer-Crowley confidence interval for its own median (via
survival::survfit's default log-log-transformed interval):
\hat\sigma_i = (\mathrm{upper}_i - \mathrm{lower}_i) / (2 z_{\alpha/2}),
combined (the two arms are independent by design) as \sqrt{\hat\sigma_T^2 +
\hat\sigma_C^2}. When either arm's median is inestimable (its Kaplan-Meier
curve never reaches 0.5) or the back-calculated bounds are non-finite, the
Wald-style confidence interval and p-value methods
($compute_asymp_confidence_interval(), $compute_asymp_two_sided_pval())
silently fall back to a nonparametric bootstrap instead of returning
NA. A convenience method,
$compute_asymp_log_rank_two_sided_pval_for_treatment_effect(), is also
provided for the log-rank p-value on the same fitted survival curves. For
left- or interval-censored data, the point estimate instead comes
from a Turnbull NPMLE median contrast (interval::icfit(), via
turnbull_npmle_stat_diff()), which has no closed-form standard error —
inference on that path relies entirely on the bootstrap fallback described
above. Randomization confidence intervals are not supported (the median
difference's units are not commensurate with the randomization CI bisection
algorithm's transformed-scale null search).
Super class
Inference -> InferenceSurvivalKMDiff
Methods
Public methods
-
InferenceSurvivalKMDiff$compute_estimate_with_bootstrap_weights() -
InferenceSurvivalKMDiff$compute_asymp_log_rank_two_sided_pval_for_treatment_effect()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalKMDiff$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize Kaplan-Meier median-difference survival inference
and prepare treatment-group survival curves used by
InferenceSurvivalKMDiff.
Usage
InferenceSurvivalKMDiff$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objThe design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseIf TRUE, print additional information.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceSurvivalKMDiff$compute_estimate()
Computes the class-specific mean or survival contrast; see
InferenceMLEorKMSummaryTable.
Usage
InferenceSurvivalKMDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
The setting-appropriate (see description) numeric estimate of the treatment effect
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival") seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10]) seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10]) seq_des$add_all_subject_responses( ys = c(4.71, NA, 4.78, 6.11, NA, 8.43), y_Ls = c(NA, 1.23, NA, NA, 5.95, NA), y_Rs = c(NA, Inf, NA, NA, Inf, NA) ) seq_des_inf = InferenceSurvivalKMDiff$new(seq_des) seq_des_inf$compute_estimate()
InferenceSurvivalKMDiff$compute_estimate_with_bootstrap_weights()
Recomputes the class-specific treatment estimate for a bootstrap sample; see
InferenceNonParamBootstrap.
Usage
InferenceSurvivalKMDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsRow weights for the bootstrap sample.
estimate_onlyIf TRUE, skip variance calculations.
InferenceSurvivalKMDiff$compute_asymp_confidence_interval()
Computes a (1 - alpha)-level confidence interval for the difference in Kaplan-Meier median survival times (treatment minus control).
The Brookmeyer-Crowley confidence interval is obtained for each group's median
separately via survival::survfit (using a log-log transformation of the
survival function by default). The per-group SE is back-calculated from the CI
half-width as \hat\sigma_i = (\text{upper}_i - \text{lower}_i) / (2 z_{\alpha/2}).
The two groups are independent by design, so the SE of the difference is
\sqrt{\hat\sigma_T^2 + \hat\sigma_C^2}, and the CI is
(\hat{m}_T - \hat{m}_C) \pm z_{\alpha/2} \cdot \sqrt{\hat\sigma_T^2 +
\hat\sigma_C^2}.
Falls back to compute_bootstrap_confidence_interval when either group's
median is not estimable (i.e., the Kaplan-Meier curve does not reach 0.5) or
when the Brookmeyer-Crowley CI bounds are NA.
Usage
InferenceSurvivalKMDiff$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaThe significance level; the confidence level is 1 -
alpha. Default is 0.05.
Returns
A numeric vector of length 2 giving the (lower, upper) confidence bounds for the difference in median survival times, on the original time scale.
InferenceSurvivalKMDiff$compute_asymp_two_sided_pval()
Computes a Wald-style 2-sided p-value based on the median difference and its back-calculated standard error.
Usage
InferenceSurvivalKMDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. Default is 0.
Returns
The approximate frequentist p-value
InferenceSurvivalKMDiff$compute_asymp_log_rank_two_sided_pval_for_treatment_effect()
Computes a 2-sided p-value via the log rank test
Usage
InferenceSurvivalKMDiff$compute_asymp_log_rank_two_sided_pval_for_treatment_effect( delta = 0 )
Arguments
deltaThe null difference to test against. For any treatment effect at all this is set to zero (the default).
Returns
The approximate frequentist p-value
InferenceSurvivalKMDiff$compute_rand_confidence_interval()
Uses the shared randomization confidence-interval contract; see
InferenceRandCI.
Usage
InferenceSurvivalKMDiff$compute_rand_confidence_interval( alpha = 0.05, r = 501, pval_epsilon = 0.005, show_progress = TRUE, ci_search_control = NULL )
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.rThe number of randomization vectors. The default is 501.
pval_epsilonThe bisection algorithm tolerance. The default is 0.005.
show_progressShow a text progress indicator.
ci_search_controlUnused.
Returns
A 1 - alpha sized frequentist confidence interval
InferenceSurvivalKMDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalKMDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kaplan, E. L., and Meier, P. (1958). "Nonparametric Estimation from Incomplete Observations." Journal of the American Statistical Association, 53(282), 457-481, doi:10.2307/2281868, for the Kaplan-Meier survival curve estimator each arm's median is read from. Brookmeyer, R., and Crowley, J. (1982). "A Confidence Interval for the Median Survival Time." Biometrics, 38(1), 29-41, doi:10.2307/2530286, for the per-arm median confidence interval this class's standard error is back-calculated from. Turnbull, B. W. (1976). "The Empirical Distribution Function with Arbitrarily Grouped, Censored and Truncated Data." Journal of the Royal Statistical Society, Series B, 38(3), 290-295, doi:10.1111/j.2517-6161.1976.tb01597.x, for the NPMLE used on the left-/interval-censored path.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalKMDiff$new(seq_des)
inf$compute_estimate()
## ------------------------------------------------
## Method `InferenceSurvivalKMDiff$compute_estimate()`
## ------------------------------------------------
seq_des = DesignSeqOneByOneBernoulli$new(n = 6, response_type = "survival")
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[1, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[2, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[3, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[4, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[5, 2 : 10])
seq_des$add_one_subject_to_experiment_and_assign(MASS::biopsy[6, 2 : 10])
seq_des$add_all_subject_responses(
ys = c(4.71, NA, 4.78, 6.11, NA, 8.43),
y_Ls = c(NA, 1.23, NA, NA, 5.95, NA),
y_Rs = c(NA, Inf, NA, NA, Inf, NA)
)
seq_des_inf = InferenceSurvivalKMDiff$new(seq_des)
seq_des_inf$compute_estimate()
Log-Rank Inference for Survival Data with Censoring
Description
Non-parametric all-subject inference for survival outcomes supporting right
censoring, based on the standard two-sample log-rank test. The treatment effect
estimate is the difference in mean martingale residuals between the treatment and
control groups under the pooled null hazard. The p-value uses the classic
log-rank score statistic with its hypergeometric tie-adjusted variance. For
left- or interval-censored data, this class dispatches instead to
interval::ictest(..., scores = "logrank1") (Sun's-scores interval-censored
generalization of the log-rank test) for both the point estimate and testing.
Super class
Inference -> InferenceSurvivalLogRank
Methods
Public methods
-
InferenceSurvivalLogRank$compute_estimate_with_bootstrap_weights() -
InferenceSurvivalLogRank$compute_asymp_confidence_interval() -
InferenceSurvivalLogRank$compute_asymp_log_rank_two_sided_pval_for_treatment_effect()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalLogRank$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize log-rank survival inference and prepare the
treatment-group survival data used by
InferenceSurvivalLogRank.
Usage
InferenceSurvivalLogRank$new(des_obj, model_formula = NULL, verbose = FALSE)
Arguments
des_objThe design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseIf TRUE, print additional information.
InferenceSurvivalLogRank$compute_estimate()
Computes the treatment-effect estimate on the martingale-residual mean-difference scale.
Under left-/interval-censored data, dispatched instead through
interval::ictest()'s Sun's-scores log-rank test (TODO-7,
interval_censored_survival_response.md); its estimate field
(mean score difference between groups) is on the same "difference of
group-mean scores" scale as the right-censored martingale-residual
difference this method otherwise returns.
Usage
InferenceSurvivalLogRank$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
InferenceSurvivalLogRank$compute_estimate_with_bootstrap_weights()
Recomputes the class-specific treatment estimate under bootstrap weights; see
InferenceBayesianBootstrap.
Usage
InferenceSurvivalLogRank$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsBootstrap weights at the subject or block level.
estimate_onlyIf TRUE, skip variance calculations.
InferenceSurvivalLogRank$compute_asymp_confidence_interval()
Computes a (1 - alpha)-level confidence interval based on the asymptotic normality of the martingale-residual mean-difference estimate. Falls back to bootstrap if the estimated standard error is unavailable.
Usage
InferenceSurvivalLogRank$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level.
InferenceSurvivalLogRank$compute_asymp_two_sided_pval()
Computes a Wald-style 2-sided p-value by inverting the confidence interval.
Usage
InferenceSurvivalLogRank$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. Default is 0.
Returns
The approximate frequentist p-value
InferenceSurvivalLogRank$compute_asymp_log_rank_two_sided_pval_for_treatment_effect()
Computes the standard two-sided log-rank p-value for a zero treatment effect.
Under left-/interval-censored data, this is interval::ictest()'s
own p-value (TODO-7) rather than a re-derived chi-squared statistic.
Usage
InferenceSurvivalLogRank$compute_asymp_log_rank_two_sided_pval_for_treatment_effect( delta = 0 )
Arguments
deltaNull treatment effect to test against. Only
0is supported.
InferenceSurvivalLogRank$compute_rand_confidence_interval()
Randomization confidence intervals are not supported for this class because the martingale-residual score scale is not commensurate with the transformed time-ratio null used by the randomization CI algorithm.
Usage
InferenceSurvivalLogRank$compute_rand_confidence_interval( alpha = 0.05, r = 501, pval_epsilon = 0.005, show_progress = TRUE, ci_search_control = NULL )
Arguments
alphaUnused.
rUnused.
pval_epsilonUnused.
show_progressUnused.
ci_search_controlUnused.
InferenceSurvivalLogRank$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalLogRank$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Mantel, N. (1966). "Evaluation of survival data and two new
rank order statistics arising in its consideration." Cancer
Chemotherapy Reports, 50(3), 163-170, for the log-rank test. Peto, R.,
and Peto, J. (1972). "Asymptotically Efficient Rank Invariant Test
Procedures." Journal of the Royal Statistical Society, Series A,
135(2), 185-207, doi:10.2307/2344317, for its asymptotic-efficiency
properties and the \rho=0 case of the Fleming-Harrington family
this class corresponds to (see also
InferenceSurvivalGehanWilcox
for the \rho=1 member of the same family). Sun, J. (1996). "A
non-parametric test for interval-censored failure time data with
application to AIDS studies." Statistics in Medicine, 15(13),
1387-1395, for the interval-censored generalization used on the
left-/interval-censored path.
Examples
set.seed(1)
x_dat <- data.frame(
x1 = c(-1.2, -0.7, -0.2, 0.3, 0.8, 1.3, 1.8, 2.3),
x2 = c(0, 1, 0, 1, 0, 1, 0, 1)
)
seq_des <- DesignSeqOneByOneBernoulli$
new(
n = nrow(x_dat),
response_type = "survival",
verbose = FALSE
)
for (i in seq_len(nrow(x_dat))) {
seq_des$
add_one_subject_to_experiment_and_assign(x_dat[i, , drop = FALSE])
}
seq_des$
add_all_subject_responses(
ys = c(1.2, 2.4, NA, 3.1, NA, 4.0, 3.3, NA),
y_Ls = c(NA, NA, 1.8, NA, 2.7, NA, NA, 4.5),
y_Rs = c(NA, NA, Inf, NA, Inf, NA, NA, Inf)
)
infer <- InferenceSurvivalLogRank$
new(
seq_des,
verbose = FALSE
)
infer
Restricted Mean Survival Time (RMST) Difference Inference for Survival Responses
Description
Fits a non-parametric treatment-effect estimator for censored survival
responses: the difference in restricted mean survival time (RMST)
between the treated and control arms, \hat\mu_T(\tau) - \hat\mu_C(\tau),
where each arm's RMST is the area under its Kaplan-Meier survival curve up to
a truncation horizon \tau (\hat\mu(\tau) = \int_0^\tau \hat S(t)\,dt),
computed by trapezoidal integration of the step-function KM curve. The
standard error of the difference comes from the Greenwood-type variance of
each arm's RMST, combined across the two (independent) arms via
get_restricted_mean_se_diff(). When that standard error is
unavailable or non-finite, $compute_asymp_confidence_interval() falls
back to a nonparametric bootstrap interval rather than returning NA.
Randomization confidence intervals are not supported (the RMST-difference
units are not commensurate with the randomization CI bisection algorithm's
transformed-scale null search).
Super class
Inference -> InferenceSurvivalRestrictedMeanDiff
Methods
Public methods
-
InferenceSurvivalRestrictedMeanDiff$compute_estimate_with_bootstrap_weights() -
InferenceSurvivalRestrictedMeanDiff$compute_asymp_confidence_interval() -
InferenceSurvivalRestrictedMeanDiff$compute_asymp_two_sided_pval() -
InferenceSurvivalRestrictedMeanDiff$compute_rand_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalRestrictedMeanDiff$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Initialize restricted-mean-survival-time difference
inference and prepare treatment-group survival summaries used by
InferenceSurvivalRestrictedMeanDiff.
Usage
InferenceSurvivalRestrictedMeanDiff$new( des_obj, model_formula = NULL, verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objThe design object.
model_formulaOptional formula for covariate adjustment. If
NULL(default), the formula from the design object is used and its pre-computed design matrix is reused. If a formula is provided, a new design matrix is constructed from the design's imputed covariates.verboseIf TRUE, print additional information.
smart_cold_start_defaultWhether to use smart cold start values by default.
InferenceSurvivalRestrictedMeanDiff$compute_estimate()
Computes the class-specific mean or survival contrast; see
InferenceMLEorKMSummaryTable.
Usage
InferenceSurvivalRestrictedMeanDiff$compute_estimate(estimate_only = FALSE)
Arguments
estimate_onlyIf TRUE, skip variance component calculations.
Returns
The setting-appropriate (see description) numeric estimate of the treatment effect
InferenceSurvivalRestrictedMeanDiff$compute_estimate_with_bootstrap_weights()
Recomputes the class-specific treatment estimate for a bootstrap sample; see
InferenceNonParamBootstrap.
Usage
InferenceSurvivalRestrictedMeanDiff$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsRow weights for the bootstrap sample.
estimate_onlyIf TRUE, skip variance calculations.
InferenceSurvivalRestrictedMeanDiff$compute_asymp_confidence_interval()
Computes a 1-\alpha level Wald confidence interval for
the RMST-difference treatment effect \hat\mu_T(\tau) -
\hat\mu_C(\tau), using its Greenwood-based standard error (see class
documentation). Falls back to a nonparametric bootstrap interval if
that standard error is unavailable or non-finite.
Usage
InferenceSurvivalRestrictedMeanDiff$compute_asymp_confidence_interval( alpha = 0.05 )
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.
Returns
A (1 - alpha)-sized frequentist confidence interval for the treatment effect
InferenceSurvivalRestrictedMeanDiff$compute_asymp_two_sided_pval()
Computes a two-sided Wald p-value testing H_0:
\mu_T(\tau) - \mu_C(\tau) = 0 (only delta = 0 is currently
supported; a non-zero null raises an error), using the RMST-difference
estimate and its Greenwood-based standard error — see class
documentation. Falls back to a nonparametric bootstrap p-value if that
standard error is unavailable.
Usage
InferenceSurvivalRestrictedMeanDiff$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaThe null difference to test against. For any treatment effect at all this is set to zero (the default).
Returns
The approximate frequentist p-value
InferenceSurvivalRestrictedMeanDiff$compute_rand_confidence_interval()
Uses the shared randomization confidence-interval contract; see
InferenceRandCI.
Usage
InferenceSurvivalRestrictedMeanDiff$compute_rand_confidence_interval( alpha = 0.05, r = 501, pval_epsilon = 0.005, show_progress = TRUE, ci_search_control = NULL )
Arguments
alphaThe confidence level in the computed confidence interval is 1 -
alpha. The default is 0.05.rThe number of randomization vectors. The default is 501.
pval_epsilonThe bisection algorithm tolerance. The default is 0.005.
show_progressShow a text progress indicator.
ci_search_controlUnused.
Returns
A 1 - alpha sized frequentist confidence interval
InferenceSurvivalRestrictedMeanDiff$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalRestrictedMeanDiff$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Royston, P., and Parmar, M. K. B. (2013). "Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome." BMC Medical Research Methodology, 13, 152, doi:10.1186/1471-2288-13-152, for RMST as a treatment-effect summary. Kaplan, E. L., and Meier, P. (1958). "Nonparametric Estimation from Incomplete Observations." Journal of the American Statistical Association, 53(282), 457-481, doi:10.2307/2281868, for the underlying survival curve estimator each arm's RMST is integrated from.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalRestrictedMeanDiff$new(seq_des)
inf$compute_estimate()
Stratified Cox PH Inference for Survival Responses
Description
Fits an auto-stratified Cox proportional hazards regression: rather
than the plain single-baseline-hazard model of
InferenceSurvivalCoxPHRegr,
this class allows a separate baseline hazard per stratum,
\lambda(t \mid x_i, s_i) = \lambda_{0,s_i}(t) \exp(x_i^\top\beta),
relaxing the proportional-hazards assumption across strata while keeping it
within each. Stratification variables are chosen automatically from
the recorded low-cardinality (categorical-like) covariates
(compute_survival_strata_ids_cpp) — no stratification variables are
specified explicitly by the caller. If no suitable stratification covariates
are found, the fit falls back to the corresponding standard (unstratified)
Cox PH model. Fitting uses survival::coxph.fit()/survival::coxph()
with strata passed through when applicable. This is a partial-likelihood
class (likelihood_tier = "partial") supporting Wald, score, gradient,
and likelihood-ratio tests, plus parametric likelihood-ratio bootstrap
calibration. Randomization confidence intervals are not supported (the
log-hazard-ratio estimator units are not commensurate with the randomization
CI bisection algorithm's log-time-ratio/AFT-effect null search).
Super class
Inference -> InferenceSurvivalStratCoxPHRegr
Methods
Public methods
-
InferenceSurvivalStratCoxPHRegr$compute_asymp_confidence_interval() -
InferenceSurvivalStratCoxPHRegr$compute_asymp_two_sided_pval() -
InferenceSurvivalStratCoxPHRegr$compute_estimate_with_bootstrap_weights() -
InferenceSurvivalStratCoxPHRegr$compute_rand_confidence_interval()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalStratCoxPHRegr$new()
Uses the shared randomization two-sided p-value contract; see
InferenceRand. Pinned from
InferenceRand for the same traced reason as
InferenceSurvivalCoxPHRegr (see that factory call's comment):
InferenceRandCI's richer override calls super$...(),
which resolves against Inference once flattened, and its only
other behavior is an incidence-only Zhang special case that never
applies to survival data.
Initialize stratified Cox proportional-hazards inference and
prepare the partial-likelihood fit used by
InferenceSurvivalStratCoxPHRegr.
Usage
InferenceSurvivalStratCoxPHRegr$new( des_obj, model_formula = NULL, use_rcpp = TRUE, optimization_alg = "lbfgs", verbose = FALSE, smart_cold_start_default = NULL )
Arguments
des_objA completed
Designobject with a survival response.model_formulaOptional formula for covariate adjustment. If
NULL(default), covariates from the design object are included. Use~ 1for univariate.use_rcppLogical. If
TRUE(default), enable internal Rcpp score/information helpers for likelihood inference. Cox optimization uses survival::coxph.fit.optimization_algOptimization algorithm:
"newton_raphson"(default) or"lbfgs".verboseWhether to print progress messages.
smart_cold_start_defaultWhether to use smart cold start values.
InferenceSurvivalStratCoxPHRegr$compute_asymp_confidence_interval()
Computes an asymptotic confidence interval using the configured likelihood-backed test.
Usage
InferenceSurvivalStratCoxPHRegr$compute_asymp_confidence_interval(alpha = 0.05)
Arguments
alphaSignificance level 1 -
alpha. Default 0.05.
InferenceSurvivalStratCoxPHRegr$compute_asymp_two_sided_pval()
Computes an asymptotic two-sided p-value using the configured likelihood-backed test.
Usage
InferenceSurvivalStratCoxPHRegr$compute_asymp_two_sided_pval(delta = 0)
Arguments
deltaNull treatment effect to test against. Default 0.
InferenceSurvivalStratCoxPHRegr$compute_estimate_with_bootstrap_weights()
Recomputes the stratified Cox PH treatment estimate under Bayesian-bootstrap weights.
Usage
InferenceSurvivalStratCoxPHRegr$compute_estimate_with_bootstrap_weights( subject_or_block_weights, estimate_only = FALSE )
Arguments
subject_or_block_weightsSubject-, block-, cluster-, or matched-set bootstrap weights.
estimate_onlyIf
TRUE, compute only the weighted point estimate.
InferenceSurvivalStratCoxPHRegr$compute_rand_confidence_interval()
Compute a randomization-based confidence interval for the
stratified Cox treatment effect by inverting the class-specific
randomization p-value. See
InferenceRandCI.
Usage
InferenceSurvivalStratCoxPHRegr$compute_rand_confidence_interval( alpha = 0.05, r = 501, pval_epsilon = 0.005, show_progress = TRUE, ci_search_control = NULL )
Arguments
alphaThe significance level (default 0.05).
rNumber of vectors to draw.
pval_epsilonThe bisection convergence tolerance.
show_progressWhether to show a progress bar.
ci_search_controlUnused.
InferenceSurvivalStratCoxPHRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalStratCoxPHRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalStratCoxPHRegr$new(seq_des)
inf$compute_estimate()
Weibull AFT Inference for Survival Responses
Description
Fits a Weibull Accelerated Failure Time (AFT) model for survival responses:
\log T_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \sigma \epsilon_i,
\epsilon_i \sim standard extreme-value (Gumbel-minimum), so that
T_i is marginally Weibull-distributed with shape 1/\sigma and
treatment-dependent scale, by maximum likelihood
(fast_weibull_regression_cpp). \hat\beta_T is a
log-time-ratio (log acceleration factor): \exp(\hat\beta_T)
is the estimated multiplicative effect of treatment on survival time (an
AFT model, not a proportional-hazards model — the Weibull distribution is
the one location where AFT and proportional-hazards parameterizations
coincide, since \exp(-\beta_T/\sigma) also equals the treatment
hazard ratio). likelihood_tier = "full": likelihood-ratio, score,
gradient, and Wald tests are all available when the model converges, plus
parametric-likelihood-bootstrap calibration of the likelihood-ratio test.
Right-censored and interval-censored observations enter the likelihood via
their appropriate survival/density contributions. Validity requires the
Weibull shape assumption for the (log-)survival-time distribution and,
when interpreted causally, the usual design-based/model-based assumptions.
Computes the randomization distribution of the treatment effect estimate under the sharp null.
Whether compute_rand_two_sided_pval() is actually
usable on this instance right now – FALSE exactly when it
would stop(): an incidence-response instance with no
custom randomization statistic and a design not eligible for
design-randomization-based incidence inference (see
private$should_use_design_randomization_for_incidence()).
TRUE for every other case, including every non-incidence
response type. Public, self-contained (only reads already-set
instance state, no side effects), so InferenceSuite can
check this before attempting the sentinel instead of relying on
the stop() being silently swallowed into a pval = NA
"ok" row – the single source of truth for both this check and
compute_rand_two_sided_pval()'s own guard, so the two can
never drift apart.
Value
When debug = FALSE (default), a numeric vector of length r. When
debug = TRUE, a list with: values, errors (list of character
vectors, one per iteration), warnings (list of character vectors, one per
iteration), num_errors, num_warnings,
prop_iterations_with_errors, prop_iterations_with_warnings, and
prop_illegal_values.
A single logical.
Super class
Inference -> InferenceSurvivalWeibullRegr
Methods
Public methods
-
InferenceSurvivalWeibullRegr$approximate_randomization_distribution_beta_hat_T() -
InferenceSurvivalWeibullRegr$supports_rand_pval_for_incidence()
+ inherited public methods from Inference
Inference$capabilities()Inference$compute_asymp_confidence_interval()Inference$compute_asymp_two_sided_pval()Inference$compute_estimate()Inference$compute_exact_confidence_interval()Inference$compute_exact_two_sided_pval_for_treatment_effect()Inference$duplicate()Inference$get_analysis_data()Inference$get_covariates()Inference$get_design_object()Inference$get_model_formula()Inference$get_nonestimable_reason()Inference$get_nonestimable_stage()Inference$get_optimization_alg()Inference$get_response()Inference$get_response_type()Inference$get_treatment()Inference$initialize()Inference$is_nonestimable()Inference$set_optimization_alg()Inference$set_seed()Inference$supports()
InferenceSurvivalWeibullRegr$approximate_randomization_distribution_beta_hat_T()
Usage
InferenceSurvivalWeibullRegr$approximate_randomization_distribution_beta_hat_T( r = 501, delta = 0, transform_responses = "none", show_progress = TRUE, permutations = NULL, debug = FALSE, zero_one_logit_clamp = .Machine$double.eps )
Arguments
rNumber of randomization vectors. Default 501.
deltaThe null difference. Default 0.
transform_responsesType of transformation. Default "none".
show_progressShow progress bar. Default TRUE.
permutationsPre-computed permutations. Default NULL.
debugIf
TRUE, return a list with the distribution values and per-iteration diagnostics including error messages, warning messages, counts of each, and summary proportions for iterations with errors, warnings, and illegal (non-finite) values. Runs serially. DefaultFALSE.zero_one_logit_clampThe clamping amount for exact 0 and 1 values when logging
InferenceSurvivalWeibullRegr$supports_rand_pval_for_incidence()
Usage
InferenceSurvivalWeibullRegr$supports_rand_pval_for_incidence()
InferenceSurvivalWeibullRegr$clone()
The objects of this class are cloneable with this method.
Usage
InferenceSurvivalWeibullRegr$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
References
Kalbfleisch, J. D., and Prentice, R. L. (2002). The
Statistical Analysis of Failure Time Data (2nd ed.). Wiley, for the
Weibull AFT model and its equivalence to a proportional-hazards model.
For the randomization confidence interval
(compute_rand_confidence_interval()), which inverts an AFT sharp
null by rescaling the recorded times of treated units — event and
censoring times alike, censoring indicators unchanged — the residual
construction and its validity under independent censoring are from
Tsiatis, A. A. (1990). Estimating regression parameters using linear rank
tests for censored data. The Annals of Statistics, 18(1), 354-372,
doi:10.1214/aos/1176347504; Wei, L. J., Ying, Z., and Lin, D. Y. (1990).
Linear regression analysis of censored survival data based on rank tests.
Biometrika, 77(4), 845-851, doi:10.1093/biomet/77.4.845; and Jin,
Z., Lin, D. Y., Wei, L. J., and Ying, Z. (2003). Rank-based inference for
the accelerated failure time model. Biometrika, 90(2), 341-353,
doi:10.1093/biomet/90.2.341.
See Also
Comparable Python API: lifelines WeibullAFTFitter. See also: Proportional hazards model (Wikipedia, for the AFT/PH equivalence note).
Examples
seq_des = DesignSeqOneByOneBernoulli$new(n = 10, response_type = 'survival')
for (i in 1:10) {
seq_des$add_one_subject_to_experiment_and_assign(data.frame(x1 = rnorm(1)))
}
seq_des$add_all_subject_responses(runif(10))
inf = InferenceSurvivalWeibullRegr$new(seq_des)
inf$compute_estimate()
inf$set_seed(1)
inf$compute_lik_ratio_bootstrap_two_sided_pval(delta = 0, B = 9, show_progress = FALSE)
Abstract class for LWA-style Marginal Cox / Standard Cox Compound Inference
Description
Initialize KK LWA Cox IVWC inference and prepare the
matched/reservoir marginal Cox partial-likelihood components used by
InferenceSurvivalKKLWACoxPHIVWC.
Returns the estimated treatment effect (log-hazard ratio).
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Compute the LWA Cox asymptotic p-value for the treatment
log-hazard ratio using the cluster-robust partial-likelihood standard
error. See InferenceAsymp.
Usage
KKLWACoxIVWCPartialLikelihoodSource
Details
This class implements a compound estimator for KK matching-on-the-fly designs with survival responses. For matched pairs, it uses a marginal Cox proportional hazards model with Lee-Wei-Amato style cluster-robust variance, treating each pair as a cluster of size two. For reservoir subjects, it uses standard Cox regression. The two estimates (both log-hazard ratios) are combined via a variance-weighted linear combination.
Under harden = TRUE, multivariate component fits preserve the treatment
column and retry reduced covariate sets after QR-based rank reduction and
correlation-based pruning. Extreme finite coefficients / standard errors are
rejected and treated as non-estimable.
Abstract class for LWA-style Marginal Cox Combined-Likelihood Inference
Description
Initialize KK LWA Cox one-likelihood inference and prepare
the combined marginal Cox partial-likelihood fit used by
InferenceSurvivalKKLWACoxPHOneLik.
Returns the model-specific combined-likelihood treatment estimate; see
InferenceAsympLik.
Recomputes the LWA one-likelihood treatment estimate under Bayesian-bootstrap weights.
Compute an asymptotic confidence interval.
Compute an asymptotic two-sided p-value.
Usage
KKLWACoxOneLikPartialLikelihoodSource
Details
Fits a single joint marginal Cox model over all KK design data for survival responses. Matched subjects share their pair ID as a cluster, and reservoir subjects are treated as independent (unique clusters). Standard errors are obtained via the Huber-White cluster-robust sandwich estimator (LWA style).
Abstract base class for KK Wilcoxon-based compound inference
Description
Shared base for all KK Wilcoxon inference classes. Overrides the per-permutation statistic used in randomization tests with standardized Wilcoxon W statistics (O(n log n), conf.int = FALSE), avoiding the O(n^2) Walsh-average computation required by the full Hodges-Lehmann estimate.
Usage
KKWilcoxIVWCSource
A Fixed Observational (Non-Randomized) Design
Description
A fixed-sample-size DesignFixed whose treatment
assignment vector w is supplied by the user (e.g. via
assign_w_to_all_subjects(w_precomputed = ...) or
overwrite_all_subject_assignments()) rather than drawn from any
randomization mechanism. Unlike every other DesignFixed subclass, there is no
prob_T to specify – treatment was not assigned by the experimenter according
to a known probability law, so no such probability exists to declare.
No draw mechanism. draw_ws_according_to_design() (and, transitively,
the fallback branch of assign_w_to_all_subjects() that would otherwise call
it) always throws. Any inference procedure that must redraw w from the design's
own randomization law – randomization tests, randomization confidence intervals, and
randomization/assignment bootstrap – therefore throws the same clear error rather
than silently fabricating a randomization mechanism that never existed. Procedures
that resample subjects instead of redrawing w (plain nonparametric
bootstrap, Bayesian bootstrap) are unaffected and remain available, since resampling
subjects with their observed, fixed assignment does not require a known randomization
probability.
No balance target. assert_even_allocation() is a no-op here (rather
than the inherited check against prob_T = 0.5): there is no targeted
allocation ratio for an observational design to be out of balance with.
Super classes
Design -> DesignFixed -> ObservationalDesign
Methods
Public methods
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_resampling()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
ObservationalDesign$new()
Initialize a fixed observational (non-randomized) design. No
treatment vector is drawn or requested here; the constructor only records
configuration (covariates, response type, etc.) and internally fixes
prob_T = 0.5 purely so the shared Design/DesignFixed
machinery has a value to store — it is never used to draw or validate an
allocation for this class (see assert_even_allocation() and class
documentation). w itself is supplied afterward via
assign_w_to_all_subjects(w_precomputed = ...) or
overwrite_all_subject_assignments().
Usage
ObservationalDesign$new( response_type, include_is_missing_as_a_new_feature = TRUE, n = NULL, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
include_is_missing_as_a_new_featureFlag for missingness indicators.
nThe sample size.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility. Note this design has no randomization mechanism to seed (see class documentation);
seedonly affects RNG-dependent behavior inherited fromDesignunrelated to treatment assignment (e.g. imputation, bootstrap resampling).
Returns
A new 'ObservationalDesign' object
ObservationalDesign$assert_even_allocation()
Observational designs have no targeted allocation ratio to
check balance against, so this is a no-op rather than an error (contrast
with the inherited Design$assert_even_allocation(), which errors if
the realized allocation deviates from prob_T = 0.5).
Usage
ObservationalDesign$assert_even_allocation()
Returns
invisible(NULL), always; never errors.
ObservationalDesign$supports_randomization_draw()
Characterization: FALSE – this design has no
randomization mechanism to redraw w from (see class documentation).
Metadata-declared replacement for the old draw_ws_raw() throwing
stub (still present below as a fallback for any caller that reaches it
without checking this first – see fix_design_hierarchy.md,
"Observational Design Migration").
Usage
ObservationalDesign$supports_randomization_draw()
Returns
Always FALSE for this class.
ObservationalDesign$supports_resampling_replay()
Characterization: FALSE – this design has no
randomization mechanism to replay against resampled data (bootstrap
randomization test eligibility). Plain nonparametric/Bayesian/m-out-of-n/
PRW-subsampling bootstrap are unaffected (see
Design$supports_resampling()'s documentation) and remain available.
Usage
ObservationalDesign$supports_resampling_replay()
Returns
Always FALSE for this class.
ObservationalDesign$clone()
The objects of this class are cloneable with this method.
Usage
ObservationalDesign$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
des = ObservationalDesign$new(n = 10, response_type = 'continuous')
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(10)))
des$assign_w_to_all_subjects(w_precomputed = rbinom(10, 1, 0.3))
A Fixed Observational (Non-Randomized) Design With Blocks
Description
An ObservationalDesign whose subjects are
additionally partitioned into user-supplied blocks (matched sets / strata), via the
block-membership vector m. As with
ObservationalDesign, there is no randomization
mechanism at all – neither the treatment assignment w nor the block membership
m is drawn by this class, both are supplied by the user – so
draw_ws_according_to_design() still always throws (inherited unchanged from
ObservationalDesign).
Why blocks, if there's no randomization to block on? Blocking here is not a
randomization restriction (there is none); it is a resampling structure. Supplying
m lets bootstrap procedures resample within blocks – exactly as they do
for DesignFixedBlocking and
DesignFixedOptimalBlocks – which is
appropriate when the observational data itself has a matched/stratified/clustered
structure (e.g. matched case-control sets, repeated measurements within site) that the
bootstrap should respect.
No auto-derived n. Unlike plain ObservationalDesign, n is not a
constructor argument here at all – it is always length(m), since a block
membership vector with one entry per subject already fixes the sample size.
Super classes
Design -> DesignFixed -> ObservationalDesign -> ObservationalDesignBlocks
Methods
Public methods
+ inherited public methods from ObservationalDesign
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_resampling()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
ObservationalDesignBlocks$new()
Initialize a fixed observational (non-randomized) design with
user-supplied block membership. Neither treatment w nor block
membership m is drawn here — both must be supplied (m at
construction, w afterward via
assign_w_to_all_subjects(w_precomputed = ...)); see class
documentation for why blocks are still useful without a randomization
mechanism to block on.
Usage
ObservationalDesignBlocks$new( response_type, m, include_is_missing_as_a_new_feature = TRUE, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
mA positive-integer vector of block (matched-set/stratum) identifiers, one entry per subject; a block may contain any number of subjects (unlike
ObservationalDesignMatching, which fixes block size at exactly 2).nis derived aslength(m)and is not a separate argument.include_is_missing_as_a_new_featureFlag for missingness indicators.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'ObservationalDesignBlocks' object
ObservationalDesignBlocks$clone()
The objects of this class are cloneable with this method.
Usage
ObservationalDesignBlocks$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
des = ObservationalDesignBlocks$new(response_type = 'continuous', m = c(1, 1, 2, 2, 3, 3))
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(6)))
des$assign_w_to_all_subjects(w_precomputed = c(1, 0, 1, 0, 1, 0))
A Fixed Observational (Non-Randomized) Matched-Pair Design
Description
An ObservationalDesign whose n subjects
are organized into n/2 matched pairs, each of size exactly 2. This is a
convenience wrapper: it builds the canonical pair-membership vector
m = (1, 1, 2, 2, \dots, n/2, n/2) internally – subject 2k - 1 and subject
2k are pair k – and installs it via set_m(), so the caller need
only supply subjects (and, later, responses) in that paired order; there is no
separate m argument to get wrong.
Why not just ObservationalDesignBlocks with block size 2? It would
look equivalent but silently behave differently: matching is a distinct capability
(private$matching_capable, checked via is_matching_design()) from
generic blocking, and several Inference components (jackknife, nonparametric
bootstrap, Bayesian bootstrap, exchangeable-resampling-unit selection) branch on it to
use pair-preserving resampling (MatchingStructure's
draw_bootstrap_indices(), via draw_matching_bootstrap_sample_cpp())
instead of generic per-stratum resampling. ObservationalDesignBlocks overrides
draw_bootstrap_indices() with the generic stratified version, so subclassing it
here would advertise is_matching_design() == TRUE while still running the
wrong bootstrap underneath. This class instead extends
ObservationalDesign directly – the same
relationship DesignFixedBinaryMatch has to
DesignFixed – so it inherits
MatchingStructure's matched-pair bootstrap machinery unmodified.
Super classes
Design -> DesignFixed -> ObservationalDesign -> ObservationalDesignMatching
Methods
Public methods
+ inherited public methods from ObservationalDesign
+ inherited public methods from DesignFixed
+ inherited public methods from Design
Design$add_one_subject_response()Design$any_censoring()Design$applicable_inference_class_names()Design$assert_all_responses_recorded()Design$assert_all_subjects_arrived()Design$assert_fixed_sample()Design$capabilities()Design$check_experiment_completed()Design$draw_ws_according_to_design()Design$duplicate()Design$get_X()Design$get_X_imp()Design$get_X_raw()Design$get_design_formula()Design$get_edi_version_created()Design$get_effective_dead()Design$get_effective_time()Design$get_missingness_method()Design$get_n()Design$get_ordinal_levels()Design$get_original_ordinal_levels()Design$get_prob_T()Design$get_response_type()Design$get_response_type_original()Design$get_t()Design$get_w()Design$get_y()Design$get_y_L()Design$get_y_R()Design$get_y_original()Design$has_general_censoring()Design$incompatible_inference_classes_due_to_design_structure()Design$is_a_bernoulli_capable()Design$is_a_cluster_capable()Design$is_a_kk_matching_capable()Design$is_blocking_design()Design$is_fixed_sample_size()Design$is_matching_design()Design$prepare_for_resampling_replay()Design$randomization_family()Design$supports()Design$supports_resampling()Design$transform_y()Design$unavailable_inference_classes_due_to_missing_packages()Design$warm_all_subject_data_cache()
ObservationalDesignMatching$new()
Initialize a fixed observational (non-randomized) design whose
subjects are organized into n / 2 matched pairs of size 2 (subjects
2k - 1 and 2k form pair k). Unlike
ObservationalDesignBlocks
there is no separate m argument — the pair structure is fixed by
subject order and installed automatically via set_m(), and
private$matching_capable is set so downstream Inference
classes use pair-preserving (not generic stratified) resampling (see class
documentation). As with every ObservationalDesign, w itself is
supplied afterward via
assign_w_to_all_subjects(w_precomputed = ...), in the same paired
subject order.
Usage
ObservationalDesignMatching$new( response_type, n, include_is_missing_as_a_new_feature = TRUE, verbose = FALSE, missingness_method = "impute", design_formula = ~., seed = NULL )
Arguments
response_type"continuous", "incidence", "proportion", "count", "survival", or "ordinal".
nThe sample size; must be even (subjects
2k - 1/2kform pairk, so an oddnwould leave one subject unpaired).include_is_missing_as_a_new_featureFlag for missingness indicators.
verboseA flag for verbosity.
missingness_methodHow to handle missing values in covariates.
design_formulaA formula object.
seedInteger seed for reproducibility.
Returns
A new 'ObservationalDesignMatching' object
ObservationalDesignMatching$clone()
The objects of this class are cloneable with this method.
Usage
ObservationalDesignMatching$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
des = ObservationalDesignMatching$new(response_type = 'continuous', n = 6)
des$add_all_subjects_to_experiment(data.frame(x1 = rnorm(6)))
des$assign_w_to_all_subjects(w_precomputed = c(1, 0, 0, 1, 1, 0))
Simulation Framework for Experimental Designs and Inference Methods
Description
An R6 class for benchmarking experimental designs and inference methods by
Monte Carlo simulation. Each replication generates synthetic covariates and
responses, runs every requested (design, inference) pair, and records
point estimates, confidence intervals, and p-values. Raw and aggregated
results are available through SimulationFrameworkReport.
Details
Covariates are drawn independently from \mathrm{Uniform}(0, 1).
-
cond_exp_func_model = "linear": the base continuous signal isy = X\betawhere\betais evenly spaced from 1 to-1. -
cond_exp_func_model = "nonlinear": the Friedman (1991) function10\sin(\pi x_1 x_2) + 20(x_3-0.5)^2 + 10x_4 + 5x_5; requiresp \ge 5.
The continuous base signal is transformed to the scale appropriate for
response_type. Treatment effects are applied per-subject: additive
on the linear/logit/ordinal scale, log-multiplicative for count and survival.
For each (design, inference) pair the framework runs whichever of the
following are supported by the inference class:
-
asymptotic (
InferenceAsympsubclasses): Wald CI and p-value. -
bootstrap (
InferenceNonParamBootstrapsubclasses): percentile CI and p-value. -
randomisation (
InferenceRandsubclasses): p-value; additionally a test-inversion CI forcontinuous,proportion, andcountresponse types (InferenceRandCIsubclasses).
Incompatible (design, inference) pairs (e.g.\ a KK-specific inference
class with a non-KK design) are silently skipped via tryCatch.
Reported summary metrics include:
-
MSE:
\overline{(\hat\beta_T - \beta_T)^2}over reps with a finite point estimate. -
coverage: proportion of reps where
\beta_Tlies inside the CI (NAwhen no CI is available for that inference type). -
power: proportion of p-values
< \alpha; equals the empirical type-I error rate whenbetaT = 0.
Methods
Public methods
SimulationFramework$new()
Create a new SimulationFramework
object that stores simulation design settings, response-generation
settings, inference methods, and replication controls.
Usage
SimulationFramework$new( response_type, design_classes_and_params = NULL, inference_classes_and_params = NULL, n = 100L, p = 5L, cond_exp_func_model = "linear", Nrep_W = 100L, Nrep_Y_w = 1L, betaT = 1, alpha = 0.05, B_boot = 201L, r_rand = 201L, pval_epsilon = 0.02, sd_noise = 1, n_ordinal_levels = 4L, proportion_epsilon = 1e-06, phi_proportion = 100, k_survival = 2, incidence_clamp = 1e-09, proportion_clamp = 1e-09, count_clamp = 1e-09, survival_clamp = 1e-09, survival_min_time = 0.1, count_min_rate = 0L, count_shift = 0, norm_sq_beta_vec = 1, X_mat = NULL, num_cores = 1L, seed = NULL, cov_draw_method = stats::rnorm, cov_draw_method_args = list(mean = 0, sd = 1), random_X_draws = TRUE, prob_censoring = 0.25, custom_replication_data_generator = NULL, custom_apply_treatment_and_noise = NULL, make_estimand_fn = NULL, dgp_params = list(), custom_dgp = NULL, verbose = TRUE, keep_all_intermediate_data = FALSE, turn_off_asserts_for_speed = TRUE, inference_types_and_params = NULL, results_filename = "simulation_framework_results.csv.bz2", continue_from_last_result_row = TRUE, reuse_cache = TRUE, stop_on_error = TRUE, save_to_disk_every_n_rep = 25L, save_model_control_fits = TRUE )
Arguments
response_type(required) Character scalar or vector. The type of outcome variable. One of
"continuous","incidence","proportion","count","survival","ordinal".design_classes_and_paramsNULL(default) or a list describing design classes and optional constructor parameters. Unnamed R6 class generators use default parameters, for examplelist(DesignSeqOneByOneKK21, DesignFixedBernoulli). Named entries use the entry name as the design class and the value as the parameter list, for examplelist(DesignSeqOneByOneUrn = list(alpha = 2, beta = 2)). Duplicate named entries are allowed for repeated designs with different parameters. Each generator must be constructable with onlyresponse_typeandnplus any extra params supplied in this list.NULLuses the package's standard design set. Designs requiringstrata_cols,cluster_col, orfactorshave sensible defaults auto-injected (first covariate column; second forcluster_col;list(treatment=2)forfactors) when not supplied in the parameter list. Example:design_classes_and_params = list( DesignSeqOneByOneKK21 = list(lambda = 0.5, t_0_pct = 0.1), DesignSeqOneByOneUrn = list(alpha = 2, beta = 2), DesignFixedBernoulli # default params )
Commonly useful design constructor parameters:
lambdaMatching-weight decay for
KK14/KK21/KK21stepwise.t_0_pctBurn-in fraction for
KK14/KK21/KK21stepwise.morrisonLogical; Morrison correction for
KK14.alpha,betaShape parameters for
DesignSeqOneByOneUrn.preferred_num_bins_for_continuous_covariateBin count for
DesignFixedBlockingandDesignFixedBlockedCluster.B_targetTarget number of blocks for
DesignFixedBlocking.
inference_classes_and_paramsNULL(default) or a list describing inference classes and optional constructor parameters. Unnamed R6 class generators use default parameters, for examplelist(InferenceContinOLS, InferenceContinKKOLSIVWC). Named entries use the entry name as the inference class and the value as the constructor parameter list, for examplelist(InferenceContinOLS = list(max_resample_attempts = 25L)). Duplicate named entries are allowed for repeated inference classes with different parameters. Supplied parameters must be accepted by the inference class constructor.NULLselects a curated set for the givenresponse_type: several universal classes that work with any design, plus representative KK-specific classes (silently skipped for non-KK designs at runtime).nInteger scalar or vector. Sample size per simulation replication. Default
100.pInteger scalar or vector. Number of covariates. Must be
\ge 5whencond_exp_func_model = "nonlinear". Default5.cond_exp_func_modelCharacter scalar or vector. How the latent continuous signal is constructed before transformation to the
response_typescale."linear"Linear combination
X\betawith coefficients evenly spaced from 1 to-1."nonlinear"Friedman (1991) function
10\sin(\pi x_1 x_2)+20(x_3-0.5)^2+10x_4+5x_5; requiresp \ge 5.
Default
"linear".Nrep_WPositive integer. Number of treatment-assignment draws (w-reps). Each w-rep generates a fresh covariate matrix
Xand a new treatment assignment vectorw. Default100L.Nrep_Y_wPositive integer. Number of draws of the response per draw of
w, the allocation vector. For eachw-rep,Nrep_Y_windependent response vectorsyare drawn from the same(X, w). The effective total number of replications recorded isNrep_W * Nrep_Y_w. Default1(standard behaviour: one outcome draw per allocation-vector draw).betaTNumeric scalar or vector. True treatment effect added to treated subjects' outcomes. The scale is response-type specific: additive for
continuous,proportion, andordinal; on the logit scale forincidence; log-multiplicative forcountandsurvival. Default1. SetbetaT = 0to check type-I error.alphaNumeric in
(0,1). Significance level used for all confidence intervals and for computing power (p < \alpha). Default0.05.B_bootPositive integer. Bootstrap resamples per CI / p-value call. Default
201.r_randPositive integer. Randomisation draws per rand p-value call, and per bisection step of the rand CI. Default
201.pval_epsilonNumeric. Bisection convergence tolerance for randomisation-based CIs (
compute_rand_confidence_interval). Default0.02.sd_noiseNumeric
> 0. Standard deviation of independent Gaussian noise added to each subject's outcome. Default1.n_ordinal_levelsPositive integer. Number of ordinal categories when
response_type = "ordinal". Default4L.proportion_epsilonNumeric scalar. Small value added to proportion base responses to avoid 0 and 1. Default
1e-6.phi_proportionPositive numeric scalar. Precision parameter for beta-distributed observed proportion outcomes. The beta mean is
y_linear_model[i] + betaT * w[i]. Default100.k_survivalPositive numeric scalar. Scale parameter passed to the Weibull draw for observed survival outcomes. Default
2.incidence_clampNumeric scalar in
(0, 0.5). Clamp applied to the Bernoulli probability for observed incidence outcomes. Default1e-9.proportion_clampNumeric scalar in
(0, 0.5). Clamp applied to the beta mean for observed proportion outcomes. Default1e-9.count_clampPositive numeric scalar. Minimum Poisson mean for observed count outcomes. Default
1e-9.survival_clampPositive numeric scalar. Minimum Weibull shape for observed survival outcomes. Default
1e-9.survival_min_timeNumeric scalar. Minimum survival time and shift for base responses. Default
0.1.count_min_rateInteger scalar. Minimum baseline rate for count responses. Default
0L.count_shiftNumeric scalar. Constant added to counts after zero-centering for base responses. Default
0.norm_sq_beta_vecPositive numeric scalar. The desired squared Euclidean norm of the latent linear coefficient vector
\beta. The generated vector is scaled to match this norm. Default1.X_matNumeric matrix of dimensions
n x p, orNULL(default). If provided, these fixed covariates are used for every replication. In this case,cov_draw_methodmust beNULL.num_coresPositive integer. Number of worker processes for parallel execution of Monte Carlo replications. Note that when
num_cores > 1, parallelization *within* individual inference routines (e.g. bootstrap, randomization) is automatically disabled to prevent thread oversubscription.Unix/Linux (recommended): A
makeForkClusterpool is created once atrun()start. Workers inherit all pre-generated design and SE caches via copy-on-write with zero serialization overhead. Parallelism operates at the replication level:num_coresreplications run simultaneously, each executing all DGP cells serially. This eliminates the per-batch dispatch overhead that would arise from cycling through cells within every replication, and keeps all cores fully subscribed regardless of the number of DGP cells. For best performance, run on a Unix/Linux machine and setnum_coresto the number of physical cores available.Non-Unix (Windows/macOS): mirai daemons are used when available. Every active (replication, DGP-cell) pair is one work unit and
num_coresunits are kept in flight continuously, so all cores stay subscribed even when the grid has fewer DGP cells than cores. Cell state is pushed to the daemons once atrun()start rather than re-serialized per dispatch. If mirai is not installed, execution falls back to serial with a warning. Default1.seedInteger or
NULL(default). Random seed for the entire simulation run.cov_draw_methodA function used to draw
n * pi.i.d. covariate values for every replication. The function must accept the total number of values as its first argument, followed by arguments incov_draw_method_args. Defaultstats::rnorm. Must beNULLwhenX_matis supplied.cov_draw_method_argsNamed list of additional arguments forwarded to
cov_draw_methodbeyond the sample-size first argument. Default islist(mean = 0, sd = 1).random_X_drawsLogical. If
TRUE(default), a new set of covariates is drawn for every single replication. IfFALSE, one set is drawn per(n, p)cell and shared across its replications.prob_censoringNumeric in
[0,1]. Per-subject independent censoring probability; applied only whenresponse_type = "survival". Default0.25.custom_replication_data_generatorOptional function for custom replication data. When supplied, it is called as
fn(state, rep)and must return a list containing at leastXandy_linear_model. Any additional fields in the returned list (e.g. latent frailty draws) are passed forward asrep_datatocustom_apply_treatment_and_noiseand the function built bymake_estimand_fn.custom_apply_treatment_and_noiseOptional function for custom response generation. Signature:
fn(y_linear_model, w, rep_data, state).wuses {0, 1} encoding, with 1 for treatment and 0 for control.rep_datais the full list returned bycustom_replication_data_generator(orNULLfor the standard path). Must return a list with componentsyanddead. Three-argument functionsfn(y_linear_model, w, state)are still accepted for backwards compatibility.make_estimand_fnOptional factory function for a custom true estimand. Signature:
fn(beta_T), returning a function with signaturefn(y_linear_model, X, w, rep_data, state). Called once per grid cell with that cell'sbeta_Tso the returned estimand function is always tied to the right effect size (important whenbetaTis a vector of multiple values). The returned function is invoked once per design class per replication after the design completes, sow(in {0, 1} encoding) andXreflect the realized assignment. Must return a numeric scalar. When supplied, its return value is used as the ground truth for all inference classes (overriding theis_mean_diffgate). Three-argument functionsfn(y_linear_model, state)are still accepted for backwards compatibility (they will not receiveXorw). DefaultNULLusesbeta_Tdirectly as the ground truth.dgp_paramsOptional named list of DGP configuration values (e.g.
list(frailty_dist = "gamma", censoring_rate = 0.8)). Injected intostateasstate\$dgp_paramsand accessible in all three custom-DGP hooks. Recommended over using closures to pass DGP parameters.custom_dgpOptional function for a fully custom DGP. Signature:
fn(n, p, rep, state)returning a list with componentsX(data.frame,nxp),w(integer vector in {0, 1}, lengthn),y(numeric, lengthn),dead(integer {0,1} orNULLfor non-survival),true_estimand(numeric scalar, optional). When supplied, the design class acts as a data container only; it does not run its own randomization or matching. Requires a fixed design class (notDesignSeqOneByOnevariants). Cannot be combined withcustom_replication_data_generatororcustom_apply_treatment_and_noise.verboseLogical. If
TRUE, prints a message for every replication and for every(design, inference)pair that is skipped due to an error. DefaultTRUE.keep_all_intermediate_dataLogical. If
TRUE, the framework saves the instantiated design and inference objects for every replication. These can be retrieved after the run using$get_all_intermediate_data(). Warning: this can consume a lot of memory for many replications. DefaultFALSE.turn_off_asserts_for_speedLogical. If
TRUE(default), all checkmate assertions across the package are globally disabled during the simulation run to improve performance.inference_types_and_paramsNULL(default) or a named list from inference type to a named list of arguments for that type's function invocation. The list names control which inference outputs are computed. Valid names are"asymp_ci","asymp_pval","exact_ci","exact_pval","boot_ci","boot_pval","rand_ci", and"rand_pval". Each value must be a named list whose names are accepted by the corresponding inference function.NULLruns all eight types with default invocation arguments. Example:inference_types_and_params = list( asymp_pval = list(delta = 0), boot_ci = list(B = 99, type = "perc"), rand_pval = list(r = 999, transform_responses = TRUE) )
When no
*_citype is requested,coverageis omitted fromSimulationFrameworkReport$summarize(). When no*_pvaltype is requested,poweris omitted.results_filenameCharacter scalar. The filename for the results file. Supported extensions are
.csvand.csv.bz2. Default"simulation_framework_results.csv.bz2".continue_from_last_result_rowLogical. If
TRUE(default), the framework loads existing results fromresults_filenameand skips previously completed replications.reuse_cacheLogical. If
TRUE(default), expensive pre-generated design / SE cache objects are loaded from disk when available. IfFALSE, these cache objects are regenerated from scratch, but each regenerated object is still saved to disk for later restarts.stop_on_errorLogical. If
TRUE(default), any error raised during a simulation path aborts the run immediately. IfFALSE, the framework records the error, skips the failing path, and continues with the remaining replications / design / inference combinations. Use$get_errors()after$run()to inspect the captured errors.save_to_disk_every_n_repPositive integer. Results are flushed to the on-disk staging file only once every this many replications, and always after the final replication. Larger values reduce disk I/O overhead at the cost of losing more progress if the run is interrupted. Default
25L.save_model_control_fitsLogical. If
TRUE(default), after the design/SE cache is built, saves the per-subject model-implied potential outcomes under treatment and control as CSV files in a subfolder named<stem>_response_values/(where<stem>isresults_filenamewith its.csv/.csv.bz2extension stripped) next toresults_filename. One file is written per unique(response_type, cond_exp_func_model, n, p, betaT)cell. Only meaningful whenrandom_X_draws = FALSE; silently skipped otherwise. Column names depend onresponse_type:"continuous","survival"columns
ytandyc"incidence","proportion"columns
ptandpc"count"columns
rtandrc
Default
TRUE.
SimulationFramework$run()
Execute the configured simulation replications, run each
requested design and inference method, collect estimates/p-values/CIs and
errors, and return a
SimulationFrameworkReport.
Usage
SimulationFramework$run()
Returns
The SimulationFramework object itself (invisibly).
SimulationFramework$get_all_intermediate_data()
Retrieve the stored intermediate data (design and inference objects)
for every replication. Only available if keep_all_intermediate_data = TRUE
was passed to the constructor.
Usage
SimulationFramework$get_all_intermediate_data()
Returns
A nested list containing the intermediate data for each replication,
or NULL if not recorded.
SimulationFramework$clear_all_intermediate_data_and_gc()
Release all stored intermediate data and invoke the garbage collector.
Useful after inspecting intermediate results to free memory before
further processing. Sets the internal store to NULL and calls
gc().
Usage
SimulationFramework$clear_all_intermediate_data_and_gc()
Returns
The SimulationFramework object itself (invisibly).
SimulationFramework$clone()
The objects of this class are cloneable with this method.
Usage
SimulationFramework$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
# Simple simulation with two designs and two inference methods.
# n/Nrep_W/num_boot/B_boot/r_rand are kept small so this example runs in a
# few seconds; a real simulation would use much larger values (this
# class's defaults, or larger still) for adequately precise estimates.
sim = SimulationFramework$new(
response_type = "continuous",
design_classes_and_params = list(
DesignSeqOneByOneKK21 = list(lambda = 0.5, num_boot = 50L),
DesignSeqOneByOneBernoulli = list()
),
inference_classes_and_params = list(
InferenceContinOLS = list(),
InferenceContinKKOLSIVWC = list()
),
n = 20, p = 3, Nrep_W = 2L, betaT = 1, B_boot = 50L, r_rand = 101L,
results_filename = tempfile(fileext = ".csv.bz2"),
continue_from_last_result_row = FALSE
)
sim$run()
report = SimulationFrameworkReport$new(sim)
report$summarize()
Reporting class for SimulationFramework results
Description
An R6 class for accessing and summarizing the results of a
SimulationFramework run. It can be constructed either from a completed
SimulationFramework object (via SimulationFrameworkReport$new(sim))
or by loading results from a previously saved CSV / CSV.BZ2
file (via SimulationFrameworkReport$new("path/to/results.csv")).
Details
When constructed from a SimulationFramework object all
design/inference parameter metadata is preserved, so $summarize() can
annotate each row with human-readable parameter strings. When constructed from
a file only the raw results are available; parameter annotation columns will be
empty strings.
Methods
Public methods
SimulationFrameworkReport$new()
Create a new
SimulationFrameworkReport
object that stores simulation results, captured errors, and summary
helpers returned by SimulationFramework.
Usage
SimulationFrameworkReport$new(sim_or_filename, alpha = NULL)
Arguments
sim_or_filenameEither a completed
SimulationFrameworkobject or a character string giving the path to a.csvor.csv.bz2results file written bySimulationFramework.alphaNumeric in
(0,1). Significance level for coverage and power calculations. Whensim_or_filenameis aSimulationFrameworkobject andalphaisNULL(default), the framework's own alpha is used. When loading from a file, defaults to0.05.
SimulationFrameworkReport$get_results()
Get the raw per-replication results.
Usage
SimulationFrameworkReport$get_results()
Returns
A data.table with one row per
(replication, design, inference class, inference type).
SimulationFrameworkReport$get_errors()
Return all errors captured during the simulation run.
Usage
SimulationFrameworkReport$get_errors()
Returns
A list of named lists, one per captured error. Each element includes the simulation cell metadata, replication number, design / inference path, user-supplied parameters, error stage, and error message. Empty when constructed from a file.
SimulationFrameworkReport$summarize()
Aggregate and summarize simulation results.
Usage
SimulationFrameworkReport$summarize()
Returns
A data.table with one row per unique
(response_type, cond_exp_func_model, n, p, betaT, design, inference,
inference_type) combination. Columns include MSE,
coverage, ci_length, and coverage_pval (when CI
types were run; coverage_pval is the exact two-sided binomial
test p-value of H0: true coverage = 1 - alpha),
power (when betaT != 0 and p-value types were run),
size and size_pval (when betaT == 0 and p-value types
were run; size_pval is the exact two-sided binomial test
p-value of H0: true size = alpha, suitable for multiplicity-corrected
calibration checks across settings), and parameter annotation strings.
SimulationFrameworkReport$print()
Print a concise summary of the report.
Usage
SimulationFrameworkReport$print()
SimulationFrameworkReport$clone()
The objects of this class are cloneable with this method.
Usage
SimulationFrameworkReport$clone(deep = FALSE)
Arguments
deepWhether to make a deep clone.
Examples
sim <- SimulationFramework$new(
response_type = "continuous",
design_classes_and_params = list(DesignFixedBernoulli),
inference_classes_and_params = list(InferenceAllSimpleAverageDiff),
n = 20L, Nrep_W = 5L, betaT = 1,
results_filename = tempfile(fileext = ".csv"),
verbose = FALSE, continue_from_last_result_row = FALSE
)
sim$run()
report <- SimulationFrameworkReport$new(sim)
report$get_results()
report$summarize()
GLM and Kaplan-Meier Inference
Description
Computes the treatment estimate using the underlying model.
Computes an asymptotic confidence interval using the configured test.
Computes an asymptotic two-sided p-value using the configured test.
Usage
StandardModelCacheSource
Details
Abstract class providing MLE/KM-based inference methods for GLM and survival models.
Value
A confidence interval.
The asymptotic p-value.
Dependent-censoring transformation component source
Description
Initialize inference for the bivariate log-normal
dependent-censoring transformation model; see
InferenceSurvivalDepCensTransformRegr
for the model form. Does not fit the model; the fit is deferred to
the first call to compute_estimate() or a method that requires
it.
Fits the joint bivariate log-normal event/censoring
transformation model by maximum likelihood and returns the event-time
log-time-ratio estimate \hat\beta_T; see
InferenceSurvivalDepCensTransformRegr
for the model form.
Recomputes the treatment estimate under subject/block-level
Bayesian-bootstrap weights, using weighted_cox_bootstrap_surrogate_fit()
— a fast weighted Cox-model surrogate fit — as an approximation to the
weighted joint dependent-censoring likelihood, rather than a full
weighted refit of the joint bivariate model.
Reports jackknife bias correction as unavailable for this model; leave-one-out bias correction is unstable for the dependent-censoring transformation likelihood on small censored samples.
Reports jackknife bias estimate as unavailable for this model.
Reports jackknife standard error as unavailable for this model.
Reports jackknife Wald two-sided p-value as unavailable for this model.
Reports jackknife Wald confidence interval as unavailable for this model.
Reports randomization inference as unavailable for this model; each randomization draw requires a full dependent-censoring likelihood refit and is not stable enough for the comprehensive suite.
Reports randomization confidence interval as unavailable for this model.
Bootstrap confidence interval, validated for this model.
Basic bootstrap confidence interval, validated for this model.
BCa bootstrap confidence interval, validated for this model.
Studentized bootstrap confidence interval, validated for this model.
Reports the randomization distribution as unavailable for this model.
Wald confidence interval for the event-submodel log-time-ratio
\beta_T using the fitted joint model's standard error; see
InferenceAsymp for the shared Wald
contract. Fits the model first if not already cached.
Two-sided Wald test of H_0: \beta_T = \code{delta} for
the event-submodel log-time-ratio, using the fitted joint model's
standard error; see InferenceAsymp
for the shared Wald contract. Fits the model first if not already
cached.
Computes a score two-sided p-value, falling back to the asymptotic test when unavailable.
Computes a likelihood-ratio confidence interval, reporting unstable inversion failures as explicitly non-estimable.
Usage
SurvivalDepCensTransformSource
Details
Source list for the dependent-censoring transformation survival component.
GLMM Weibull log-gamma-frailty IVWC component source
Description
Initialize KK Clayton-copula survival inference and prepare
the matched/reservoir likelihood components used by
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC.
Returns the model-specific log-time-ratio treatment estimate; see
Inference.
Recomputes the IVWC Clayton-copula treatment estimate under Bayesian-bootstrap weights.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Compute the Clayton-copula survival asymptotic p-value for
the treatment effect using the fitted frailty/dependence model. See
InferenceAsymp for shared p-value
semantics.
Usage
SurvivalGLMMWeibullFrailtyLoggammaIVWCSource
Details
Source list for the KK inverse-variance-weighted-combination (IVWC) Weibull log-gamma-frailty survival component.
Clayton Copula Combined-Likelihood Inference for KK Designs
Description
Initialize KK Clayton-copula one-likelihood survival
inference and prepare the combined likelihood used by
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik.
Returns the model-specific log-time-ratio treatment estimate; see
Inference.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Compute the one-likelihood Clayton-copula survival
asymptotic p-value for the treatment effect using the fitted dependence
model. See InferenceAsymp.
Duplicates this subclass while preserving fit caches; see
Inference.
Usage
SurvivalGLMMWeibullFrailtyLoggammaOneLikSource
Details
Gamma-frailty (Clayton copula) Weibull estimator; see
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC
for the frailty-distribution details and contrast with the log-normal-frailty
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik
alternative.
Abstract class for Weibull Frailty / Standard Weibull Compound Inference
Description
Initialize KK Weibull-frailty IVWC survival inference and
prepare matched/reservoir parametric survival components used by
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC.
Returns the model-specific log-time-ratio treatment estimate; see
Inference.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Compute the Weibull-frailty asymptotic p-value for the
treatment effect using the fitted parametric survival likelihood. See
InferenceAsymp.
Usage
SurvivalGLMMWeibullFrailtyNormalIVWCSource
Details
This class implements a compound estimator for KK matching-on-the-fly designs with survival responses using a Weibull AFT GLMM for matched pairs. The matched-pair component uses a shared log-normal random intercept per pair, fitted by the package's native Rcpp likelihood optimizer. The reservoir component uses standard Weibull AFT regression. The two treatment-effect estimates are combined by inverse-variance weighting.
This compound estimator accounts for the dependence within matched pairs by modeling it as a shared frailty.
Frailty distribution. The matched-pair likelihood is an AFT
(accelerated failure time) parameterization, log(T) = X beta + u +
sigma_eps * epsilon, where epsilon is standard extreme-value (giving
Weibull margins) and the pair-shared random intercept u is Gaussian,
u ~ N(0, sigma_u^2). On the natural time scale this is a multiplicative
log-normal frailty, exp(u). A Gaussian random effect has no
closed-form marginal likelihood under a Weibull baseline, so the pair
likelihood is evaluated by Gauss-Hermite quadrature (fast_weibull_frailty_cpp)
rather than in closed form.
This is a different (and equally standard) frailty assumption from the classic
gamma-frailty Weibull model (Clayton 1978; Vaupel, Manton & Stallard
1979; Hougaard 2000), which multiplies the hazard (not the AFT error) by a
shared Gamma(1/theta, 1/theta) term and has a closed-form marginal
survival function via the frailty's Laplace transform. That model is implemented
in this package as the Clayton copula of
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC /
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik
(the Clayton copula with Weibull margins is exactly the closed-form bivariate
survival function obtained by integrating a shared gamma frailty out of two
conditionally-independent Weibull hazards). Prefer this log-normal-frailty class
for a Gaussian-random-intercept / GLMM-style dependence structure; prefer the
Clayton-copula class for the classic gamma-frailty / proportional-hazards
dependence structure.
Univariate (ncol(as.matrix(private$X)) == 0): uses the native
Rcpp Weibull frailty likelihood with formula = survival::Surv(y, dead) ~ w
and a pair-level random intercept.
Multivariate (ncol(as.matrix(private$X)) > 0): fits the same
native Rcpp likelihood with covariate adjustment, dropping rank-deficient
columns when needed.
See Also
InferenceSurvivalGLMMWeibullFrailtyLoggammaIVWC
for the corresponding gamma-frailty (Clayton copula) IVWC estimator.
Weibull Frailty Combined-Likelihood Inference for KK Designs
Description
Initialize the one-likelihood Weibull-frailty inference object.
Usage
SurvivalGLMMWeibullFrailtyNormalOneLikLeafSource
Details
Log-normal (Gaussian random-intercept) frailty Weibull AFT estimator; see
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik
for the frailty-distribution details and contrast with the gamma-frailty
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik
(Clayton copula) alternative.
Abstract class for Weibull Frailty Combined-Likelihood Inference
Description
Initialize KK Weibull-frailty one-likelihood survival
inference and prepare the combined parametric survival likelihood used by
InferenceSurvivalGLMMWeibullFrailtyNormalOneLik.
Returns the model-specific combined-likelihood treatment estimate; see
InferenceAsympLik.
Recomputes the one-likelihood Weibull-frailty treatment estimate under Bayesian-bootstrap weights.
Computes an asymptotic confidence interval for the treatment effect.
Returns a 2-sided p-value for H0: beta_T = delta.
Usage
SurvivalGLMMWeibullFrailtyNormalOneLikSource
Details
One-likelihood (combined matched-pair + reservoir) analog of
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC:
same Weibull-AFT-with-log-normal-random-intercept (Gaussian, Gauss-Hermite
quadrature) frailty assumption for matched pairs, but the matched-pair and
reservoir contributions are fit as a single combined likelihood rather than
combined by inverse-variance weighting. See
InferenceSurvivalGLMMWeibullFrailtyNormalIVWC
for the frailty-distribution details and its contrast with the gamma-frailty
Clayton copula model implemented by
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik.
See Also
InferenceSurvivalGLMMWeibullFrailtyLoggammaOneLik
for the corresponding gamma-frailty (Clayton copula) combined-likelihood estimator.
Abstract class for Survival Rank-based Regression (AFT) Compound Inference
Description
Initialize KK survival rank-regression IVWC inference and
prepare matched/reservoir Gehan-Wilcoxon rank-regression components used
by InferenceSurvivalKKRankRegrIVWC.
Returns the model-specific log-time-ratio treatment estimate; see
Inference.
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Compute the survival rank-regression asymptotic p-value for
the treatment effect using the fitted rank-regression estimate and
standard error. See InferenceAsymp.
Usage
SurvivalKKRankRegrIVWCSource
Details
This class implements a robust compound estimator for KK matching-on-the-fly designs with survival responses using rank-based estimating equations via the aftgee package. For matched pairs, it fits a rank-based AFT model with clustering. For reservoir subjects, it fits a standard rank-based AFT model. The two estimates (both log-time ratios) are combined via a variance-weighted linear combination.
This class requires the aftgee package. Under harden = TRUE,
multivariate component fits preserve the treatment column and retry reduced
covariate sets after QR-based rank reduction and correlation-based pruning.
Extreme finite coefficients / standard errors are rejected and treated as
non-estimable.
KK stratified Cox IVWC component source
Description
Initialize KK stratified Cox IVWC inference and prepare the
matched/reservoir partial-likelihood components used by
InferenceSurvivalKKStratCoxPHIVWC.
Returns the estimated treatment effect (log-hazard ratio).
Uses the shared asymptotic confidence-interval contract; see
InferenceAsymp.
Compute the stratified-Cox asymptotic p-value for the
treatment log-hazard ratio using the fitted partial-likelihood standard
error. See InferenceAsymp.
Usage
SurvivalKKStratCoxIVWCSource
Details
Source list for the KK stratified Cox inverse-variance-weighted-combination (IVWC) survival component.
Stratified Cox Combined-Likelihood Compound Inference for KK Designs
Description
Initialize KK stratified Cox one-likelihood inference and
prepare the combined partial-likelihood fit used by
InferenceSurvivalKKStratCoxPHOneLik.
Returns the model-specific combined-likelihood treatment estimate; see
InferenceAsympLik.
Recomputes the one-likelihood stratified Cox treatment estimate under Bayesian-bootstrap weights.
Computes an asymptotic confidence interval for the treatment effect.
Returns a 2-sided p-value for H0: beta_T = delta.
Usage
SurvivalKKStratCoxOneLikPartialLikelihoodSource
Weibull likelihood component source
Description
Initialize inference for the Weibull AFT model
\log T_i = \beta_0 + \beta_T W_i + X_i^\top \gamma + \sigma
\epsilon_i, \epsilon_i \sim standard extreme-value (so T_i
is marginally Weibull); see
InferenceSurvivalWeibullRegr
for the model form. Does not fit the model; the fit is deferred to the
first call to compute_estimate() or a method that requires it.
Fits the Weibull AFT model by maximum likelihood and returns
the log-time-ratio estimate \hat\beta_T. Handles right-, left-,
and interval-censored observations via their appropriate
survival/density likelihood contributions.
Recomputes the treatment estimate under subject/block-level
Bayesian-bootstrap weights via
weighted_weibull_bootstrap_surrogate_fit(), a fast weighted
Weibull surrogate fit, as an approximation to the weighted AFT
likelihood. Only supported for ordinary right-censored data:
throws an error for left-/interval-censored designs, since the
surrogate fit assumes ordinary right-censoring semantics.
Wald confidence interval for the log-time-ratio \beta_T
using the fitted model's standard error; see
InferenceAsymp for the shared Wald
contract. Fits the model first if not already cached.
Two-sided Wald test of H_0: \beta_T = \code{delta}
using the fitted model's standard error; see
InferenceAsymp for the shared Wald
contract. Fits the model first if not already cached.
Bayesian-bootstrap two-sided p-value; see
Inference. Blocked outright (rather than
letting the underlying weighted re-estimate fail once per replicate,
inside a per-iteration tryCatch that silently converts the
stop() in compute_estimate_with_bootstrap_weights() into an
all-NA bootstrap distribution) under left-/interval-censored data.
Bartlett-corrected likelihood-ratio two-sided p-value; see
Inference. Blocked outright under left-/
interval-censored data for the same reason as
compute_bayesian_bootstrap_two_sided_pval(): the underlying
simulate_under_lik_null() guard would otherwise only surface as a
silently all-NA calibration distribution.
Randomization-test two-sided p-value; see
Inference. A nonzero null shift
(delta != 0) is blocked outright under left-/interval-censored
data for the same reason as the other bootstrap guards in this class:
the underlying shift-template guard in
setup_randomization_template_and_shifts() would otherwise only
surface as a silently all-NA randomization distribution.
Usage
SurvivalWeibullLikelihoodSource
Details
Source list for the SurvivalWeibullLikelihood component composed by InferenceSurvivalWeibullRegr.
Bisection loop for computing confidence interval bounds by inverting randomization tests
Description
This function implements the bisection algorithm to find CI bounds by inverting the randomization test. It repeatedly calls the p-value computation function until convergence.
Usage
bisection_ci_loop_cpp(
pval_fn,
r,
l,
u,
pval_th,
tol,
transform_responses,
lower
)
Arguments
pval_fn |
R function that computes two-sided p-value given delta |
r |
Number of randomization iterations |
l |
Initial lower bound |
u |
Initial upper bound |
pval_th |
P-value threshold (typically alpha/2 for two-sided CI) |
tol |
Tolerance for convergence (in p-value space) |
transform_responses |
String: "none", "log", or "logit" |
lower |
Logical: TRUE for lower CI bound, FALSE for upper |
Value
The CI bound value
Sequential computation of both CI bounds (called from R-level parallelism)
Description
This function is kept for backwards compatibility but the outer parallelism (running lower/upper bounds simultaneously) is now handled at the R level via parallel::mclapply, which is safe. Calling R functions from OpenMP threads is undefined behaviour in R and caused process crashes.
Usage
bisection_ci_parallel_cpp(
pval_fn,
r,
l_lower,
u_lower,
l_upper,
u_upper,
pval_th,
tol,
transform_responses,
num_cores = 1L
)
Arguments
pval_fn |
R function: pval_fn(nsim, delta, transform_responses, num_cores) |
r |
Number of randomization iterations |
l_lower |
Initial lower bound for lower CI bound search |
u_lower |
Initial upper bound for lower CI bound search (typically the estimate) |
l_upper |
Initial lower bound for upper CI bound search (typically the estimate) |
u_upper |
Initial upper bound for upper CI bound search |
pval_th |
P-value threshold (typically alpha/2 for two-sided CI) |
tol |
Tolerance for convergence (in p-value space) |
transform_responses |
String: "none", "log", or "logit" |
num_cores |
Passed through to pval_fn for inner parallelism |
Value
Numeric vector of length 2: [lower_bound, upper_bound]
Single-threaded helper for computing one CI bound
Description
Single-threaded helper for computing one CI bound
Usage
bisection_ci_single_bound_cpp(
pval_fn,
r,
l,
u,
pval_th,
tol,
transform_responses,
lower,
num_cores = 1L
)
Build a Reusable Unstratified Cox Data Cache (C++ Backend)
Description
Precomputes and caches the sorted risk-set structure needed to evaluate the
Cox partial-likelihood, score, and Hessian, so that repeated Newton-Raphson
or L-BFGS fits on the same (X, y, dead) data (e.g. across
bootstrap/randomization replicates of the treatment column, or successive
estimate_only vs. full-variance calls) do not repeat the
O(n \log n) sort and event-time tabulation on every call. This is the
unstratified counterpart of build_stratified_cox_data_cache_cpp;
the returned cache is consumed by fast_coxph_regression_prebuilt_cpp,
which implements the same Cox partial-likelihood as
fast_coxph_regression and fast_coxph_regression_cpp.
Usage
build_cox_data_cache_cpp(X, y, dead)
Arguments
X |
A numeric matrix of predictor variables (no intercept column);
|
y |
A numeric vector of length |
dead |
A numeric vector of length |
Details
Model. No fitting happens here; this function only prepares the single risk set (stratum) used by the Breslow-tied Cox partial likelihood
\ell(\beta) = \sum_{k} \Big[ \big(\textstyle\sum_{i \in D_k} x_i\big)^\top \beta - d_k \log\!\big(\textstyle\sum_{j \in R_k} e^{x_j^\top \beta}\big) \Big],
where D_k is the set of subjects with an event at the k-th
unique event time and R_k is the risk set (all subjects with
y >= that event time). See
fast_coxph_regression for the full model, scale, and
optimizer contract; this page documents only the cache construction.
What is cached. Internally builds a single CoxData record
that: (1) sorts subjects by ascending y (observed/censoring time),
breaking ties by placing events (dead == 1) before censored
observations at the same time; (2) stores the sort-permuted y,
dead, and row-major copy of X; and (3) tabulates the vector
of unique event times and, for each, the number of tied events
(event_counts), which drives the Breslow tie-handling correction in
the partial-likelihood, score, and Hessian.
Input conventions. X is an n \times p design matrix
with one row per subject and no intercept column (Cox models are
fit on the partial likelihood, which has no intercept). y is the
observed time (event or censoring time), and dead is a 0/1 event
indicator (1 = event, 0 = right-censored); both must have length n
matching nrow(X). Tied event times are handled via the Breslow
approximation (via event_counts), not the exact (Cox) or Efron
method. There is no support for left- or interval-censored data at this
layer; classes with more general censoring (see
supports_interval_or_left_censored_data() in
InferenceEngine) bypass this cache and dispatch to
icenReg instead. NA/non-finite values in X, y,
or dead and negative values of y are not checked or handled
at this layer; callers are responsible for filtering or imputing upstream.
Return value and object lifetime. Returns an externalptr
(Rcpp::XPtr) wrapping a heap-allocated std::vector<CoxData>
of length 1 (a single unstratified stratum), so that
fast_coxph_regression_prebuilt_cpp can share the same stratified
interface as the stratified cache. The pointer owns its memory (finalizer
registered via XPtr(..., true)) and is freed automatically by R's
garbage collector; it must not be serialized (e.g. via saveRDS)
or reused after the R session that created it exits. The cache is
immutable once built and safe to reuse across many calls to
fast_coxph_regression_prebuilt_cpp as long as X, y,
and dead have not changed; callers (e.g.
InferenceCoxPH$private$cox_data_cache) are responsible for
invalidating and rebuilding the cache when the treatment assignment or
covariates change.
Complexity. O(n \log n) for the sort plus O(n) for
the tabulation pass, where n is the number of subjects; memory use
is O(np) for the row-major copy of X plus O(n) for the
sorted y/dead and event-time tables.
Value
An externalptr to a cached, sorted Cox risk-set
representation of (X, y, dead), for use as the cox_data_xptr
argument of fast_coxph_regression_prebuilt_cpp.
See Also
build_stratified_cox_data_cache_cpp for the
stratified (multiple risk-set) analog, fast_coxph_regression_prebuilt_cpp
for the fitting routine that consumes this cache, and
fast_coxph_regression for the full Cox model documentation,
including the partial-likelihood derivation, tie-handling, and references.
Analogous Python API: lifelines CoxPHFitter.
Build a Reusable Stratified Cox Data Cache (C++ Backend)
Description
Precomputes and caches, per stratum, the sorted risk-set structure needed
to evaluate the stratified Cox partial-likelihood, score, and Hessian, so
that repeated Newton-Raphson or L-BFGS fits on the same
(X, y, dead, strata) data (e.g. across bootstrap/randomization
replicates of the treatment column) do not repeat the per-stratum sort and
event-time tabulation on every call. This is the stratified counterpart of
build_cox_data_cache_cpp; the returned cache is consumed by
fast_coxph_regression_prebuilt_cpp, which implements the
same partial-likelihood family as fast_coxph_regression and
fast_coxph_regression_cpp, extended to sum the log-partial-
likelihood, score, and Hessian across independent strata-specific risk
sets while sharing a single regression coefficient vector \beta
across all strata (a shared-\beta, per-stratum-baseline-hazard
stratified Cox model).
Usage
build_stratified_cox_data_cache_cpp(X, y, dead, strata)
Arguments
X |
A numeric matrix of predictor variables (no intercept column);
|
y |
A numeric vector of length |
dead |
A numeric vector of length |
strata |
An integer vector of length |
Details
Model. No fitting happens here; this function only partitions subjects into per-stratum risk sets and prepares each one for the Breslow-tied stratified Cox partial likelihood
\ell(\beta) = \sum_{s} \sum_{k \in s} \Big[ \big(\textstyle\sum_{i \in D_{sk}} x_i\big)^\top \beta - d_{sk} \log\!\big(\textstyle\sum_{j \in R_{sk}} e^{x_j^\top \beta}\big) \Big],
where the outer sum runs over strata s (distinct baseline hazards),
D_{sk} is the set of subjects with an event at the k-th unique
event time within stratum s, and R_{sk} is the corresponding
within-stratum risk set (subjects in stratum s with y >=
that event time). Risk sets never cross strata, so subjects are only ever
compared to other subjects in the same stratum; \beta is shared
across strata while the baseline hazard is allowed to differ arbitrarily
by stratum. See fast_coxph_regression for the unstratified
model, scale, and optimizer contract; this page documents only the
per-stratum cache construction.
What is cached. Subjects are grouped into strata by their
integer strata label (via an ordered std::map<int, ...>,
so strata are processed and stored in ascending label order). Within each
stratum, a CoxData record is built exactly as in
build_cox_data_cache_cpp: subjects are sorted by ascending
y, ties broken by placing events before censored observations,
and the per-stratum unique event times and tied-event counts
(event_counts) are tabulated to drive the within-stratum Breslow
tie-handling correction.
Input conventions. X is an n \times p design matrix
with one row per subject and no intercept column. y is
the observed time (event or censoring time) and dead is a 0/1
event indicator (1 = event, 0 = right-censored); both have length
n matching nrow(X). strata is a length-n
integer vector of stratum (block/cluster) labels; labels need not be
contiguous or start at 1, and a stratum with a single subject or with no
observed events contributes an empty or degenerate risk set that carries
no information to the partial likelihood, score, or Hessian (its rows
are still stored, but they will not affect the fit). Callers are
expected to have already dropped uninformative strata upstream (see
get_informative_rows() in the stratified-Cox inference class) when
that matters for numerical stability; this function does not filter
strata itself. As with build_cox_data_cache_cpp, tied
event times use the Breslow approximation, there is no support for left-
or interval-censored data at this layer, and NA/non-finite values
in X, y, dead, or strata are not checked or
handled here.
Return value and object lifetime. Returns an externalptr
(Rcpp::XPtr) wrapping a heap-allocated
std::vector<CoxData> with one element per distinct stratum label
(ordered ascending by label), consumed by
fast_coxph_regression_prebuilt_cpp exactly like the length-1
vector returned by build_cox_data_cache_cpp. Lifetime,
garbage-collection, and non-serializability semantics are identical to
build_cox_data_cache_cpp: the cache is immutable once
built and safe to reuse across repeated fits as long as X,
y, dead, and strata have not changed; callers (e.g.
InferenceStratifiedCoxPH$private$strat_cox_data_cache) are
responsible for invalidating and rebuilding it when the treatment
assignment, covariates, or stratification changes.
Complexity. O(n \log(n/S)) for the per-stratum sorts plus
O(n) for the tabulation passes, where n is the number of
subjects and S is the number of distinct strata; memory use is
O(np) for the per-stratum row-major copies of X plus
O(n) for the sorted y/dead and event-time tables.
Value
An externalptr to a cached, per-stratum sorted Cox
risk-set representation of (X, y, dead, strata), for use as the
cox_data_xptr argument of
fast_coxph_regression_prebuilt_cpp.
See Also
build_cox_data_cache_cpp for the unstratified
(single risk-set) analog, fast_coxph_regression_prebuilt_cpp
for the fitting routine that consumes this cache, and
fast_coxph_regression for the full Cox model documentation,
including the partial-likelihood derivation, tie-handling, and references.
Analogous Python API:
lifelines CoxPHFitter
(stratified fits via the strata argument) and
statsmodels duration models.
Check Whether a Suggested Package Is Installed (Memoized)
Description
Tests whether package_name is installed via
requireNamespace and memoizes the result in a
package-level environment (package_cache), so that repeated checks
for the same package within an R session pay the namespace-lookup cost
only once. EDI uses this to guard optional code paths that depend on
Suggests-only packages (e.g. quantreg, betareg,
nbpMatching, geepack, icenReg) that are not installed
automatically with the package, issuing an informative
stop()/warning() and falling back to an internal
implementation when the dependency is absent, rather than failing with an
opaque "could not find function" error.
Usage
check_package_installed(package_name)
Arguments
package_name |
Character scalar. The name of the package to check
(as passed to |
Details
Caching / mutation semantics. This function has a side effect:
on the first call for a given package_name in the current R
session, it assigns the boolean result of requireNamespace() into
the package-global environment package_cache, keyed by
package_name. All subsequent calls for the same
package_name (from anywhere in the package, or from user code)
read the cached value directly and do not re-query the namespace
registry. This means the result reflects whether the package was
installed at the time of the first call; installing or removing
package_name later in the same R session will not be picked up.
The cache is a plain environment (not an R6 object) shared by all
callers within the process and is not reset between calls; it is,
however, re-initialized fresh in each new R session.
Determinism. For a fixed installed-package state,
check_package_installed() is deterministic and side-effect-free
beyond the memoization described above; it does not consume random
number generator state and involves no numerical computation.
Lifecycle. Internal utility (exported for reuse within the
package's own R6 classes across files, not intended as a general-purpose
user-facing API); prefer requireNamespace() directly for
one-off checks in user code.
Value
Logical scalar. TRUE if the package is installed and its
namespace can be loaded, FALSE otherwise.
See Also
requireNamespace, on which this function is
a thin memoizing wrapper.
Delete this machine's saved EDI tuning and return to shipped defaults
Description
Removes the per-user config file written by
tune_EDI_for_this_machine (so the next library(EDI)
starts from the package's built-in performance-policy defaults) and
resets the in-session cold-start, warm-start, and optimizer-algorithm
dispatch policies to those defaults right away.
Usage
clear_local_EDI_optimization()
Value
Invisibly, TRUE if a saved tuning existed and was removed,
FALSE if there was none.
See Also
tune_EDI_for_this_machine,
get_local_EDI_optimization.
Examples
clear_local_EDI_optimization()
Conditional logistic regression for matched pairs
Description
Internal method. Replaces bclogit::clogit. For matched pairs (exactly 2 subjects per stratum), the conditional log-likelihood depends only on discordant pairs. Within each discordant pair the contribution reduces to ordinary logistic regression on signed within-pair differences with no intercept.
Usage
clogit_helper(y_m, X_m, w_m, strata_m)
Arguments
y_m |
Binary outcome vector (0/1) for matched subjects. |
X_m |
Covariate matrix or data.frame (may have 0 columns). |
w_m |
Treatment indicator (0/1) for matched subjects. |
strata_m |
Integer stratum IDs (pair labels) for matched subjects. |
Value
Result list from fast_logistic_regression_with_var: b[1] = beta_T, ssq_b_j = Var(beta_T). NULL on failure.
Fast Bai Adjusted T Statistic for Multiple Permutations
Description
Fast Bai Adjusted T Statistic for Multiple Permutations
Usage
compute_bai_distr_parallel_cpp(
w_mat,
m_mat,
y,
delta,
halves_idx,
convex_flag,
num_cores
)
Arguments
w_mat |
Integer matrix of permuted treatment assignments (n x r). |
m_mat |
Integer matrix of match indicators (n x r). |
y |
Numeric response vector. |
delta |
Null treatment effect shift. |
halves_idx |
Integer matrix of half-sample indices. |
convex_flag |
Logical flag for convex combination. |
num_cores |
Number of OpenMP threads. |
Value
Numeric vector of Bai adjusted T statistics.
Randomization/Bootstrap Reference Distribution of the Treatment Log-Hazard-Ratio for a Treatment-Only Cox PH Model (C++ Backend, Single-Covariate)
Description
Builds an empirical reference (null or shifted-null) distribution of the treatment
log-hazard-ratio \hat\beta_T from a treatment-only (single-covariate, no other
adjustment covariates) Cox proportional-hazards model, by refitting the model on
B = ncol(i_mat) pre-generated resample-and-reassign draws. This backs bootstrap
randomization test (BRT) inversion for confidence intervals and p-values on the Cox
log-hazard-ratio: repeated calls at different delta values (or a single call at
delta = 0 for a null/reference distribution) let the caller invert the empirical
distribution of \hat\beta_T against a target quantile or tail probability. This is
a single-covariate special case; compute_coxph_rand_bootstrap_parallel_cpp (used
by fast_coxph_regression's survival inference class) generalizes this to
models with additional adjustment covariates and to Gaussian smoothing noise on the
resampled log-hazard-ratio, and is the version actually wired into
InferenceCoxPH; this treatment-only function currently has no in-package
caller and should be treated as a lighter-weight standalone utility or superseded
building block rather than part of the primary inference path.
Usage
compute_coxph_rand_bootstrap_cpp(y0, dead, i_mat, w_mat, delta, num_cores)
Arguments
y0 |
Numeric vector of original survival times (event or censoring time), length
|
dead |
Numeric vector of length |
i_mat |
Integer matrix ( |
w_mat |
Integer matrix ( |
delta |
Sharp-null log-time shift applied multiplicatively ( |
num_cores |
Number of OpenMP threads to use for parallelizing across draws (ignored, and draws run sequentially, when the package is built without OpenMP support). |
Details
Per-draw model. For each draw b = 1, \dots, B, this function forms a
resampled dataset of size n = length(y0) by taking row indices
i_mat[, b] (1-based, into the original y0/dead) and treatment
labels w_mat[, b], applies the sharp-null time shift (see below), and fits an
unstratified, single-covariate (treatment-only) Cox partial-likelihood model — the
p = 1 case of the model documented in build_cox_data_cache_cpp and
fast_coxph_regression — via Newton-Raphson (cox_fit() with
estimate_only = true, maxit = 20, tol = 1e-9, no warm start,
smart_cold_start = false). Only the fitted coefficient \hat\beta_T (the log
hazard ratio for treatment) is returned per draw; no variance-covariance matrix is
computed (this function is for building a resampling distribution, not for single-fit
inference).
Sharp-null shift. delta encodes a sharp null hypothesis of a constant
multiplicative shift on the time scale for treated subjects: for draw b, subject
i with resampled treatment label w_i \in \{0, 1\}, the working survival time
is y_i \cdot e^{\delta} if w_i = 1 and y_i (unchanged) if w_i = 0;
dead status is carried over from the original row unchanged. This is an
accelerated-failure-time-style sharp null (\delta = 0 recovers the unshifted
resample), not a proportional-hazards sharp null; delta is on the same log-time
scale used elsewhere in the package's AFT/Weibull machinery, not the log-hazard-ratio
scale of the returned \hat\beta_T.
Resampling scheme. i_mat and w_mat are assumed pre-generated by
the caller (e.g. via the package's bootstrap or randomization-draw machinery) and are
not validated here: i_mat need not be a permutation (indices may repeat, as in a
nonparametric bootstrap draw with replacement, or may be a permutation, as in a
randomization test) and w_mat need not respect any particular design's
assignment-probability structure — whatever exchangeability or randomization-validity
properties the resulting reference distribution has are entirely a property of how the
caller generated i_mat/w_mat, not of this function.
Non-convergence and missingness. A draw's entry in the output is NA if
the Newton-Raphson fit fails to converge within maxit iterations or the fitted
coefficient is non-finite; callers must handle NA entries (e.g. by omission) when
computing empirical quantiles or tail probabilities from the returned vector.
Parallelism and reproducibility. When compiled with OpenMP support and
num_cores > 1, draws are processed in parallel via
#pragma omp parallel for schedule(dynamic); each draw is a self-contained fit with
no shared mutable state across draws (aside from writing to disjoint output slots), so
results are deterministic given i_mat/w_mat regardless of the number of
threads or scheduling order — this function consumes no RNG state itself, since the
randomness lives entirely in how the caller generated i_mat and w_mat.
Complexity. O(B \cdot n \log n) for the B independent single-
covariate Cox fits (dominated by the per-draw CoxData sort), parallelized across
num_cores threads when available; memory use is O(n) per in-flight draw.
Value
Numeric vector of length B = ncol(i_mat) with the fitted treatment
log-hazard-ratio \hat\beta_T for each draw, or NA_real_ for draws whose
Cox fit failed to converge.
See Also
fast_coxph_regression for the underlying Cox partial-likelihood
model and its references; build_cox_data_cache_cpp for the per-draw
sorted risk-set construction each draw performs internally.
See also randomization test
and bootstrap for
background on resampling-based reference distributions.
Fast KK Wilcoxon Statistic for Multiple Permutations
Description
Fast KK Wilcoxon Statistic for Multiple Permutations
Usage
compute_matching_wilcox_distr_parallel_cpp(
w_mat,
m_mat,
y,
delta,
transform_code,
zero_one_logit_clamp,
is_fixed_matching,
num_cores
)
Arguments
w_mat |
Integer matrix of permuted treatment assignments (n x r). |
m_mat |
Integer matrix of match indicators (n x r). |
y |
Numeric response vector. |
delta |
Null treatment effect shift. |
transform_code |
Integer code for response transformation. |
zero_one_logit_clamp |
Clamp value for logit transformation. |
is_fixed_matching |
Logical flag for fixed matching designs. |
num_cores |
Number of OpenMP threads. |
Value
Numeric vector of KK Wilcoxon statistics.
Parallel Stereotype Logit Randomization Distribution
Description
Parallel Stereotype Logit Randomization Distribution
Usage
compute_stereotype_logit_distr_parallel_cpp(X, y, w_mat, delta, num_cores)
Arguments
X |
Matrix of covariates (without intercept or treatment). |
y |
Numeric vector of response values (pre-null-shifted for treated). |
w_mat |
Integer matrix of permuted treatment assignments (n x nsim). |
delta |
Null treatment effect (additive shift). |
num_cores |
Number of OpenMP threads. |
Value
Numeric vector of length nsim with treatment coefficients.
Parallel BRT kernel for KM-diff (median) and RMST-diff. Each replicate resamples rows i_mat(.,b) and pairs them with assignment w_mat(.,b). Sharp-null shift is multiplicative on treated times (exp(delta)). Uses an inline pure-C++ KM calculator — no R objects inside the loop, so OpenMP is safe.
Description
Parallel BRT kernel for KM-diff (median) and RMST-diff. Each replicate resamples rows i_mat(.,b) and pairs them with assignment w_mat(.,b). Sharp-null shift is multiplicative on treated times (exp(delta)). Uses an inline pure-C++ KM calculator — no R objects inside the loop, so OpenMP is safe.
Usage
compute_survival_stat_diff_rand_bootstrap_parallel_cpp(
y0,
dead,
i_mat,
w_mat,
delta,
do_rmst,
noise_mat,
num_cores
)
Arguments
do_rmst |
TRUE for RMST-diff, FALSE for median (KM-diff). |
Compute automatic survival strata IDs from low-cardinality covariates
Description
Selects numeric covariate columns with a small number of observed levels and combines them into a single all-subject stratum identifier. This is used to support automatic stratified Cox models when the design object stores observed covariates but no explicit stratum variable.
Usage
compute_survival_strata_ids_cpp(
X,
max_unique_per_col = 4L,
max_strata_cols = 4L,
min_count_per_level = 2L
)
Arguments
X |
Numeric covariate matrix. |
max_unique_per_col |
Maximum number of unique values allowed for a column to be considered a stratification candidate. |
max_strata_cols |
Maximum number of candidate columns to combine. |
min_count_per_level |
Minimum frequency required for every level in a candidate column. |
Value
A list with 'strata_id', 'selected_cols', and 'num_strata'.
Fast Wilcoxon HL Statistic for Multiple Permutations
Description
Fast Wilcoxon HL Statistic for Multiple Permutations
Usage
compute_wilcox_hl_distr_parallel_cpp(
w_mat,
y,
delta,
transform_code,
zero_one_logit_clamp,
num_cores
)
Arguments
w_mat |
Integer matrix of permuted treatment assignments (n x r). |
y |
Numeric response vector. |
delta |
Null treatment effect shift. |
transform_code |
Integer code for response transformation. |
zero_one_logit_clamp |
Clamp value for logit transformation. |
num_cores |
Number of OpenMP threads. |
Value
Numeric vector of HL statistics.
Build an Intercept-Free, Full-Rank Covariate Design Matrix from a Formula
Description
Expands formula against data (via model.matrix) into
a purely numeric covariate design matrix suitable for the package's own fast_*
GLM/survival/ordinal fitting routines, which manage their own intercept and treatment
columns separately rather than relying on the formula/model-matrix machinery for them.
This is the standard covariate-matrix builder used throughout EDI's inference classes
(e.g. Inference$private$X) whenever adjustment covariates need to go from a
user-facing formula/data-frame representation to a numeric matrix the C++ backends can
consume.
Usage
create_model_matrix_from_features(formula, data)
Arguments
formula |
A formula object giving the covariate specification to expand (e.g.
|
data |
A data frame or data table supplying the variables referenced in
|
Details
What it does. (1) If data has zero columns, returns a numeric
nrow(data) x 0 matrix immediately (no covariates to expand). (2) Otherwise calls
model.matrix(formula, data = data), which performs standard
formula expansion: factor variables are dummy-coded against their reference level (the
first level of levels, or the level ordering already present in
data), interactions (a:b, a*b) are expanded to product columns, and
any model.matrix contrasts option in effect at call time applies. (3) If the
first resulting column is named "(Intercept)" (i.e. the formula was not given
- 1 / + 0), that column is dropped — this function always returns a
covariate-only matrix with no intercept column, since EDI's design and inference classes
add their own intercept/treatment columns at a fixed position. (4) The result is passed
through drop_linearly_dependent_cols, which detects the numeric rank of
the matrix (via matrix_rank_cpp() at tolerance 1e-7) and, if the matrix is
rank-deficient, greedily retains a full-rank subset of columns using the pivot order
from qr(M, tol = 1e-7) (dropping the same tolerance's worth of
redundant/aliased columns, e.g. from collinear dummy expansions or an over-specified
interaction structure); this rank-reduction step is silent — no warning is issued when
columns are dropped, and the dropped columns' identity/names are not returned to the
caller, only the reduced matrix.
Input conventions. data is expected to already be free of missing
values at call time (imputation, when configured, happens upstream in the design/
inference class before this function is called); this function does not impute or warn
about NAs, and model.matrix will drop incomplete rows or error, per its
own na.action default, if NAs remain. Column order and names in the
returned matrix follow model.matrix's expansion order (all factor/interaction
columns for a term before the next term), possibly reduced by the rank-deficiency step;
callers relying on stable column identity (e.g. warm-starting coefficients across calls)
should not assume the set or order of columns is invariant if data's factor
levels or rank change between calls.
Failure semantics. If drop_linearly_dependent_cols detects the matrix is
non-numeric or contains non-finite values, it returns the matrix unchanged (rank
reduction is skipped rather than erroring); a downstream fitting routine operating on a
rank-deficient or non-finite design matrix may then fail to converge or report
non-finite coefficients/standard errors, which is where such problems will actually
surface to the user.
Value
A numeric matrix with nrow(data) rows and one column per retained,
full-rank expanded covariate term (no intercept column). Has zero columns if
data has zero columns.
See Also
model.matrix, which performs the formula expansion this
function wraps; drop_linearly_dependent_cols (internal, same file) for the
rank-deficiency cleanup step. Analogous Python API:
patsy/
statsmodels formula API
for formula-based design matrix construction.
Return EDI Build Information (C++ Backend)
Description
Returns the compiler and package build metadata that was baked into the currently
loaded EDI shared object at compile time (via preprocessor macros defined
in edi_build_flags.h, generated by the package's build tooling), not
anything queried live from the running system. This is intended for benchmark
reports and reproducibility audits where the exact compiler, flags, and build
environment used to produce the installed binary matter (e.g. explaining
performance differences between two installations of the same package version, or
confirming a binary was built with a particular optimization/vectorization
configuration).
Usage
edi_build_info_cpp()
Details
Every field is a fixed string or boolean baked in when the C++ source was compiled (via preprocessor macro substitution); calling this function multiple times within the same R session always returns identical values, and the values reflect the build environment, not the environment the function happens to be called from. If EDI is reinstalled/recompiled, existing R sessions that already loaded the old shared object continue to report the old build's metadata until they restart and load the new one.
Value
A named list with the following fields, all length-1 character strings unless noted otherwise:
- capture_method
How the build metadata below was captured by the package's build tooling (
EDI_BUILD_CAPTURE_METHOD).- build_timestamp
Timestamp of compilation (
EDI_BUILD_TIMESTAMP).- build_host
Hostname of the machine that compiled this binary (
EDI_BUILD_HOST).- r_home
R_HOMEof the R installation used to build the package.- r_version
R version string used to build the package.
- r_cxx20
The C++20 compiler command R's build configuration reported.
- r_cxx20std
The C++ standard flag (e.g.
-std=gnu++20) used.- r_cxx20flags
Additional C++20 compiler flags from R's configuration.
- r_shlib_openmp_cxxflags
OpenMP C++ flags R's configuration supplies for shared-library builds (empty if OpenMP support was not available/enabled).
- env_edi_portable, env_edi_disable_vectorization, env_edi_native_speed, env_edi_native_lto
The values (or default/unset indicator) of the corresponding
EDI_*environment variables that were set at build time to control portable-vs-native code generation, vectorization, and link-time optimization; see the package's build documentation for what each controls.- pkg_cppflags, pkg_cxxflags, pkg_libs
The package-level
PKG_CPPFLAGS/PKG_CXXFLAGS/PKG_LIBSused when compiling/linking EDI's C++ sources.- compiler
The compiler identification string (
__VERSION__).- compiler_optimize_macro
Logical;
TRUEiff the compiler defined__OPTIMIZE__(i.e. optimizations were enabled) when this translation unit was compiled.- compiler_fast_math_macro
Logical;
TRUEiff__FAST_MATH__was defined (i.e. non-IEEE-compliant fast-math optimizations were enabled), which can affect floating-point reproducibility/edge-case behavior (NaN/Inf handling, exact rounding) relative to a standard-compliant build.- eigen_dont_vectorize_macro
Logical;
TRUEiffEIGEN_DONT_VECTORIZEwas defined, disabling Eigen's SIMD vectorization for this build (e.g. for portability to CPUs lacking the vector instructions a native build would target).
Examples
info = edi_build_info_cpp()
info$pkg_cxxflags
Re-bind already-installed lazy-component methods on a freshly cloned
Inference object to that clone's own self/private.
Description
install_lazy_inference_component() permanently binds each real
(non-stub) implementation it installs to whichever object triggered the
install, via environment(value) = parent.frame(). R6's
clone() correctly rebinds every method present in the class
generator's original method list, but a lazily-installed method is
injected into private/self at runtime and is invisible to
that bookkeeping, so a clone keeps calling back into the ORIGINAL
object's data (e.g. a Bayesian-bootstrap worker clone silently reading
the pre-clone object's current_bayesian_bootstrap_context, always
NULL, instead of its own). Call this right after self$clone()
to repoint every already-installed lazy-component method (public and
private) at the clone's own enclosing environment; state fields
(owns_state) are left untouched since clone() already
copies their current values correctly.
Usage
edi_rebind_lazy_components_after_clone(i, source_private = NULL)
Arguments
i |
The freshly cloned |
source_private |
The pre-clone source object's own |
Exact Two-Group Jonckheere-Terpstra Test via Full Randomization Enumeration (C++ Backend)
Description
Computes the exact randomization-distribution p-value and a probabilistic-index
effect size for the two-group Jonckheere-Terpstra statistic — which, with exactly
two groups (w in {0, 1}), coincides with the Wilcoxon-Mann-Whitney
U statistic generalized to handle ties (repeated ordinal levels in y):
U = \sum_{k} t_k \big(2 L_k + (n_k - t_k)\big) / 2,
summed over the K distinct observed levels of y (in increasing order),
where n_k is the total count at level k, t_k is the observed count
of w == 1 subjects at level k, and L_k is the number of subjects
at strictly lower levels (this is the standard "number of favorable comparisons"
Mann-Whitney statistic, adapted for tied/grouped ordinal data — a tie at the same
level contributes 1/2 rather than 0 or 1). The exact (not asymptotic,
not Monte Carlo) null/reference distribution of this statistic under the sharp null
of no treatment effect is obtained by enumerating, via a dynamic-programming
recursion over levels (recurse_jt_distribution()), every way to distribute
n_treat "treated" labels among the n subjects consistent with the fixed
per-level totals n_k — i.e. the exact multivariate hypergeometric
randomization distribution of the statistic conditional on the observed level
counts, with each configuration's probability computed in log-space from
log-binomial-coefficient weights to avoid overflow for larger n.
Usage
exact_jonckheere_terpstra_pval_cpp(y, w)
Arguments
y |
Integer (or integer-coercible) vector of length |
w |
Integer (or integer-coercible) vector of length |
Details
Randomization test, not a model-based test. This is a randomization
(permutation) test in the Fisherian sense: it conditions on the observed marginal
level counts n_k and the group sizes n_treat/n_control, and asks
how extreme the observed statistic is relative to every other way those same labels
could have been randomly assigned, so its validity does not depend on any
distributional assumption about y beyond exchangeability under the null. The
two-sided p-value is p_{\mathrm{exact}} = \min(1, 2 \min(p_{\mathrm{lower}},
p_{\mathrm{upper}})), where p_{\mathrm{lower}}/p_{\mathrm{upper}} are
the exact one-sided tail probabilities of the randomization distribution at or below
/ at or above the observed statistic.
Effect size (superiority). superiority is the probabilistic
index \Pr(Y_T > Y_C) + \tfrac{1}{2}\Pr(Y_T = Y_C) (equivalently U
rescaled to [0, 1] by dividing by n_{\mathrm{treat}} \cdot
n_{\mathrm{control}}), the probability a randomly chosen treated subject's ordinal
outcome exceeds a randomly chosen control subject's, counting ties as half a win;
0.5 indicates no stochastic ordering between groups, and 1/0
indicate the treated group's outcomes are uniformly higher/lower.
Input conventions. y is coerced to integer and treated as an ordinal
(or any orderable-by-integer-value) response with an arbitrary number of tied
levels; w must be an integer/coercible-to-integer vector of {0, 1}
values with both groups non-empty. NA in either y or w is not
permitted and raises an error, as does a non-{0,1} value in w or an
empty input.
Complexity. The recursion's state space scales with the number of distinct
possible statistic values (O(n_{\mathrm{treat}} \cdot n_{\mathrm{control}})
many), and thread-local buffers are reused (not reallocated) across repeated calls
within the same thread for the same or smaller problem sizes; this exact enumeration
is exponential in the number of distinct levels/group sizes in the worst case
(unlike an asymptotic normal-approximation JT test), so it is intended for small-to-
moderate n where exactness matters more than raw speed.
Value
A list with components stat2 (twice the observed U statistic,
an integer, used internally to keep the enumeration in integer arithmetic),
n_treat, n_control (the two group sizes), superiority (the
probabilistic-index effect size described above), p_lower, p_upper
(the exact one-sided randomization-distribution tail probabilities), and
p_exact (the exact two-sided p-value).
See Also
Mann-Whitney
U test and
Jonckheere's trend
test for background; analogous Python API:
SciPy
mannwhitneyu (method="exact" for the same exact-enumeration
approach, though SciPy's exact path does not handle ties the same way).
Expand Ordinal Data into Stacked Binary Comparisons for Adjacent-Category Logit Regression (C++ Backend)
Description
Reshapes an ordinal response y (levels 1, \dots, K) into the stacked
binary-outcome, per-cut-stratified form required to fit an adjacent-category logit
model as a single conditional (stratified) logistic regression, so the package's
existing binary/conditional-logit fitting backends can be reused unchanged for
ordinal adjacent-category models rather than needing a bespoke ordinal solver.
Usage
expand_adjacent_category_data_cpp(y, w, strata, K)
Arguments
y |
Integer vector of length |
w |
Integer vector of length |
strata |
Integer vector of length |
K |
Integer; the number of ordinal categories (so there are |
Details
Model. The adjacent-category logit model compares each pair of
consecutive categories j and j+1 (j = 1, \dots, K-1) via
\log\frac{\Pr(Y = j+1 \mid Y \in \{j, j+1\})}{\Pr(Y = j \mid Y \in \{j, j+1\})} = \alpha_j + \beta^\top x,
i.e. a logistic model for "category j+1 vs. category j" fit using only
the subjects actually observed in one of those two categories, with a
cut-specific intercept \alpha_j and covariate effects \beta constrained
equal across all K-1 cuts (the proportional/parallel adjacent-category
assumption). This differs from the cumulative-logit (proportional-odds) model,
which instead compares Y \le j vs. Y > j using every subject at
every cut.
Expansion mechanics. For each subject i and each cut
j = 1, \dots, K-1 (n_alpha = K - 1), a stacked row is emitted
only if y[i] equals j or j+1; subjects at any other
level contribute nothing to that cut's comparison (so each subject contributes to
at most 2 of the K-1 cuts: the ones immediately adjacent to their observed
level, and exactly 1 cut if at an extreme level). The stacked binary outcome is
1 if y[i] == j + 1 (upper category) and 0 if
y[i] == j (lower category). The stacked stratum ID is
strata[i] + (j - 1) * num_strata (where num_strata = max(strata)),
i.e. the original stratum crossed with the cut index j: fitting a
conditional logistic regression stratified on this combined ID and pooling all
stacked rows together estimates a single shared treatment coefficient
\beta across all cuts, while allowing each (original stratum, cut)
combination to absorb its own nuisance intercept via strata conditioning (the same
stratified-conditional-logit trick used elsewhere in the package, e.g. for
continuation-ratio models via expand_continuation_ratio_data_cpp()).
Input conventions. y must take integer values in
1:K (1-based category labels); w is typically the treatment
indicator/covariate to estimate a coefficient for, passed through unchanged per
stacked row (not itself expanded/transformed); strata must be positive
integers with max(strata) == num_strata (no gaps assumed beyond that
maximum, since combined stratum IDs are computed by simple integer arithmetic on
num_strata, not by re-indexing distinct values). No input validation is
performed at this layer (no range/type checks on y/strata); passing
out-of-range values silently produces incorrect stratum IDs or drops rows rather
than erroring.
Value
A list with components y (stacked 0/1 binary outcome), w
(stacked covariate, passed through unchanged), and strata (stacked
combined stratum-by-cut ID); all three are integer vectors of the same,
generally-longer-than-n length (each subject contributes 0, 1, or 2
stacked rows depending on their observed category).
See Also
expand_continuation_ratio_data_cpp() for the analogous expansion
used by continuation-ratio ordinal models.
Ordinal regression for
orientation; analogous Python API:
statsmodels discrete
models (no direct adjacent-category equivalent; the closest analog is fitting
the expanded data as a conditional/grouped logit).
Expand Ordinal Data into Stacked Binary Comparisons for Continuation-Ratio Regression (C++ Backend)
Description
Reshapes an ordinal response y (levels 1, \dots, K) into the stacked
binary-outcome, per-cut-stratified form required to fit a (forward) continuation-
ratio logit model as a single conditional (stratified) logistic regression —
the discrete-time-hazard analog for ordinal data — so the package's existing
binary/conditional-logit fitting backends can be reused unchanged rather than
needing a bespoke ordinal solver. This is the continuation-ratio counterpart of
expand_adjacent_category_data_cpp(); the two share the same stacking and
combined-stratum trick but differ in which rows each subject contributes (see
Details).
Usage
expand_continuation_ratio_data_cpp(y, w, strata, K)
Arguments
y |
Integer vector of length |
w |
Integer vector of length |
strata |
Integer vector of length |
K |
Integer; the number of ordinal categories (so there are |
Details
Model. The continuation-ratio model treats reaching each successive
category as a sequence of conditional "continue past this cut" events, analogous
to a discrete-time survival/hazard model: for cut j = 1, \dots, K-1, among
subjects who have reached at least category j (Y \ge j),
\log\frac{\Pr(Y > j \mid Y \ge j)}{\Pr(Y = j \mid Y \ge j)} = \alpha_j + \beta^\top x,
i.e. the log-odds of "continuing" past category j versus "stopping" (being
observed) exactly there, given the subject has reached at least j, with
a cut-specific intercept \alpha_j and covariate effects \beta
constrained equal across cuts (the proportional continuation-ratio assumption).
This orientation — numerator is the "continue" event — keeps a positive
\beta meaning "pushes toward higher categories of y", matching
fast_continuation_ratio_regression_cpp and every other ordinal
estimator in the package. Unlike the adjacent-category model (which only
compares the two categories immediately flanking a cut), every subject
contributes to every cut up to and including the one at which they are observed
to stop.
Expansion mechanics. For each subject i with observed category
y[i], a stacked row is emitted for every cut
j = 1, \dots, \min(\code{y[i]}, K-1): the stacked binary outcome is 0
("stopped here") if y[i] == j, and 1 ("continued past") for every
earlier cut the subject passed through. A subject observed at the top category
(y[i] == K) contributes a 1 at every one of the K - 1 cuts
(having "survived" all of them without stopping); a subject observed at category
j <= K - 1 contributes 1s for cuts 1:(j-1) and a single
0 at cut j, then no further rows (later cuts are irrelevant once a
subject has already stopped). As in expand_adjacent_category_data_cpp(),
the stacked stratum ID is strata[i] + (j - 1) * num_strata (with
num_strata = max(strata)): fitting a conditional logistic regression
stratified on this combined ID and pooling all stacked rows estimates a single
shared treatment coefficient \beta across all cuts, while each (original
stratum, cut) combination absorbs its own nuisance intercept via strata
conditioning.
Input conventions. y must take integer values in 1:K;
w is passed through unchanged into each stacked row for that subject
(typically the treatment indicator/covariate to estimate a coefficient for);
strata must be positive integers, with max(strata) used as the
per-cut stratum-ID offset. No input validation is performed at this layer.
Value
A list with components y (stacked 0/1 "continued past this cut"
outcome),
w (stacked covariate, passed through unchanged), and strata
(stacked combined stratum-by-cut ID); all three are integer vectors of the
same, generally-longer-than-n length (each subject contributes between 1
and K - 1 stacked rows, depending on their observed category).
See Also
expand_adjacent_category_data_cpp() for the analogous expansion
used by adjacent-category ordinal models.
Ordinal regression for
orientation; analogous Python API:
statsmodels discrete
models (no direct continuation-ratio equivalent; the closest analog is fitting
the expanded data as a conditional/grouped logit, or discrete-time survival
packages).
Fast Adjacent-Category Logit Regression, Direct MLE (C++ Backend)
Description
Fits the adjacent-category logit ordinal regression model
\log\frac{\Pr(Y = k+1)}{\Pr(Y = k)} = \alpha_k + \beta^\top x, \quad k = 1, \dots, K-1,
by direct maximum likelihood on the full multinomial likelihood of y,
rather than via the stacked-binary / stratified-conditional-logit reduction
implemented by expand_adjacent_category_data_cpp() elsewhere in the
package. \beta (the covariate effects, shared across all K - 1 cuts)
and the K - 1 cut-specific intercepts \alpha_k are estimated jointly
by numerically optimizing the exact multinomial log-likelihood, which is
generally more accurate and can be faster than fitting the row-stacked expansion
as a stratified logistic regression, at the cost of a custom (rather than reused)
optimizer implementation.
Usage
fast_adjacent_category_logit_cpp(
X,
y,
maxit = 100L,
tol = 1e-08,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
warm_start_params = NULL,
warm_start_beta = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of length |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
fixed_idx |
Optional integer indices (into the
|
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see Details. |
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
the full |
warm_start_params |
Optional starting values for the full parameter vector
|
warm_start_beta |
Optional starting values for just the covariate
coefficients |
Details
Category coding. y need not already be coded 1:K: the
distinct values of y are extracted and sorted (get_levels()), and
each observation is remapped to its 1-based rank among those sorted
distinct values (map_y_to_1K()) — e.g. y = c(10, 30, 20, 10) is
treated identically to y = c(1, 3, 2, 1), with K = 3. K is
therefore the number of distinct observed values, not any externally
supplied category count, and requires at least 2 (an error is raised otherwise).
Parameterization and likelihood. Internally, category probabilities are
computed via a numerically stable log-space recurrence: unnormalized
log-probabilities \log \tilde p_k = \sum_{j=k}^{K-2} (\alpha_j - \eta)
(\eta = x^\top \beta) are accumulated additively — never by exponentiating
\alpha_k or -\eta directly — and normalized with a standard
log-sum-exp, so every exp() call sees an argument \le 0 and
\Pr(Y = k) is bounded to [0, 1] regardless of how extreme
\alpha/\eta get during optimization (fixed 2026-08-27: the prior
right-to-left product recurrence in terms of raw e^{-\eta} and
e^{\alpha_k} could each individually overflow to Inf before
normalization, corrupting the objective/gradient to Inf/NaN and
leaving the optimizer's line search unable to recover — confirmed via direct
testing to reliably exhaust the full iteration budget without converging on
ordinary synthetic data at every sample size and seed tried, and to diverge
outright to NaN parameters from an all-zero start). The returned
neg_loglik is the resulting exact multinomial negative log-likelihood
(-\sum_i \log \Pr(Y_i = y_i)), with the analytic gradient computed in the
same pass and used internally for optimization. See
fast_adjacent_category_logit_with_var_cpp for the variant that
additionally returns the variance-covariance matrix of the estimates.
Parameter vector layout. The optimizer's parameter vector (returned as
params) is c(alpha_1, ..., alpha_{K-1}, beta_1, ..., beta_p) — the
K - 1 cut intercepts first, then the p shared covariate
coefficients (p = ncol(X)).
Optimization. Optimized via optimization_alg ("lbfgs"
default; see .normalize_optimizer_algorithm for the supported set),
for at most maxit iterations at tolerance tol. When no warm start is
supplied, smart_cold_start = TRUE (default) seeds the optimizer from an
OLS-based initial guess rather than a naive zero/arbitrary start; supplying
warm_start_params (the full parameter vector) or warm_start_beta
(just the covariate coefficients, with cut intercepts initialized separately)
overrides smart_cold_start entirely. fixed_idx/fixed_values
allow holding specific parameters (by index into the layout above) fixed at
supplied values during optimization rather than estimating them, and
warm_start_fisher_info allows reusing a previously computed Fisher
information matrix to warm-start curvature information for faster convergence.
Value
A list with components b (the shared covariate coefficients
\hat\beta, length p), alpha (the K - 1 estimated
cut intercepts \hat\alpha_k), params (the full
c(alpha, b) parameter vector, as optimized), neg_loglik (the
multinomial negative log-likelihood at convergence), and converged
(logical).
See Also
fast_adjacent_category_logit_with_var_cpp for the
variance-augmented variant; expand_adjacent_category_data_cpp() for the
alternative stacked-binary reduction of the same model.
Ordinal regression for
orientation.
Fast Adjacent-Category Logit with Variance (C++)
Description
Adjacent-category logit model fitting with full variance-covariance matrix.
Usage
fast_adjacent_category_logit_with_var_cpp(
X,
y,
maxit = 100L,
tol = 1e-08,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
warm_start_params = NULL,
warm_start_beta = NULL
)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of responses (categorical). |
maxit |
Maximum number of iterations. Fast Adjacent-Category Logit Regression with Variance, Direct MLE (C++ Backend) Fits the same adjacent-category logit model as
|
tol |
Convergence tolerance. |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
fixed_idx |
Optional integer indices (into the
|
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see Details. |
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
the full |
warm_start_params |
Optional starting values for the full parameter vector
|
warm_start_beta |
Optional starting values for just the covariate
coefficients |
Details
Variance computation. The observed Fisher information (the Hessian of the
negative log-likelihood, via AdjacentCategoryLogitNegLogLik::hessian()) is
evaluated at the fitted parameter vector over all n_alpha + p
parameters (returned in full as fisher_information), then restricted to the
free (non-fixed_idx) parameters and inverted via a rank-aware
(symmetric_pseudo_inverse(), not a plain Cholesky/LDLT solve) inverse
before being expanded back to full (n_alpha + p) x (n_alpha + p) size as
vcov. The pseudo-inverse is used deliberately: adjacent-category fits can
have an estimable treatment effect even when nuisance columns make the full
information matrix rank-deficient, a case where a standard Cholesky/LDLT solve
can report spurious success with an invalid (sometimes negative) variance rather
than failing cleanly. vcov is only populated when converged is
TRUE; otherwise it is NULL.
First-covariate variance shortcut. ssq_b_1 (aliased as
ssq_b_j for interface consistency with the package's other
fast_*_with_var_cpp functions) is the variance of \hat\beta_1, the
coefficient on the first column of X — by the package's usual
convention, the treatment-effect column — extracted directly from the free-parameter
covariance block rather than requiring the caller to index into the full
vcov matrix; it is NA if that coefficient was fixed
(via fixed_idx) or if its estimated variance is non-finite or non-positive.
Value
A list with all the components of
fast_adjacent_category_logit_cpp (b, alpha,
params, neg_loglik, converged), plus ssq_b_1
(equivalently ssq_b_j, the variance of the first covariate's
coefficient), vcov (the full parameter variance-covariance matrix, or
NULL if not converged), and fisher_information (the full observed
information matrix at the fitted parameters, over all parameters regardless of
fixed_idx).
See Also
fast_adjacent_category_logit_cpp for the estimate-only
variant and the full model/parameterization documentation.
Fast Beta Regression (R Wrapper)
Description
Fits the beta regression model of Ferrari and Cribari-Neto (2004) for a
continuous response strictly in (0, 1), with mean linked to the covariates
via the logit link and a single (constant) precision parameter \phi. See
fast_beta_regression_cpp for the full model equation, parameter
layout, and optimizer contract implemented by the C++ backend this function
wraps; this page documents only the R-level fallback chain and response-scale
conventions.
Usage
fast_beta_regression(
X,
y,
start_phi = 10,
optimization_alg = "lbfgs",
warm_start_beta = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, with values strictly between 0 and 1.
See |
start_phi |
A numeric value, the starting value for the precision parameter phi. Defaults to 10. |
optimization_alg |
Optimization algorithm: |
warm_start_beta |
Optional starting values for coefficients |
warm_start_fisher_info |
Optional initial Fisher Information matrix, used to warm-start curvature information for the optimizer. |
Details
Fallback chain. The primary implementation uses the C++ backend
(fast_beta_regression_cpp). If that fails to converge or errors, the
function falls back to betareg, which is listed in Suggests and is not
installed automatically with EDI. If betareg is also unavailable (or
itself fails), a final fallback of OLS on logit(y) is used — this last
resort is always available (no external dependency) but does not respect the beta
distribution's mean-variance relationship or estimate \phi at all, so its
coefficients should be treated as an approximate, non-model-based summary rather
than a true beta-regression fit. Install betareg manually to
enable the intermediate fallback. A warning() is issued whenever a fallback
stage is used, naming which stage and the triggering error, so callers can detect
when the primary fit failed even though a result was still returned.
Value
A list containing the following components:
b |
A numeric vector of the estimated beta regression coefficients |
phi |
The estimated precision parameter |
fisher_information |
The working-weights Fisher information matrix from the
C++ backend (see |
Examples
X = matrix(rnorm(500), 100, 5)
y = runif(100)
fast_beta_regression(X, y)
Fast Beta Regression (C++ Backend)
Description
Fits the beta regression model of Ferrari and Cribari-Neto (2004) for a continuous response strictly between 0 and 1 (proportions, rates, and similar bounded outcomes), via direct maximum likelihood on the reparameterized beta density
f(y_i; \mu_i, \phi) = \frac{\Gamma(\phi)}{\Gamma(\mu_i \phi)\Gamma((1-\mu_i)\phi)} y_i^{\mu_i \phi - 1} (1 - y_i)^{(1-\mu_i)\phi - 1}, \quad 0 < y_i < 1,
with mean E[Y_i] = \mu_i and variance
\mathrm{Var}(Y_i) = \mu_i(1-\mu_i) / (1 + \phi), where \phi > 0 is a
single (constant-across-observations) precision parameter and the mean is linked
to the covariates via the logit link \mathrm{logit}(\mu_i) = x_i^\top \beta
(fixed; no alternative link functions are supported by this backend).
Usage
fast_beta_regression_cpp(
X,
y,
warm_start_beta = NULL,
smart_cold_start = TRUE,
start_phi = 10,
compute_std_errs = FALSE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of responses, strictly in |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
start_phi |
Starting value for the precision parameter |
compute_std_errs |
Deprecated; has no effect on this estimate-only entry
point. Use |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see Details. |
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
|
estimate_only |
Logical; if |
Details
Parameterization and optimization. The optimizer's parameter vector is
c(beta, log(phi)) — \phi is optimized on the log scale to keep it
unconstrained (\phi > 0 enforced automatically by exponentiating back),
initialized from start_phi (or a value derived from it via
smart_cold_start). \mu_i is clipped to [10^{-8}, 1 - 10^{-8}]
internally during likelihood/gradient/Hessian evaluation to avoid boundary blowup
when x_i^\top \beta is extreme; this affects only numerical evaluation, not
the returned \hat\beta itself. Optimized via optimization_alg
("lbfgs" default; see .normalize_optimizer_algorithm); when
no warm_start_beta is supplied, smart_cold_start = TRUE (default)
seeds \beta from an OLS-based initial guess.
fixed_idx/fixed_values allow holding specific parameters (by index
into the c(beta, log(phi)) layout) fixed rather than estimated.
compute_std_errs is a legacy/deprecated argument with no effect in this
estimate-only entry point; use fast_beta_regression_with_var_cpp to
obtain standard errors.
Reported likelihood and information. neg_loglik is the exact beta
negative log-likelihood re-evaluated at the fitted parameters (not merely the
optimizer's internal objective trace); fisher_information is the
X^\top W X-style working-weights curvature matrix from the fit
(fit.XtWX) — the same expected-information approximation classical IRLS
uses for GLM standard errors, and exactly what
fast_beta_regression_with_var_cpp inverts to produce vcov —
rather than a fresh evaluation of the exact observed-information Hessian from
get_beta_regression_hessian_cpp. It is also suitable for
warm-starting a subsequent fit via warm_start_fisher_info.
Value
A list containing the following components:
- coefficients
A numeric vector of the estimated beta regression coefficients.
- phi
The estimated precision parameter phi.
- neg_ll
The negative log-likelihood at the final iteration.
- converged
A logical value indicating whether the algorithm converged.
A list with components coefficients (\hat\beta, length
p), phi (\hat\phi, on its natural scale), neg_loglik
(the exact beta negative log-likelihood at the fitted parameters),
converged (logical), and fisher_information (an approximate
working curvature matrix; see Details).
References
Ferrari, S., and Cribari-Neto, F. (2004). "Beta regression for
modelling rates and proportions." Journal of Applied Statistics, 31(7),
799-815, doi:10.1080/0266476042000214501. Analogous Python API:
statsmodels GLM (via the
Beta family, statsmodels.othermod.betareg).
See Also
fast_beta_regression_weighted_cpp for the row-weighted
variant; fast_beta_regression_with_var_cpp for the
variance-augmented variant; fast_beta_regression for the R-level
wrapper with betareg fallback; get_beta_regression_score_cpp/
get_beta_regression_hessian_cpp for standalone score/Hessian
evaluation at arbitrary parameter values.
Examples
X = matrix(rnorm(100), 10, 10)
y = runif(10)
fast_beta_regression_cpp(X, y)
Fast Weighted Beta Regression, Estimate Only (C++ Backend)
Description
Fits the same beta regression model as
fast_beta_regression_cpp (see that page for the full model,
parameterization, and optimizer contract), with each observation's contribution
to the log-likelihood, score, and Hessian multiplied by a nonnegative row weight
weights[i]. Setting all weights to 1 recovers
fast_beta_regression_cpp exactly; this is the backend the package's
Inference classes use whenever the beta regression must be fit on
bootstrap-reweighted or otherwise weighted data (e.g. Bayesian bootstrap weights)
without physically resampling rows.
Usage
fast_beta_regression_weighted_cpp(
X,
y,
weights,
warm_start_beta = NULL,
smart_cold_start = TRUE,
start_phi = 10,
compute_std_errs = FALSE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of responses, strictly in |
weights |
A nonnegative, finite numeric vector of length |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when no warm start is provided. |
start_phi |
Starting value for the precision parameter |
compute_std_errs |
Deprecated; has no effect on this estimate-only entry point. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see
|
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information. |
estimate_only |
If TRUE, skip Fisher information calculation. |
Details
Input validation. weights must have length nrow(X), be
finite and non-negative, and sum to a strictly positive value; violating any of
these raises an error immediately rather than producing a degenerate fit. A
weight of 0 for a given row contributes nothing to the likelihood (effectively
excludes that row) without changing n in downstream index bookkeeping.
Value
A list with the same components as
fast_beta_regression_cpp: coefficients, phi,
neg_loglik (the weighted negative log-likelihood), converged, and
fisher_information.
See Also
fast_beta_regression_cpp for the unweighted model and full
parameterization documentation; fast_beta_regression_with_var_cpp
for the (unweighted) variance-augmented variant.
Fast Beta Regression with Variance Calculation (R Wrapper)
Description
Fits the same beta regression model as fast_beta_regression (see
fast_beta_regression_cpp for the full model equation and
parameterization) and additionally reports the estimated variance of a
caller-selected coefficient, extracted from the fitted parameter
variance-covariance matrix (see
fast_beta_regression_with_var_cpp for how that matrix is computed
and its plain-inverse numerical caveat on rank-deficient designs).
Usage
fast_beta_regression_with_var(
X,
y,
start_phi = 10,
j = 2,
optimization_alg = "lbfgs",
warm_start_beta = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, with values strictly between 0 and 1.
See |
start_phi |
A numeric value, the starting value for the precision parameter phi. Defaults to 10. |
j |
The 1-based index (into |
optimization_alg |
Optimization algorithm: |
warm_start_beta |
Optional starting values for coefficients |
warm_start_fisher_info |
Optional initial Fisher Information matrix, used to warm-start curvature information for the optimizer. |
Details
The primary implementation uses a C++ backend. If that fails, the function falls back
to betareg, which is listed in Suggests and is not installed automatically
with EDI. If betareg is also unavailable, a final fallback of OLS on
logit(y) is used. Install betareg manually to
enable the intermediate fallback.
Value
A list containing the following components:
b |
A numeric vector of the estimated beta regression coefficients |
ssq_b_j |
The estimated variance (squared standard error) of the |
ssq_b_2 |
The estimated variance of the second coefficient specifically
( |
Examples
X = matrix(rnorm(100), 10, 10)
y = runif(10)
fast_beta_regression_with_var(X, y)
Fast Beta Regression with Variance Calculation (C++ Backend)
Description
Fits the same beta regression model as fast_beta_regression_cpp
(see that page for the full model, parameterization, and optimizer contract) and
additionally computes the variance-covariance matrix and standard errors of the
fitted parameters, via the same working-weights (X^\top W X) curvature
matrix documented there.
Usage
fast_beta_regression_with_var_cpp(
X,
y,
warm_start_beta = NULL,
smart_cold_start = TRUE,
start_phi = 10,
compute_std_errs = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of responses, strictly in |
warm_start_beta |
Optional starting values for coefficients. If provided, |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
start_phi |
Starting value for the precision parameter |
compute_std_errs |
Deprecated; standard errors are always computed by this entry point regardless of this argument's value. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see
|
Details
Variance computation. The fit's working-information matrix
(fit.XtWX, over all p + 1 parameters c(beta, log(phi))) is
restricted to the free (non-fixed_idx) parameters and inverted via a
plain matrix inverse (.inverse(), not a rank-aware pseudo-inverse
as used by, e.g., fast_adjacent_category_logit_with_var_cpp) before
being expanded back to the full (p + 1) x (p + 1) size as vcov;
std_errs is sqrt(diag(vcov)). Because this uses a plain inverse, a
rank-deficient or near-singular X (after restricting to free parameters)
will produce numerically unstable or NaN standard errors rather than a
graceful fallback — callers should ensure X is full rank on the free
parameters (e.g. via the package's shared drop_linearly_dependent_cols()
preprocessing) before calling this function if that is not already guaranteed.
Value
A list containing the following components:
- coefficients
A numeric vector of the obtained Poisson regression coefficients.
- phi
The estimated precision parameter phi.
- vcov
The variance-covariance matrix.
- neg_ll
The negative log-likelihood at the final iteration.
- converged
A logical value indicating whether the algorithm converged.
A list with components coefficients (\hat\beta),
phi (\hat\phi), neg_loglik, vcov (the full
(p + 1) x (p + 1) parameter variance-covariance matrix), std_errs
(sqrt(diag(vcov))), converged (logical), and
fisher_information (the working-weights curvature matrix vcov was
inverted from).
See Also
fast_beta_regression_cpp for the estimate-only variant and
the full model/parameterization documentation;
fast_beta_regression_weighted_cpp for the row-weighted
estimate-only variant.
Examples
X = matrix(rnorm(100), 10, 10)
y = runif(10)
fast_beta_regression_with_var_cpp(X, y)
Fast Continuation-Ratio Regression, Direct MLE via Row Augmentation (C++ Backend)
Description
Fits the (forward) continuation-ratio logit ordinal regression model: for cut
j = 1, \dots, K-1, \log \Pr(Y > j \mid Y \ge j) / \Pr(Y = j \mid Y \ge j)
= \alpha_j + \beta^\top x, i.e. the log-odds of continuing past cut
j (rather than stopping there), among subjects who have reached it. This
orientation — numerator is the higher-category event — keeps a positive
\beta meaning "pushes toward higher categories of y", consistent
with every other ordinal estimator in the package (contrast the cumulative-logit
-x^T \beta convention in fast_ordinal_regression.cpp and the
adjacent-category model's \Pr(Y = j+1 \mid \cdot) numerator), and matches
expand_continuation_ratio_data_cpp() (a separate, standalone
row-expansion utility not used by this backend, but documenting the same
"continue past this cut" = 1 orientation). This backend fits the model as a
single unconditional logistic regression MLE
on an internally-built augmented design: build_continuation_ratio_augmented_data()
constructs an augmented matrix X_aug with one dummy column per cut
(n_alpha = K - 1 columns) followed by the original p covariate
columns, and an augmented binary response z (1 = "continued past this
cut", 0 = "stopped here"). Because the cut effects \alpha_j are simply
K - 1 ordinary coefficients on dummy columns (not nuisance parameters
requiring conditioning), an unconditional logistic fit on the augmented data is
exactly equivalent to the continuation-ratio likelihood — no
stratification/conditioning machinery is needed for this standalone use case.
Usage
fast_continuation_ratio_regression_cpp(
X,
y,
maxit = 100L,
tol = 1e-08,
warm_start_beta = NULL,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of length |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
warm_start_beta |
Optional starting values for the full
|
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when no warm start is provided. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see Details. |
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
the full |
Details
Category coding. As in expand_continuation_ratio_data_cpp(),
distinct values of y are extracted and sorted; K is the number of
distinct observed values (not an externally supplied count), and each
observation contributes min(observed_level + 1, K - 1) augmented rows.
Parameter vector layout. The optimizer's parameter vector (returned as
params/beta_full) is c(alpha_1, ..., alpha_{K-1}, beta_1,
..., beta_p). Optimized via optimization_alg ("lbfgs" default),
for at most maxit iterations at tolerance tol; when no warm start
is supplied, smart_cold_start = TRUE seeds the optimizer via OLS on the
augmented binary response z. fixed_idx/fixed_values hold
specific parameters (by index into this layout) fixed rather than estimated, and
warm_start_fisher_info warm-starts curvature information.
Degenerate case. If y has fewer than 2 distinct observed values
(K < 2), no model can be fit: the function returns early with b
zeroed (length p) and an empty alpha, without attempting
optimization or setting converged/neg_loglik/etc.
Value
A list with components b (the shared covariate coefficients
\hat\beta, length p), alpha (the K - 1 estimated
cut intercepts), params/beta_full (the full
c(alpha, b) parameter vector, identical to each other),
neg_loglik (the augmented-data logistic negative log-likelihood, which
equals the continuation-ratio model's negative log-likelihood), X_aug/
z (the augmented design matrix and binary response actually fit, exposed
for reuse, e.g. by
get_continuation_ratio_regression_hessian_cpp()), converged
(logical), and fisher_information (the exact observed information
Hessian at the fitted parameters). See Details for the degenerate
fewer-than-2-categories case, which returns a reduced subset of these fields.
See Also
expand_continuation_ratio_data_cpp() for the full continuation-
ratio model equation and the shared row-augmentation logic;
fast_continuation_ratio_regression_with_var_cpp for the
variance-augmented variant; fast_adjacent_category_logit_cpp for
the analogous direct-MLE fit of the adjacent-category (rather than
continuation-ratio) ordinal model.
Ordinal regression for
orientation.
Fast Weighted Continuation-Ratio Regression, Direct MLE (C++ Backend)
Description
Fits the same continuation-ratio likelihood as
fast_continuation_ratio_regression_cpp(), weighting every augmented
binary row for subject i by that subject's nonnegative weight
w_i. This entry point is intended for bootstrap and other weighted
refits whose estimates must retain the continuation-ratio coefficient
convention.
Usage
fast_continuation_ratio_regression_weighted_cpp(
X,
y,
weights,
maxit = 100L,
tol = 1e-08,
warm_start_beta = NULL,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of length |
weights |
A finite, nonnegative subject-level weight vector of length
|
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
warm_start_beta |
Optional starting values for the full
|
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when no warm start is provided. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see Details. |
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
the full |
Value
The same result fields as
fast_continuation_ratio_regression_cpp(), plus
weights_aug, the weights copied onto the augmented binary rows.
Export of C++ function fast_continuation_ratio_regression_with_var_cpp
Description
Fits the same continuation-ratio model as
fast_continuation_ratio_regression_cpp (see that page for the full
model, row-augmentation mechanics, category coding, and parameter layout) and
additionally computes the variance of the first covariate coefficient and (when
converged) the full parameter variance-covariance matrix, from the same observed
information Hessian.
Usage
fast_continuation_ratio_regression_with_var_cpp(
X,
y,
maxit = 100L,
tol = 1e-08,
warm_start_beta = NULL,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric vector of length |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
warm_start_beta |
Optional starting values for the full
|
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when no warm start is provided. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm; see Details. |
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
the full |
Details
Variance computation. The observed information (Hessian of the
augmented-data logistic negative log-likelihood, evaluated at the fitted
parameters over all n_alpha + p parameters) is restricted to the free
(non-fixed_idx) parameters. ssq_b_j — the variance of
\hat\beta_1 (the coefficient on the first covariate column of
X, the package's usual treatment-effect position) — is obtained via a
single targeted diagonal-entry inversion (compute_diagonal_inverse_entry()),
not a full matrix inverse, and is NA if that coefficient is fixed via
fixed_idx. The full vcov (over all n_alpha + p parameters,
expanded back from the free-parameter block) is computed only when
converged is TRUE (via covariance_from_information());
otherwise vcov is NULL.
Degenerate case. As in fast_continuation_ratio_regression_cpp,
if y has fewer than 2 distinct observed values, the function returns early
with b = NA_real_, ssq_b_j = NA_real_, and converged = FALSE,
without vcov/params/fisher_information.
Value
A list with components b (the shared covariate coefficients
\hat\beta), ssq_b_j (the variance of the first covariate's
coefficient), neg_loglik, vcov (the full parameter
variance-covariance matrix, or NULL if not converged), converged
(logical), params (the full c(alpha, b) parameter vector — cut
intercepts are recoverable as params[1:n_alpha] but are not returned as
a separate alpha field, unlike
fast_continuation_ratio_regression_cpp), and
fisher_information (the full observed information Hessian). See Details
for the degenerate fewer-than-2-categories case, which returns a reduced subset
of these fields.
See Also
fast_continuation_ratio_regression_cpp for the
estimate-only variant and the full model/row-augmentation documentation.
Fast Cox Proportional Hazards Regression (R Wrapper)
Description
Fits the Cox proportional-hazards partial-likelihood model documented in full at
build_cox_data_cache_cpp (model equation, Breslow tie-handling, and
input conventions). This R-level wrapper dispatches to either the package's own
native C++ implementation (fast_coxph_regression_cpp, the default
and recommended path) or, for cross-checking or when the Rcpp path is
unavailable, an elastic-net-with-zero-penalty Cox fit via glmnet
(use_rcpp = FALSE) — not survival::coxph, despite that
being the more commonly used reference implementation for Cox models in R.
Usage
fast_coxph_regression(
X,
y,
dead,
use_rcpp = TRUE,
estimate_only = FALSE,
optimization_alg = "lbfgs",
warm_start_beta = NULL,
warm_start_fisher_info = NULL,
smart_cold_start = TRUE
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept term
is handled implicitly by the Cox model and should not be included in |
y |
A numeric vector representing the observed time (event time or censoring time). |
dead |
A numeric vector (0 or 1) indicating event status (1 for event, 0 for censored). |
use_rcpp |
Logical. If |
estimate_only |
Logical. If |
optimization_alg |
Optimization algorithm: |
warm_start_beta |
Optional starting values for coefficients. If provided, |
warm_start_fisher_info |
Optional initial Fisher Information matrix. Only
affects the |
smart_cold_start |
Logical. If |
Details
Failure semantics. If the C++ fit (use_rcpp = TRUE) errors or
fails to report converged, this function stops with an error rather than
silently falling back to glmnet — the two code paths are alternative
caller choices, not an automatic fallback chain (contrast with, e.g.,
fast_beta_regression's automatic betareg fallback).
glmnet dependency. When use_rcpp = FALSE, this function requires the glmnet package,
which is listed in Suggests and is not installed automatically with EDI; it errors immediately if
glmnet is not installed.
Value
A list. When use_rcpp = TRUE (default), a list with components
- b, coefficients
A numeric vector of the estimated log-hazard-ratio coefficients
\hat\beta(bandcoefficientsare identical; both are populated for interface consistency with the package's otherfast_*wrappers).- vcov
The variance-covariance matrix of
\hat\beta, orNULLwhenestimate_only = TRUE.- neg_log_lik
The negative Cox partial log-likelihood at the fitted coefficients.
- fisher_information
The Hessian of the negative partial log-likelihood at the fitted coefficients.
When use_rcpp = FALSE, only a single component,
b (the glmnet-fitted coefficient vector via coef(), in
glmnet's own sparse-matrix representation rather than a plain numeric
vector) — none of coefficients/vcov/neg_log_lik/
fisher_information are present on this path.
See Also
build_cox_data_cache_cpp for the full Cox partial-likelihood
model, Breslow tie-handling, and input conventions;
fast_coxph_regression_cpp for the native C++ backend this
wrapper calls by default.
Examples
X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
fast_coxph_regression(X, y, dead)
Fast Cox Proportional Hazards Regression, One-Shot Fit (C++ Backend)
Description
Fits the unstratified Cox proportional-hazards partial-likelihood model
documented in full at build_cox_data_cache_cpp — the same model,
Breslow tie-handling, and input conventions — in a single call that internally
builds the sorted risk-set cache, runs the optimizer, and discards the cache
afterward. Use this entry point for a one-off fit; use
build_cox_data_cache_cpp plus
fast_coxph_regression_prebuilt_cpp instead when fitting the same
(X, y, dead) repeatedly (e.g. across bootstrap/randomization replicates),
to avoid rebuilding the risk-set cache on every call. fast_coxph_regression
is the R-level wrapper around this backend (with an survival-free-of-Rcpp
fallback path via glmnet).
Usage
fast_coxph_regression_cpp(
X,
y,
dead,
warm_start_beta = NULL,
smart_cold_start = TRUE,
estimate_only = FALSE,
maxit = 20L,
tol = 1e-9,
cluster = NULL,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "newton_raphson",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables (no intercept column; see
|
y |
Numeric vector of observed (event or censoring) times. |
dead |
Numeric vector with values in |
warm_start_beta |
Optional starting values for the coefficients |
smart_cold_start |
Logical. If |
estimate_only |
Logical. If |
maxit |
Maximum number of Newton-Raphson/L-BFGS iterations. |
tol |
Convergence tolerance. |
cluster |
Optional clustering variable; when supplied, the returned variance-covariance matrix uses a cluster-robust (grouped) sandwich correction instead of the naive model-based inverse-information variance, i.e. one that remains asymptotically valid under within-cluster correlation of the martingale residuals. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information for the optimizer. |
Value
A list containing the following components:
- coefficients
A numeric vector of the estimated log-hazard-ratio coefficients
\hat\beta.- vcov
The variance-covariance matrix of
\hat\beta(naive inverse-information, or cluster-robust sandwich ifclusteris supplied); omitted/not computed whenestimate_only = TRUE.- neg_ll
The negative Cox partial log-likelihood at the final iteration.
- converged
A logical value indicating whether the algorithm converged.
- iterations
The number of optimizer iterations performed.
- fisher_information
The Hessian of the negative partial log-likelihood at the fitted coefficients (the observed information matrix).
- gradient_norm
The norm of the score (gradient) vector at convergence, a diagnostic of how tightly the convergence criterion was met.
See Also
build_cox_data_cache_cpp for the full Cox partial-likelihood
model, Breslow tie-handling, and input conventions this function implements;
fast_coxph_regression_prebuilt_cpp for the cache-reusing variant;
fast_coxph_regression for the R-level wrapper.
Fast Cox Proportional Hazards Regression, Cache-Reusing Fit (C++ Backend)
Description
Fits the same unstratified-or-stratified Cox partial-likelihood model
documented at build_cox_data_cache_cpp /
build_stratified_cox_data_cache_cpp, but takes a
pre-built risk-set cache (cox_data_xptr, an externalptr
produced by one of those two functions) instead of raw (X, y, dead)
data, skipping the sort/tabulation step on every call. This is the entry point
the package's Cox inference classes (e.g. InferenceCoxPH,
InferenceStratifiedCoxPH) use for repeated fits on the same data
(successive estimate_only vs. full-variance calls, or bootstrap/
randomization replicates that only change the treatment column of X,
rebuilding the cache only when the covariates or assignment actually change).
fast_coxph_regression_cpp is the equivalent one-shot entry point
that builds and discards the cache internally, for callers that only need a
single fit.
Usage
fast_coxph_regression_prebuilt_cpp(
cox_data_xptr,
warm_start_beta = NULL,
smart_cold_start = TRUE,
estimate_only = FALSE,
maxit = 20L,
tol = 1e-9,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "newton_raphson",
warm_start_fisher_info = NULL
)
Arguments
cox_data_xptr |
An |
warm_start_beta |
Optional starting values for the coefficients |
smart_cold_start |
Logical. If |
estimate_only |
Logical. If |
maxit |
Maximum number of Newton-Raphson/L-BFGS iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information for the optimizer. |
Details
Because cox_data_xptr carries a fixed, already-sorted risk-set structure
(whether unstratified — one risk set — or stratified — one risk set per
stratum, depending on which cache-building function produced it), the number
and identity of subjects/strata are entirely determined by the cache; only the
optimization behavior (warm starts, convergence, algorithm) is configurable
through this function's own arguments. Passing a stale cache (built from data
that has since changed) silently fits the model to the cached data, not
the caller's current X/y/dead — callers are responsible for
invalidating and rebuilding the cache when the underlying data changes; see
build_cox_data_cache_cpp for the exact caching/mutation contract.
Value
A list containing the following components:
- coefficients
A numeric vector of the estimated log-hazard-ratio coefficients
\hat\beta.- vcov
The variance-covariance matrix of
\hat\beta; omitted whenestimate_only = TRUE.- neg_ll
The negative Cox partial log-likelihood at the final iteration.
- converged
A logical value indicating whether the algorithm converged.
- iterations
The number of optimizer iterations performed.
- fisher_information
The Hessian of the negative partial log-likelihood at the fitted coefficients.
- gradient_norm
The norm of the score (gradient) vector at convergence.
See Also
build_cox_data_cache_cpp/
build_stratified_cox_data_cache_cpp for building the required
cache and the full Cox partial-likelihood model documentation;
fast_coxph_regression_cpp for the one-shot (build-and-discard)
variant.
Fast Combined Conditional-Poisson + Poisson Regression for KK Matched-Pair/ Reservoir Designs, with Variance (C++ Backend)
Description
Jointly fits a single treatment-effect coefficient \beta_T (and shared
covariate effects \beta_{xs}) across two structurally different count
likelihoods at once — the matched-pair (conditional Poisson) component from
subjects paired on-the-fly by a KK matching design (e.g.
DesignSeqOneByOneKK14) and the
marginal Poisson component from unmatched "reservoir" subjects — rather than
fitting the two subsets separately and combining estimates afterward (as an
inverse-variance-weighted combination does elsewhere in the package). This
one-likelihood joint fit is what backs
InferenceCountKKCondPoissonOneLik-style estimators.
Usage
fast_cpoisson_combined_with_var_cpp(
yT_v_r,
n_k_v_r,
X_diff_v_r,
y_r_r,
w_r_r,
X_r_r,
maxit = 100L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL,
warm_start_params = NULL,
warm_start_beta = NULL,
estimate_only = FALSE
)
Arguments
yT_v_r |
Numeric vector of length |
n_k_v_r |
Numeric vector of length |
X_diff_v_r |
Numeric matrix, |
y_r_r |
Numeric vector of length |
w_r_r |
Numeric vector of length |
X_r_r |
Numeric matrix, |
maxit |
Maximum number of Newton iterations. |
tol |
Convergence tolerance (on the norm of the parameter update step). |
fixed_idx |
Optional integer indices (into the
|
fixed_values |
Optional values to fix the parameters named by
|
warm_start_fisher_info |
Optional initial Fisher Information matrix (over
the full |
warm_start_params |
Optional starting values for the full parameter
vector |
warm_start_beta |
Optional starting values for just
|
estimate_only |
Logical; if |
Details
Matched-pair component (conditional Poisson). For pair k with
total count n_k (sum of both members' counts) and treated-member count
y_{T,k}, conditioning on n_k (the sufficient statistic that
eliminates the pair's nuisance baseline rate) reduces the joint Poisson
likelihood of the pair to a Binomial: y_{T,k} \mid n_k \sim
\mathrm{Binomial}(n_k, p_k), p_k = \mathrm{logit}^{-1}(\beta_T +
x_{\Delta,k}^\top \beta_{xs}), where x_{\Delta,k} is the pair's
covariate difference (treated minus control). This is exactly the
count-response analog of conditional logistic regression for matched pairs —
no per-pair intercept is estimated (it is conditioned out entirely), so only
\beta_T and \beta_{xs} appear in this component.
Reservoir component (marginal Poisson). Unmatched reservoir subjects
contribute an ordinary Poisson log-linear likelihood,
y_i \sim \mathrm{Poisson}(\mu_i), \log \mu_i = \beta_0 + w_i \beta_T +
x_i^\top \beta_{xs}, sharing the same \beta_T and
\beta_{xs} as the pair component but additionally estimating an
intercept \beta_0 (which the conditional pair likelihood has no use
for).
Combined likelihood and optimization. The total log-likelihood is the
simple sum of the pair (conditional Poisson/Binomial) and reservoir (Poisson)
log-likelihoods, jointly maximized over c(beta_0, beta_T, beta_xs)
(length p + 2) via Newton's method using the analytic Fisher
information as the Hessian (quadratic convergence near the optimum, typically
very few iterations). fixed_idx/fixed_values hold specific
parameters fixed rather than estimated; warm_start_params (full vector)
or warm_start_beta (either the full vector, or just
c(beta_T, beta_xs) when of length p + 1, in which case
beta_0 is initialized separately) seed the optimizer, with a
log-mean-based default cold start for beta_0 when neither is supplied.
Variance. ssq_b_j is the variance of \hat\beta_T
specifically (index 1, 0-based, in the parameter layout — the package's usual
single-treatment-coefficient convention), obtained via a targeted diagonal
inverse of the observed/Fisher information restricted to free parameters;
NA if \beta_T was itself fixed via fixed_idx.
Estimate-only mode. If estimate_only = TRUE, optimization
still runs to convergence but the score/information/variance computation is
skipped entirely, returning only b, params, and
converged.
Value
A list with components b/params (the fitted
c(beta_0, beta_T, beta_xs) vector), converged (logical), and,
unless estimate_only = TRUE: ssq_b_j (the variance of
\hat\beta_T), score (the score vector at the fitted
parameters), observed_information/fisher_information/
information (three aliases for the same Fisher information matrix,
also tagged by information_type = "fisher"), hessian (the
Hessian of the negative log-likelihood, i.e. -information),
neg_loglik/neg_ll (aliases for the combined negative
log-likelihood at the fitted parameters), and loglik (its negation).
See Also
Conditional logistic regression for the matched-pair likelihood's structural analog; Poisson regression for the reservoir component; analogous Python API: statsmodels ConditionalPoisson for the conditional-Poisson matched-set likelihood alone (not the combined pair+reservoir model implemented here).
Fast Digamma Function, Vectorized (C++ Backend)
Description
Computes the digamma function \psi(x) = d/dx \log \Gamma(x) elementwise
over x, via an asymptotic expansion with a recurrence (reflection)
shift for small arguments to keep the expansion accurate — the standard
technique for evaluating digamma/trigamma to double precision without a
lookup table. Used internally inside the package's negative-binomial, beta,
zero-inflated/hurdle, and KK21 count-response likelihood, score, and Hessian
kernels (wherever a Poisson/NegBin/Beta log-likelihood derivative requires
\psi), and exported standalone because it is consistently faster than
base R's digamma — measured at 6.78x on a length-5000
vector (see the
"Utility
/ Math Kernel Performance" benchmark report for the full methodology and
per-kernel results).
Usage
fast_digamma_vec_cpp(x)
Arguments
x |
Numeric vector of arguments (should be finite and, per the digamma
function's domain, not a non-positive integer, where |
Value
A numeric vector of \psi(x) values, the same length as x.
References
Abramowitz, M., and Stegun, I. A. (1972). Handbook of
Mathematical Functions, Section 6.3, for the asymptotic expansion and
recurrence relation used. See also
digamma function for
orientation. Analogous Python API:
SciPy
digamma.
Fast Mean-Parameterized Negative-Binomial Density, Vectorized (C++ Backend)
Description
Computes the negative-binomial probability mass function, in its mean/dispersion parameterization,
f(x; \mathrm{size}, \mu) = \binom{x + \mathrm{size} - 1}{x} \left(\frac{\mathrm{size}}{\mathrm{size} + \mu}\right)^{\mathrm{size}} \left(\frac{\mu}{\mathrm{size} + \mu}\right)^{x},
elementwise over x, with E[X] = \mu and
\mathrm{Var}(X) = \mu + \mu^2/\mathrm{size} (size is the
dispersion/shape parameter; smaller size means more overdispersion
relative to Poisson). This matches R::dnbinom_mu(x, size, mu, give_log)
semantics exactly, but evaluates the three required lgamma calls per
observation via fast_lgamma_vec_cpp's kernel instead of R's own
lgamma dispatch, making it faster than base R's
stats::dnbinom(x, size, mu = mu, log = ...) — measured at 1.35x on
a length-5000 vector (see the
"Utility
/ Math Kernel Performance" benchmark report) — while returning
numerically identical values. Used internally inside the package's
negative-binomial regression likelihood, score, and Hessian kernels.
Usage
fast_dnbinom_mu_vec_cpp(x, size, mu, return_log)
Arguments
x |
Numeric vector of non-negative integer counts (non-integer or
negative values are not validated by this function and will produce
incorrect or non-finite results, matching |
size |
Dispersion (shape) parameter |
mu |
Mean parameter |
return_log |
Logical. If |
Value
A numeric vector of (log-)density values, the same length as
x.
References
Negative
binomial distribution for the mean/dispersion parameterization used here.
Analogous Python API:
SciPy stats
distributions index (scipy.stats.nbinom, in its
number-of-successes/probability parameterization — convert via
p = \mathrm{size}/(\mathrm{size}+\mu)).
See Also
fast_lgamma_vec_cpp, whose kernel this function calls
three times per observation.
Fast Hurdle Negative-Binomial Regression, with Variance (C++ Backend)
Description
Fits a two-part hurdle negative-binomial model for count data with excess
zeros: (1) a hurdle part — logistic regression of the binary
indicator I(Y_i > 0) on X_hurdle — models whether the hurdle is
crossed at all, and (2) a count part — a zero-truncated
negative-binomial regression fit only on the subset of subjects with
Y_i > 0, using X — models the count given the hurdle is
crossed. Unlike a zero-inflated model (which mixes a point mass at zero with
an untruncated count distribution that can itself also produce zeros),
the hurdle model's two parts are a clean partition: every zero comes from the
hurdle part, and every positive count's distribution is exactly the
negative-binomial conditional on being positive (left-truncated at 1).
Usage
fast_hurdle_negbin_with_var_cpp(
X_r,
y_r,
X_hurdle_r,
j = 2L,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 1000L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
warm_start_hurdle_fisher_info = NULL
)
Arguments
X_r |
Numeric matrix of predictors for the count component (the
zero-truncated negative-binomial part), |
y_r |
Numeric vector of length |
X_hurdle_r |
Numeric matrix of predictors for the hurdle (zero-vs-
positive) logistic component, |
j |
1-based index (into the count model's |
warm_start_params |
Optional starting values for the count model's full
|
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
maxit |
Maximum number of count-model optimizer iterations. |
tol |
Convergence tolerance (count model). |
fixed_idx |
Optional integer indices (into the count model's
|
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm for the count model; the hurdle logistic part always uses the same algorithm internally. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the count model's first optimizer iteration. |
warm_start_hurdle_fisher_info |
Optional initial Fisher Information matrix for the hurdle logistic model's first optimizer iteration. |
Details
Hurdle part. Fit via fast_logistic_regression_cpp's
internal engine on y_pos_ind = as.numeric(y > 0) regressed on
X_hurdle; if y_pos_ind has no variation (all-zero or
all-positive y), the hurdle part is skipped (hurdle_b is all
NA, hurdle_converged = FALSE) rather than erroring.
Count part. Fit on the positive-count subset
(n_+ = \sum_i I(y_i > 0) rows) via maximum likelihood on the
zero-truncated negative-binomial density with mean-parameterized dispersion
\theta: parameter vector c(beta, log(theta)) (length
p + 1), with \theta optimized on the log scale for positivity
and reported back as theta_hat = exp(params[p+1]). If n_+ \le p
(too few positive observations to identify the count-model coefficients), the
count part returns b as all NA and converged = FALSE
with an explanatory failure_message, while the hurdle part (which does
not depend on n_+) is still fit and returned normally.
Variance. ssq_b_j/ssq_b_2 are the variances of the
j-th and 2nd count-model coefficients (from the count part's observed
information, restricted to free/non-fixed_idx parameters);
hurdle_ssq_b_j/hurdle_ssq_b_2 are the analogous variances for
the hurdle-model coefficients (index j into that model's own
coefficient vector, from the hurdle logistic regression's own information
matrix — no fixed_idx applies to the hurdle part). Both use a targeted
diagonal-entry inversion rather than a full matrix inverse, and are NA
if the relevant coefficient was fixed, out of range, or its model failed to
converge/produce a finite information matrix.
Value
A list with components b (count-model coefficients
\hat\beta), theta_hat (the zero-truncated NB dispersion),
converged (count model), hurdle_b (hurdle logistic
coefficients), hurdle_converged, ssq_b_j/ssq_b_2
(count-model coefficient variances), hurdle_ssq_b_j/
hurdle_ssq_b_2 (hurdle-model coefficient variances),
observed_information/fisher_information/information
(three aliases for the count model's observed information, over
c(beta, log(theta))), information_type = "observed",
hessian (the negative of that information),
hurdle_fisher_information (the hurdle model's own information
matrix), and failure_message (empty on success, otherwise an
explanatory string for a degenerate count-part fit).
See Also
fast_logistic_regression_cpp for the hurdle
component's fitting engine.
Negative
binomial distribution for orientation. Analogous Python API:
statsmodels
discrete models (HurdleCountModel with a negative-binomial count
distribution).
Fast Identity-Link Binomial Regression, Estimate Only (C++ Backend)
Description
Fits a binary-response GLM with the identity link (a linear
probability / risk-difference model), \mu_i = \Pr(Y_i = 1) = x_i^\top \beta
(constrained to (10^{-8}, 1 - 10^{-8}); no other link transformation is
applied), via Fisher scoring (IRLS) with a step-halving line search that
rejects any Newton step whose resulting \eta_i = x_i^\top \beta would
leave the valid probability range or decrease the log-likelihood — this
boundary-constrained line search, not a link-function transform, is what
keeps fitted probabilities in (0, 1) for this otherwise-unconstrained
linear-in-\beta model. Regression coefficients on the identity-link
scale are directly interpretable as risk differences: \beta_j
is the change in \Pr(Y = 1) per unit change in covariate j, in
contrast to fast_log_binomial_regression_cpp's log-link coefficients
(interpretable as log relative risks) or a standard logit-link model's
log-odds-ratio coefficients.
Usage
fast_identity_binomial_regression_cpp(
X,
y_r,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_beta = NULL,
smart_cold_start = TRUE,
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y_r |
A binary (0/1) numeric vector of responses, length |
maxit |
Maximum number of Fisher-scoring iterations. |
tol |
Convergence tolerance, on the relative norm of the coefficient update step. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
warm_start_beta |
Optional starting values for coefficients. If provided, |
smart_cold_start |
Logical. If |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
Details
Optimization. Each Fisher-scoring iteration solves a weighted
least-squares step using working weights
w_i = 1 / \max(\mu_i(1-\mu_i), 10^{-8}) (the inverse Bernoulli variance,
clamped away from 0 for stability near the boundary), then backtracks
(halving the step size, down to a minimum step of 10^{-8}) until the
resulting \eta stays within (10^{-8}, 1-10^{-8}) for every
observation and the log-likelihood does not decrease; a step that
cannot be accepted at any halving depth terminates iteration without
converged = TRUE. fixed_idx/fixed_values hold specific
coefficients fixed rather than estimated; warm_start_beta (or, when
absent, an OLS-based guess if smart_cold_start = TRUE) seeds the
first iteration, and warm_start_weights/warm_start_fisher_info
warm-start the first IRLS working-weights/curvature computation.
No guarantee of a feasible solution. Because the identity link has
no inherent boundary protection, some (X, y) configurations (e.g.
extreme covariate values, near-perfect separation, or an ill-conditioned
X) may have no interior maximum-likelihood solution reachable by this
constrained line search; such cases surface as converged = FALSE
rather than a silently invalid (out-of-range) fitted probability.
Value
A list with components b (estimated coefficients
\hat\beta, on the risk-difference/identity scale), mu_hat
(fitted probabilities \hat\mu_i, length n), working_weights
(the final IRLS weights w_i), iterations (number of Fisher-
scoring iterations performed), converged (logical; also requires all
of b, mu_hat, working_weights to be finite), and
fisher_information (the working-weights curvature matrix
X^\top W X).
See Also
fast_identity_binomial_regression_with_var_cpp for the
variance-augmented variant; fast_identity_binomial_regression_weighted_cpp
for the row-weighted variant; fast_log_binomial_regression_cpp for
the log-link (relative-risk) analog of this model.
Generalized
linear model for orientation. Analogous Python API:
statsmodels GLM
(families.Binomial(link=identity())).
Fast Weighted Identity-Link Binomial Regression, Estimate Only (C++ Backend)
Description
Fits the same identity-link (risk-difference) binomial regression as
fast_identity_binomial_regression_cpp (see that page for the
full model, boundary-constrained IRLS line search, and interpretation), with
each observation's contribution to the log-likelihood and IRLS working
weights multiplied by a nonnegative row weight weights_r[i]. Setting
all weights to 1 recovers fast_identity_binomial_regression_cpp
exactly; this is the backend used when the identity-link model must be fit on
bootstrap-reweighted or otherwise weighted data.
Usage
fast_identity_binomial_regression_weighted_cpp(
X,
y_r,
weights_r,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_beta = NULL,
smart_cold_start = TRUE,
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y_r |
A binary (0/1) numeric vector of responses, length |
weights_r |
A nonnegative numeric vector of length |
maxit |
Maximum number of Fisher-scoring iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
warm_start_beta |
Optional starting values for coefficients. If provided, |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
Value
A list with the same components as
fast_identity_binomial_regression_cpp: b,
mu_hat, working_weights, iterations, converged,
and fisher_information (all reflecting the weighted log-likelihood).
See Also
fast_identity_binomial_regression_cpp for the
unweighted model and full documentation;
fast_identity_binomial_regression_with_var_cpp for the
(unweighted) variance-augmented variant.
Fast Identity-Link Binomial Regression with Targeted Variance (C++ Backend)
Description
Fits the same identity-link (risk-difference) binomial regression as
fast_identity_binomial_regression_cpp (see that page for the
full model and boundary-constrained IRLS line search) and additionally
computes the variance of a single caller-selected coefficient, via a targeted
diagonal-entry inversion of the working-weights Fisher information — this
entry point does not compute or return a full variance-covariance
matrix or a vector of standard errors for every coefficient, despite its
name; only the one coefficient named by j gets a variance
(ssq_b_j).
Usage
fast_identity_binomial_regression_with_var_cpp(
X,
y_r,
j = 2L,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_beta = NULL,
smart_cold_start = TRUE,
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y_r |
A binary (0/1) numeric vector of responses, length |
j |
1-based index (into |
maxit |
Maximum number of Fisher-scoring iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
warm_start_beta |
Optional starting values for coefficients. If provided, |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
Details
Variance computation. The IRLS working-weights Fisher information
X^\top W X (reused from the underlying fit if finite and correctly
sized, else recomputed from the final working weights) is restricted to the
free (non-fixed_idx) parameters and factorized via LDLT; ssq_b_j
is then obtained from a single targeted diagonal-entry inversion
(compute_diagonal_inverse_entry()) at the free-parameter position
corresponding to j, not a full matrix inverse. If the underlying fit
did not converge, or the LDLT factorization fails (e.g. a rank-deficient
free-parameter information matrix), the function returns early with
converged = FALSE, ssq_b_j = NA, and empty
(zero-length/zero-dimension) vcov/std_err/z_vals
placeholders — these three fields are only ever populated as empty
placeholders, on both the success and failure paths; no caller should rely
on them containing actual values.
Value
A list with components b (estimated coefficients
\hat\beta), ssq_b_j (the variance of \hat\beta_j, or
NA on failure), converged (logical), fisher_information
(the working-weights curvature matrix used for ssq_b_j, present only
on the success path), neg_ll/logLik (the negative/positive
log-likelihood at \hat\beta, present only on the success path), and
the always-empty vcov/std_err/z_vals placeholders
described in Details.
See Also
fast_identity_binomial_regression_cpp for the
estimate-only variant and the full model documentation.
Fast Log-Beta Function, Vectorized (C++ Backend)
Description
Computes the log of the Beta function,
\log B(a, b) = \log\Gamma(a) + \log\Gamma(b) - \log\Gamma(a+b),
elementwise, via three calls into fast_lgamma_vec_cpp's kernel
rather than R's own lgamma dispatch — faster than base R's
lbeta — measured at 2.43x on a length-5000 vector (see
the
"Utility
/ Math Kernel Performance" benchmark report) — while returning numerically
identical values (up to the Lanczos/Stirling approximation's own
precision). Used internally inside the
package's beta-regression and beta-distribution-based (zero-one-inflated
beta) likelihood, score, and Hessian kernels, wherever a Beta-density
normalizing constant is required.
Usage
fast_lbeta_vec_cpp(a, b)
Arguments
a |
Numeric vector of first shape arguments (should be positive; not validated by this function). |
b |
Numeric vector of second shape arguments (should be positive; not
validated), recycled against |
Value
A numeric vector of \log B(a, b) values, the same length as
a/b.
References
Beta
function for orientation. Analogous Python API:
SciPy
betaln.
See Also
fast_lgamma_vec_cpp, whose kernel this function calls.
Fast Log-Gamma Function, Vectorized (C++ Backend)
Description
Computes \log \Gamma(x) elementwise over x, via a Lanczos
approximation (with a Stirling-series tail for large arguments) — faster
than base R's lgamma while matching it to within
the approximation's own precision. Used pervasively throughout the package's
likelihood kernels (beta, negative-binomial, Poisson/count, and other
Gamma-function-based densities) wherever a log-factorial-like normalizing
term is required, and exported standalone for the same reason as
fast_digamma_vec_cpp — measured at 2.18x over
lgamma on a length-5000 vector (see the
"Utility
/ Math Kernel Performance" benchmark report for the full methodology and
per-kernel results).
Usage
fast_lgamma_vec_cpp(x)
Arguments
x |
Numeric vector of arguments (should be positive, or a non-positive non-integer if the reflection formula is supported by the underlying kernel; not validated by this function — see the package's C++ source for the exact domain the Lanczos kernel handles). |
Value
A numeric vector of \log \Gamma(x) values, the same length as
x.
References
Lanczos
approximation and Stirling's
approximation for the numerical techniques used; see also
Gamma function for
orientation. Analogous Python API:
SciPy
gammaln.
See Also
fast_digamma_vec_cpp, fast_trigamma_vec_cpp,
fast_lbeta_vec_cpp (built on this function's kernel).
Fast Log-Link Binomial Regression, Estimate Only (C++ Backend)
Description
Fits a binary-response GLM with the log link (a relative-risk
model), \log \mu_i = \log \Pr(Y_i = 1) = x_i^\top \beta (equivalently
\mu_i = e^{x_i^\top \beta}, constrained to stay below
1 - 10^{-8} so it remains a valid probability), via Fisher scoring
(IRLS) with a step-halving line search that rejects any Newton step whose
resulting \eta_i = x_i^\top \beta would push \mu_i out of range
or decrease the log-likelihood — the same boundary-constrained-line-search
mechanism documented in full at
fast_identity_binomial_regression_cpp (see that page for the
IRLS/line-search mechanics, which are shared verbatim between the log and
identity links here; only the link function itself, and hence the
coefficient scale, differs). Regression coefficients are directly
interpretable as log relative risks: e^{\beta_j} is the
multiplicative change in \Pr(Y = 1) per unit change in covariate
j — in contrast to
fast_identity_binomial_regression_cpp's risk-difference scale,
or a logit-link model's odds-ratio scale.
Usage
fast_log_binomial_regression_cpp(
X,
y_r,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_beta = NULL,
smart_cold_start = TRUE,
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y_r |
A binary (0/1) numeric vector of responses, length |
maxit |
Maximum number of Fisher-scoring iterations. |
tol |
Convergence tolerance, on the relative norm of the coefficient update step. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
warm_start_beta |
Optional starting values for coefficients. If provided, |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
Value
A list with components b (estimated coefficients
\hat\beta, on the log-relative-risk scale), mu_hat (fitted
probabilities, length n), working_weights (final IRLS
weights), iterations, converged (logical), and
fisher_information (the working-weights curvature matrix
X^\top W X).
See Also
fast_identity_binomial_regression_cpp for the
identity-link (risk-difference) analog and the full IRLS/line-search
mechanics; fast_log_binomial_regression_with_var_cpp for the
variance-augmented variant; fast_log_binomial_regression_weighted_cpp
for the row-weighted variant.
Poisson
regression's log link is the closest common orientation point for a
log-link GLM. Analogous Python API:
statsmodels GLM
(families.Binomial(link=log())).
Fast Weighted Log-Link Binomial Regression, Estimate Only (C++ Backend)
Description
Fits the same log-link (relative-risk) binomial regression as
fast_log_binomial_regression_cpp (see that page for the full
model and boundary-constrained IRLS line search), with each observation's
contribution to the log-likelihood and IRLS working weights multiplied by a
nonnegative row weight weights_r[i]. Setting all weights to 1
recovers fast_log_binomial_regression_cpp exactly.
Usage
fast_log_binomial_regression_weighted_cpp(
X,
y_r,
weights_r,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_beta = NULL,
smart_cold_start = TRUE,
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y_r |
A binary (0/1) numeric vector of responses, length |
weights_r |
A nonnegative numeric vector of length |
maxit |
Maximum number of Fisher-scoring iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
warm_start_beta |
Optional starting values for coefficients. If provided, |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
Value
A list with the same components as
fast_log_binomial_regression_cpp: b, mu_hat,
working_weights, iterations, converged, and
fisher_information (all reflecting the weighted log-likelihood).
See Also
fast_log_binomial_regression_cpp for the unweighted
model and full documentation.
Fast Log-Link Binomial Regression with Targeted Variance (C++ Backend)
Description
Fits the same log-link (relative-risk) binomial regression as
fast_log_binomial_regression_cpp (see that page for the full
model) and additionally computes the variance of a single caller-selected
coefficient — the log-link analog of
fast_identity_binomial_regression_with_var_cpp, sharing
exactly the same targeted-diagonal-entry variance mechanism and the same
caveat: this entry point does not compute or return a full
variance-covariance matrix or per-coefficient standard errors, despite its
name; only the coefficient named by j gets a variance
(ssq_b_j), and the returned vcov/std_err/z_vals
fields are always empty placeholders (see
fast_identity_binomial_regression_with_var_cpp's Details for
the exact mechanics, identical here up to the link function).
Usage
fast_log_binomial_regression_with_var_cpp(
X,
y_r,
j = 2L,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_beta = NULL,
smart_cold_start = TRUE,
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, |
y_r |
A binary (0/1) numeric vector of responses, length |
j |
1-based index (into |
maxit |
Maximum number of Fisher-scoring iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
warm_start_beta |
Optional starting values for coefficients. If provided, |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
Value
A list with components b, ssq_b_j, converged,
fisher_information, neg_ll/logLik (present only on
the success path), and the always-empty vcov/std_err/
z_vals placeholders; see
fast_identity_binomial_regression_with_var_cpp for the exact
field semantics (shared verbatim here).
See Also
fast_log_binomial_regression_cpp for the
estimate-only variant; fast_identity_binomial_regression_with_var_cpp
for the identity-link analog with the same targeted-variance mechanism.
Fast Log Standard Normal Density, Vectorized (C++ Backend)
Description
Computes \log \phi(x) = -\tfrac{1}{2}\log(2\pi) - x^2/2, the log-density
of the standard normal distribution, elementwise over x, via a
direct closed-form evaluation — no series expansion or special-function
dispatch is needed since the standard normal log-density has an exact
elementary closed form. Faster than base R's
dnorm(x, log = TRUE) — measured at 5x on a
length-5000 vector (see the
"Utility
/ Math Kernel Performance" benchmark report) — while returning
numerically identical values. Used internally inside the package's probit
regression and other Gaussian-likelihood kernels wherever a standard normal
log-density is required.
Usage
fast_log_dnorm_vec_cpp(x)
Arguments
x |
Numeric vector of arguments. |
Value
A numeric vector of \log \phi(x) values, the same length as
x.
References
Normal
distribution for orientation. Analogous Python API:
SciPy stats
distributions index (scipy.stats.norm.logpdf).
See Also
fast_log_pnorm_vec_cpp for the corresponding log-CDF
kernel; fast_qnorm_vec_cpp for the standard normal quantile
function.
Fast Log Standard Normal CDF, Vectorized (C++ Backend)
Description
Computes \log \Phi(x), the log of the standard normal cumulative
distribution function, elementwise over x, via the complementary
error function kernel fast_erfc (\Phi(x) = \tfrac{1}{2}
\mathrm{erfc}(-x/\sqrt{2}), evaluated in a form stable for large negative
x, where \Phi(x) underflows in ordinary (non-log) arithmetic
long before the true log-probability does), avoiding R's own
pnorm dispatch overhead. Faster than base R's
pnorm(x, log.p = TRUE) — measured at 2.49x on a
length-5000 vector (see the
"Utility
/ Math Kernel Performance" benchmark report). Used internally inside the
package's probit regression and other likelihood kernels that need a
numerically stable normal log-CDF, e.g. for censored/truncated Gaussian
contributions.
Usage
fast_log_pnorm_vec_cpp(x)
Arguments
x |
Numeric vector of arguments. |
Value
A numeric vector of \log \Phi(x) values, the same length as
x.
References
Normal
distribution for orientation. Analogous Python API:
SciPy stats
distributions index (scipy.stats.norm.logcdf).
See Also
fast_log_dnorm_vec_cpp for the corresponding
log-density kernel; fast_qnorm_vec_cpp for the standard
normal quantile function.
Fast Logistic Regression, Estimate Only (R Wrapper)
Description
Fits the logistic regression model documented in full at
fast_logistic_regression_cpp (log-odds-ratio interpretation,
IRLS/L-BFGS/Newton-Raphson optimization) via that C++ backend, returning
only the point estimate \hat\beta — no variance-covariance matrix or
per-coefficient standard errors are computed. Unlike
fast_logistic_regression_with_var, this function does
not attempt to detect or retry on (quasi-)complete separation; if
the underlying C++ fit errors for any reason, this function silently
returns b as a vector of NAs (of length ncol(X))
rather than raising an error or retrying with fewer covariates.
Usage
fast_logistic_regression(
X,
y,
optimization_alg = "lbfgs",
warm_start_beta = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be binary (0 or 1). |
optimization_alg |
Optimization algorithm: |
warm_start_beta |
Optional starting values for the coefficients. |
warm_start_fisher_info |
Optional initial Fisher Information matrix. |
Value
A list containing the following component:
- b
A numeric vector of the estimated logistic regression coefficients
\hat\beta, or a vector ofNA_real_(lengthncol(X)) if the underlying fit errored.
See Also
fast_logistic_regression_cpp for the underlying
backend and full model documentation;
fast_logistic_regression_with_var for the variance-
augmented, separation-retrying variant.
Examples
X = matrix(rnorm(500), 100, 5)
y = rbinom(100, 1, 0.5)
fast_logistic_regression(X, y)
Fast Logistic Regression, Estimate Only (C++ Backend)
Description
Fits the standard binary logistic regression model,
\mathrm{logit}(\mu_i) = \Pr(Y_i = 1) \text{'s log-odds} = x_i^\top \beta,
\mu_i = \mathrm{logit}^{-1}(x_i^\top \beta), via maximum likelihood.
Coefficients are directly interpretable as log odds ratios:
e^{\beta_j} is the multiplicative change in the odds
\mu_i / (1 - \mu_i) per unit change in covariate j. This is the
package's baseline binary-response fitting backend, used wherever an
incidence/binary outcome needs a logit-link fit (as opposed to the log-link
or identity-link constrained binomial models in
fast_log_binomial_regression_cpp/
fast_identity_binomial_regression_cpp, which target relative
risk / risk difference scales instead of odds ratios).
Usage
fast_logistic_regression_cpp(
X,
y,
warm_start_beta = NULL,
smart_cold_start = FALSE,
maxit = 100L,
tol = 1e-8,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be binary (0 or 1). |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations for the algorithm. Defaults to 100. |
tol |
Convergence tolerance. Defaults to 1e-8. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Optional initial IRLS working weights for the first iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information. |
estimate_only |
Logical. If |
Value
A list containing the following components:
- b
A numeric vector of the estimated logistic regression coefficients
\hat\beta.- w
The IRLS working weights
\hat\mu_i(1-\hat\mu_i)at the final iteration (the Bernoulli variance function evaluated at the fitted probabilities); omitted whenestimate_only = TRUE.- num_iter
The number of optimizer iterations performed.
- fisher_information
The working-weights curvature matrix
X^\top W X; omitted whenestimate_only = TRUE.- score
The score (gradient of the log-likelihood) vector at the fitted coefficients; omitted when
estimate_only = TRUE.- neg_ll
The negative log-likelihood at the fitted coefficients; omitted when
estimate_only = TRUE.- converged
A logical value indicating whether the final gradient norm was below
tol(gradient_norm < tol); uniform across the"irls"/"lbfgs"optimizers.- hit_iteration_cap
A logical value, mutually exclusive with
converged:TRUEiff the optimizer exhaustedmaxititerations without meeting the gradient-norm convergence criterion.- gradient_norm
The norm of the score vector at the returned coefficients, a diagnostic of how tightly the convergence criterion was met.
See Also
fast_logistic_regression_with_var_cpp for the
variance-augmented variant; fast_logistic_regression for the
R-level wrapper; fast_log_binomial_regression_cpp/
fast_identity_binomial_regression_cpp for the log-link/
identity-link analogs targeting relative-risk/risk-difference scales.
Fast Weighted Logistic Regression, Estimate Only (C++ Backend)
Description
Fits the same logistic regression model as
fast_logistic_regression_cpp (see that page for the full
model, log-odds-ratio interpretation, and optimizer contract), with each
observation's contribution to the log-likelihood and IRLS working weights
multiplied by a row weight weights[i]. Setting all weights to 1
recovers fast_logistic_regression_cpp exactly; this is the
backend used when the logistic model must be fit on bootstrap-reweighted or
otherwise weighted data.
Usage
fast_logistic_regression_weighted_cpp(
X,
y,
weights,
warm_start_beta = NULL,
smart_cold_start = FALSE,
maxit = 100L,
tol = 1e-8,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be binary (0 or 1). |
weights |
A numeric vector of weights for each observation. |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations for the IRLS algorithm. Defaults to 100. |
tol |
Convergence tolerance. Defaults to 1e-8. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Optional initial IRLS working weights for the first iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information. |
Value
A list containing the following components:
- b
A numeric vector of the estimated logistic regression coefficients
\hat\beta.- mu
The fitted probabilities
\hat\mu_i.- XtWX, fisher_information
Two aliases for the same working-weights curvature matrix
X^\top W Xat the final iteration.- score
The (weighted) score vector at the fitted coefficients.
- neg_ll
The weighted negative log-likelihood at the fitted coefficients.
- converged
A logical value indicating whether the final gradient norm was below
tol(gradient_norm < tol); uniform across the"irls"/"lbfgs"optimizers.- num_iter
The number of optimizer iterations performed.
- hit_iteration_cap
A logical value, mutually exclusive with
converged:TRUEiff the optimizer exhaustedmaxititerations without meeting the gradient-norm convergence criterion.- gradient_norm
The norm of the score vector at the returned coefficients.
See Also
fast_logistic_regression_cpp for the unweighted model
and full documentation.
Fast Logistic Regression with Variance, Auto-Retrying on Separation (R Wrapper)
Description
Fits the logistic regression model documented in full at
fast_logistic_regression_cpp (log-odds-ratio interpretation,
IRLS/L-BFGS/Newton-Raphson optimization) via
fast_logistic_regression_with_var_cpp, and additionally
detects and automatically retries on (quasi-)complete separation —
the well-known logistic-regression failure mode where the MLE does not
exist because some linear combination of covariates perfectly (or
near-perfectly) predicts the outcome, causing the optimizer's coefficient
estimates to diverge to a large-but-finite value that would otherwise
silently pass ordinary is.finite() convergence checks and corrupt
downstream confidence intervals.
Usage
fast_logistic_regression_with_var(
X,
y,
j = 2,
optimization_alg = "lbfgs",
warm_start_beta = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be binary (0 or 1). |
j |
The index of the coefficient to compute the variance for. Defaults to 2. |
optimization_alg |
Optimization algorithm: |
warm_start_beta |
Optional starting values for the coefficients. |
warm_start_fisher_info |
Optional initial Fisher Information matrix. |
Details
Separation detection and retry. After each fit attempt,
is_separated_coefficient_magnitude() checks whether any fitted
coefficient exceeds a fixed separation-detection threshold
(EDI_SEPARATION_THRESHOLD); if so, the fit is treated as
converged = FALSE regardless of what the underlying C++ optimizer
itself reported. On non-convergence (including detected separation), the
covariate (column index \ge 3, i.e. never the intercept in
column 1 or the treatment column in column 2) with the largest absolute
fitted coefficient is dropped, and the model is refit on the reduced
design; this repeats until either the fit converges or only the intercept
and treatment columns remain. If separation persists even with just those
two columns, this function stop()s with an explicit
"complete separation detected" error rather than returning a corrupted
variance estimate.
Interpretation caveat. Because covariates can be silently dropped
by this retry loop, the returned b may have fewer coefficients than
ncol(X) implies, and the fitted model's covariate adjustment set can
differ from what was requested; callers relying on a specific covariate
being present in the final fit should check for this rather than assume it
always is.
Value
A list containing the following components:
- b
A numeric vector of the obtained logistic regression coefficients
\hat\beta, from whichever (possibly covariate-reduced) fit in the retry sequence ultimately converged.- ssq_b_j
The squared standard error (variance) of the j-th estimated coefficient.
- ssq_b_2
The squared standard error (variance) of the second estimated coefficient, which typically corresponds to the treatment effect.
See Also
fast_logistic_regression_with_var_cpp for the
underlying single-fit (no retry) backend and its variance-computation
details; fast_logistic_regression_cpp for the full model
documentation.
Examples
X = matrix(rnorm(100), 10, 10)
y = rbinom(10, 1, 0.5)
fast_logistic_regression_with_var(X, y)
Fast Logistic Regression with Targeted Variance (C++ Backend)
Description
Fits the same logistic regression model as
fast_logistic_regression_cpp (see that page for the full
model and log-odds-ratio interpretation) and additionally computes the
variance of two coefficients — the caller-selected j-th coefficient
and, separately, the 2nd coefficient (the package's usual treatment-effect
position) — via a targeted diagonal-entry inversion of the working-weights
Fisher information, rather than a full matrix inverse. Unlike
fast_logistic_regression_cpp, maxit and tol are
not exposed here: they are fixed internally at 100 iterations and
1e-8 tolerance.
Usage
fast_logistic_regression_with_var_cpp(
X,
y,
j = 2L,
warm_start_beta = NULL,
smart_cold_start = FALSE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be binary (0 or 1). |
j |
1-based index (into |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Optional initial IRLS working weights for the first iteration. |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information. |
Details
Variance computation. The working-weights Fisher information
X^\top W X is restricted to the free (non-fixed_idx)
parameters; ssq_b_j and ssq_b_2 are each obtained via a
single targeted diagonal-entry inversion
(compute_diagonal_inverse_entry()) at the free-parameter position
corresponding to j and to column 2, respectively — NA if the
relevant coefficient is out of range or was itself fixed via
fixed_idx. ssq_b_2 is always computed (when ncol(X) >= 2)
regardless of what j is, so a caller interested in the treatment
effect's variance does not need to pass j = 2 explicitly.
Value
A list with components b/params (estimated
coefficients \hat\beta, identical to each other),
ssq_b_j/ssq_b_2 (the two targeted coefficient variances
described in Details), score (the score vector at
\hat\beta), observed_information/fisher_information/
information (three aliases for the same working-weights curvature
matrix, tagged information_type = "fisher"), hessian (the
negative of that matrix), neg_loglik/neg_ll (aliases for the
negative log-likelihood), loglik (its negation), converged
(logical, gradient_norm < tol, uniform across optimizers),
num_iter, hit_iteration_cap (logical, mutually exclusive
with converged), and gradient_norm.
See Also
fast_logistic_regression_cpp for the estimate-only
variant and the full model documentation.
Fast Negative Binomial Regression, Estimate Only (C++ Backend)
Description
Fits the negative-binomial regression model in its mean/dispersion
parameterization documented in full at
fast_dnbinom_mu_vec_cpp: log link
\log \mu_i = x_i^\top \beta (so e^{\beta_j} is a multiplicative
change in the mean count, as in Poisson regression), with a single
dispersion parameter \theta shared across all observations and
\mathrm{Var}(Y_i) = \mu_i + \mu_i^2/\theta (smaller \theta means
more overdispersion relative to Poisson; \theta \to \infty recovers
Poisson). The optimizer's parameter vector is c(beta, log(theta))
(\theta optimized on the log scale for positivity).
High-performance negative binomial regression fitting.
Usage
fast_neg_bin_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = FALSE,
maxit = 1000L,
eps_f = 1e-08,
eps_g = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of responses. |
warm_start_params |
Optional starting values for coefficients and dispersion. If provided, |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
maxit |
Maximum number of iterations. |
eps_f |
Convergence tolerance for function value. |
eps_g |
Convergence tolerance for gradient. |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
optimization_alg |
Optimization algorithm. |
warm_start_fisher_info |
Optional initial Fisher Information matrix for the first IRLS iteration. |
estimate_only |
Logical; if |
Value
A list containing the following components:
- b
A numeric vector of the estimated negative binomial regression coefficients
\hat\beta.- theta_hat
The estimated dispersion parameter
\hat\theta(natural scale).- logLik
The model's log-likelihood at the fitted parameters.
- converged
A logical value indicating whether the final gradient norm was below the convergence tolerance (
gradient_norm < tol); uniform across the"lbfgs"/"newton_raphson"optimizers.- num_iter
The number of optimizer iterations performed.
- hit_iteration_cap
A logical value, mutually exclusive with
converged:TRUEiff the optimizer exhaustedmaxititerations without meeting the gradient-norm convergence criterion.- gradient_norm
The norm of the gradient at the returned parameters.
- fisher_information
The working-weights curvature matrix used during fitting.
A list containing coefficients, theta, and convergence status.
See Also
fast_dnbinom_mu_vec_cpp for the mean/dispersion
density parameterization used here; fast_neg_bin_with_var_cpp
for the variance-augmented variant.
Examples
X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_neg_bin_cpp(X, y)
Fast Weighted Negative Binomial Regression, Estimate Only (C++ Backend)
Description
Fits the same mean/dispersion-parameterized negative-binomial regression as
fast_neg_bin_cpp (see that page, and
fast_dnbinom_mu_vec_cpp, for the full model), with each
observation's contribution to the log-likelihood multiplied by a
nonnegative row weight weights[i]. Setting all weights to 1 recovers
fast_neg_bin_cpp exactly.
Usage
fast_neg_bin_weighted_cpp(
X,
y,
weights,
warm_start_params = NULL,
smart_cold_start = FALSE,
maxit = 1000L,
eps_f = 1e-08,
eps_g = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric (integer-valued) vector of non-negative observed
counts, length |
weights |
A nonnegative numeric vector of length |
warm_start_params |
Optional starting values for coefficients and dispersion. If provided, |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
maxit |
Maximum number of optimizer iterations. |
eps_f |
Convergence tolerance on the objective (log-likelihood) value. |
eps_g |
Convergence tolerance on the gradient norm. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information. |
estimate_only |
If TRUE, skip Fisher information calculation. |
Value
A list with the same components as fast_neg_bin_cpp:
b, theta_hat, logLik, converged,
iterations, and fisher_information (all reflecting the
weighted log-likelihood).
See Also
fast_neg_bin_cpp for the unweighted model and full
documentation.
Examples
X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_neg_bin_weighted_cpp(X, y, weights = rep(1, 10))
Fast Negative Binomial Regression with Variance Calculation (C++ Backend)
Description
Fits the same mean/dispersion-parameterized negative-binomial regression as
fast_neg_bin_cpp (see that page, and
fast_dnbinom_mu_vec_cpp, for the full model) and additionally
computes the full variance-covariance matrix of c(beta, log(theta)).
Usage
fast_neg_bin_with_var_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = FALSE,
maxit = 1000L,
eps_f = 1e-08,
eps_g = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors, |
y |
A numeric (integer-valued) vector of non-negative observed
counts, length |
warm_start_params |
Optional starting values for coefficients and dispersion. If provided, |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess when starting from scratch (a "cold start") with no prior knowledge. This is ignored if a warm start is provided. |
maxit |
Maximum number of optimizer iterations. |
eps_f |
Convergence tolerance on the objective (log-likelihood) value. |
eps_g |
Convergence tolerance on the gradient norm. |
fixed_idx |
Optional integer indices (into the |
fixed_values |
Optional values to fix the parameters named by
|
optimization_alg |
Optimization algorithm: |
warm_start_fisher_info |
Optional initial Fisher Information matrix to warm-start curvature information. |
Details
Variance computation. The working-weights curvature matrix
res.XtWX (over all p + 1 parameters) is restricted to the
free (non-fixed_idx) parameters and inverted via a plain
matrix inverse (.inverse(), not a rank-aware pseudo-inverse as used
by, e.g., fast_adjacent_category_logit_with_var_cpp) before
being expanded back to the full (p + 1) x (p + 1) size as
vcov. A rank-deficient or near-singular design (after restricting to
free parameters) will therefore produce numerically unstable or NaN
variances rather than a graceful fallback.
Value
A list containing the following components:
- b
A numeric vector of the obtained Poisson regression coefficients.
- hess_fisher_info_matrix
The Fisher information matrix.
- neg_ll
The negative log-likelihood at the final iteration.
- converged
A logical value indicating whether the final gradient norm was below the convergence tolerance (
gradient_norm < tol); uniform across the"lbfgs"/"newton_raphson"optimizers.- num_iter
The number of optimizer iterations performed.
- hit_iteration_cap
A logical value, mutually exclusive with
converged:TRUEiff the optimizer exhaustedmaxititerations without meeting the gradient-norm convergence criterion.- gradient_norm
The norm of the gradient at the returned parameters.
A list with components b (\hat\beta),
theta_hat (\hat\theta), logLik, vcov (the full
(p + 1) x (p + 1) parameter variance-covariance matrix),
converged, iterations, and hess_fisher_info_matrix
(the working-weights curvature matrix vcov was inverted from).
See Also
fast_neg_bin_cpp for the estimate-only variant and
full model documentation; fast_neg_bin_weighted_cpp for the
row-weighted estimate-only variant.
Examples
X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_neg_bin_with_var_cpp(X, y)
Fast Negative Binomial Regression, Estimate-Only (R Wrapper)
Description
This function provides a fast implementation of mean/dispersion-parameterized
negative binomial regression, wrapping a C++ backend
(fast_neg_bin_cpp; see that page, and
fast_dnbinom_mu_vec_cpp, for the full model). It returns point
estimates only (no standard errors) — see fast_negbin_regression_with_var
for the variance-computing counterpart. Columns 1 and 2 of X (conventionally
the intercept and treatment indicator) are always kept; this wrapper adds automatic,
silent handling of rank-deficient or numerically unstable covariate sets
beyond those first two columns: (1) if warm_start_params is supplied, a single
fit is attempted on the full matrix using the warm start, and its result is returned
if successful; (2) otherwise, upfront, any covariate columns (3 onward) found
rank-deficient by qr are dropped before the first fit attempt;
(3) if the C++ fit still fails (e.g. an L-BFGS line-search failure), covariates are
dropped one at a time, in reverse QR-pivot order (most redundant first), retrying
after each drop, until the fit succeeds or only the intercept and treatment columns
remain — at which point, if it still fails, this function stop()s with an
explicit error rather than returning a corrupted fit.
Usage
fast_negbin_regression(
X,
y,
optimization_alg = "lbfgs",
warm_start_params = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, representing count data. |
optimization_alg |
Optimization algorithm: |
warm_start_params |
Optional starting values for coefficients and |
warm_start_fisher_info |
Optional initial Fisher information matrix, used only together
with |
Value
A list containing the following components:
b |
A numeric vector of the estimated negative binomial regression coefficients
|
fisher_information |
The C++ backend's returned Fisher information matrix for the successful fit. |
See Also
fast_negbin_regression_with_var for the variance-computing,
similarly retry-hardened wrapper; fast_neg_bin_cpp for the underlying
backend and full model documentation.
Examples
X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_negbin_regression(X, y)
Fast Negative Binomial Regression with Variance Calculation (R Wrapper)
Description
This function provides a fast implementation of negative binomial regression, wrapping
a C++ backend (fast_neg_bin_with_var_cpp; see that page, and
fast_dnbinom_mu_vec_cpp, for the full mean/dispersion
negative-binomial model). Columns 1 and 2 of X (conventionally the
intercept and treatment indicator) are always kept; this wrapper adds
automatic, silent handling of rank-deficient or numerically
unstable covariate sets beyond those first two columns: (1) upfront, any
covariate columns (3 onward) found rank-deficient by qr
are dropped before the first fit attempt; (2) if the C++ fit still fails
(e.g. an L-BFGS line-search failure), covariates are dropped one at a time,
in reverse QR-pivot order (most redundant first), retrying after each drop,
until the fit succeeds or only the intercept and treatment columns remain —
at which point, if it still fails, this function stop()s with an
explicit error rather than returning a corrupted fit.
Usage
fast_negbin_regression_with_var(X, y, j = 2, optimization_alg = "lbfgs")
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, representing count data. |
j |
The index of the coefficient to compute the variance for. Defaults to 2. |
optimization_alg |
Optimization algorithm: |
Value
A list containing the following components:
- b
A numeric vector of the estimated negative binomial regression coefficients
\hat\beta, from whichever (possibly covariate-reduced) fit in the retry sequence ultimately succeeded.- ssq_b_j
The variance of the j-th estimated coefficient, computed by inverting the C++ backend's returned Hessian/Fisher-information matrix directly (via
solve, not the backend's ownvcov);NAif that inversion fails orjexceeds the number of columns remaining after covariate-dropping.- ssq_b_2
The variance of the second estimated coefficient specifically (typically the treatment effect), computed the same way, regardless of what
jis.
See Also
fast_neg_bin_with_var_cpp for the underlying
backend (no automatic rank-deficiency retry) and full model
documentation; fast_negbin_regression for the
estimate-only, similarly retry-hardened wrapper.
Examples
X = matrix(rnorm(100), 10, 10)
y = rpois(10, 2)
fast_negbin_regression_with_var(X, y)
Fast Ordinary Least Squares (OLS) Regression, Estimate-Only (C++ Backend)
Description
Solves the ordinary least squares normal equations
\hat\beta = (X^\top X)^{-1} X^\top y via Eigen's LDLT Cholesky
decomposition of X^\top X; if that decomposition fails (e.g. X
is rank-deficient), it falls back to a column-pivoted QR decomposition of
X directly (Eigen::ColPivHouseholderQR), which returns a
minimum-norm least-squares solution even when X is not full rank.
Estimate-only: no standard errors or covariance matrix are computed, only
the coefficient vector \hat\beta.
Usage
fast_ols_cpp(X, y, fixed_idx = NULL, fixed_values = NULL)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the (continuous) response variable. |
fixed_idx |
Optional integer vector of 1-indexed columns of |
fixed_values |
Optional numeric vector, parallel to |
Value
A list containing the following component:
- b
A numeric vector of the estimated regression coefficients
\hat\beta(with anyfixed_idxentries set tofixed_values); if the solve produces non-finite values, this is instead a vector ofNaN.
Fixed (offset) coefficients
fixed_idx (1-indexed columns
of X) and fixed_values optionally hold a subset of coefficients at
caller-supplied constant values rather than estimating them: those columns'
contribution X_{\mathrm{fixed}} \beta_{\mathrm{fixed}} is subtracted out of
y first, and only the remaining ("free") columns are fit by least squares;
the fixed coefficients are then copied back into \hat\beta unchanged. This is
used, e.g., to fit a model with a known/offset intercept without re-estimating it.
See Also
fast_ols_with_var_cpp for the variance-computing counterpart.
Fast Ordinary Least Squares (OLS) Regression with Variance (C++ Backend)
Description
As fast_ols_cpp, but additionally computes the classical OLS
variance estimate \hat\sigma^2 = \mathrm{SSE} / (n - p) (with
\mathrm{SSE} = y^\top y - \hat\beta^\top X^\top y on the fixed-parameter-adjusted
response, and p the number of free — non-fixed — columns) and the
sampling variance of two coefficients, \widehat{\mathrm{Var}}(\hat\beta_k) =
\hat\sigma^2 \, [(X^\top X)^{-1}]_{kk}, obtained from the same Cholesky (LDLT)
factorization used to solve for \hat\beta, without forming the full inverse
matrix. If the LDLT decomposition fails (rank-deficient X), this function
falls back to a QR solve exactly as fast_ols_cpp does, but in that case
no variance quantities are computed: ssq_b_j, ssq_b_2, and
XtX are omitted from the result and converged is FALSE.
Usage
fast_ols_with_var_cpp(X, y, j = 2L, fixed_idx = NULL, fixed_values = NULL)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the (continuous) response variable. |
j |
This function will compute the variance of the jth (1-indexed) coefficient estimator. Default is 2 (conventionally the treatment effect). |
fixed_idx |
Optional integer vector of 1-indexed columns of |
fixed_values |
Optional numeric vector, parallel to |
Value
A list containing the following components (the last four only present
when the LDLT solve succeeds):
- b
A numeric vector of the estimated regression coefficients
\hat\beta.- converged
TRUEif theLDLTsolve succeeded,FALSEif the QR fallback was used.- sigma2_hat
The estimated residual variance
\hat\sigma^2.- XtX
The (free-coefficient)
X^\top Xmatrix, expanded back to fullp \times pshape with zeros in the fixed-coefficient rows/columns.- ssq_b_j
The variance of the
j-th coefficient estimator,NAifjindexes a fixed coefficient.- ssq_b_2
The variance of the second coefficient estimator specifically (typically the treatment effect), regardless of what
jis; equal tossq_b_jwhenj = 2.NAif the second column is a fixed coefficient.
Fixed (offset) coefficients
fixed_idx (1-indexed columns
of X) and fixed_values optionally hold a subset of coefficients at
caller-supplied constant values rather than estimating them: those columns'
contribution X_{\mathrm{fixed}} \beta_{\mathrm{fixed}} is subtracted out of
y first, and only the remaining ("free") columns are fit by least squares;
the fixed coefficients are then copied back into \hat\beta unchanged. This is
used, e.g., to fit a model with a known/offset intercept without re-estimating it.
See Also
fast_ols_cpp for the estimate-only counterpart.
Fast Cumulative Ordinal Regression with a Cauchit Link (C++)
Description
Fits a cumulative-link ordinal regression model with the cauchit link
(the inverse CDF of the standard Cauchy distribution) via direct maximum
likelihood, jointly optimizing the category thresholds and regression
coefficients. y's distinct values (in sorted order, whatever their
original coding) are treated as K ordered categories; for observation
i in category k (k = 0, \ldots, K-1),
\Pr(Y_i \le k \mid x_i) = F(\alpha_k - x_i^\top \beta), \qquad
F(z) = \frac{1}{2} + \frac{\arctan(z)}{\pi},
with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing)
category thresholds and \beta the regression coefficients on X
(no separate intercept column is needed — the thresholds serve that role);
the category probability is the corresponding CDF difference,
\Pr(Y_i = k \mid x_i) = F(\alpha_k - x_i^\top\beta) - F(\alpha_{k-1} - x_i^\top\beta)
(with F(\alpha_{-1} - \cdot) := 0 and F(\alpha_{K-1} - \cdot) := 1
at the boundaries), each clamped below at 10^{-12} before taking logs
for numerical safety. This is a direct-likelihood analogue of ordinal
logistic regression with the logit link swapped for the heavier-tailed
Cauchy CDF, which is more robust to outlying/misclassified extreme
categories at the cost of less standard interpretability (no proportional-odds
log-odds-ratio reading of \beta). If y has fewer than 2 distinct
levels, an empty result list is returned (the model is degenerate).
Usage
fast_ordinal_cauchit_regression_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-06,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see Details). |
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
estimate_only |
If |
Value
A list with components b (the \beta coefficients), alpha
(the K-1 category thresholds), params (the concatenated
[\alpha, \beta] vector), n_params, converged, and iterations;
when estimate_only = FALSE (the default) and the fit converged, additionally
neg_loglik, observed_information/fisher_information/information
(all the same observed-information matrix), information_type (always
"observed"), vcov (the parameter covariance matrix), and ssq_b_j
(the variance of b[1], i.e. the coefficient on X's first column —
conventionally the treatment effect, since X carries no separate intercept
column here; the thresholds alpha play that role). Empty if y has
fewer than 2 distinct levels.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise, when
smart_cold_start = TRUE (the default), starting values come from an
OLS-based heuristic, and when FALSE, thresholds start at
\tan(\pi(k/K - 1/2)) (an inverse-cauchit spacing of the empirical
marginal category proportions) with \beta at zero. Optimization runs
via optimization_alg (default "lbfgs") for up to maxit
iterations or until the parameter/gradient change falls below tol;
warm_start_fisher_info, if supplied, seeds the first iteration's
curvature estimate.
See Also
fast_ordinal_cauchit_regression_with_var_cpp, which always computes
the variance quantities (equivalent to calling this function with
estimate_only = FALSE) and additionally guards against the degenerate/non-converged
case by returning NA placeholders instead of an empty list.
Fast Cumulative Ordinal Regression with a Cauchit Link, with Variance (C++)
Description
As fast_ordinal_cauchit_regression_cpp (see that page for the full
cumulative cauchit-link model, \Pr(Y_i \le k \mid x_i) = F(\alpha_k -
x_i^\top\beta)), but always fits with estimate_only = FALSE (equivalent
to calling that function with its default), so the observed information matrix
and variance-covariance matrix are always computed. It additionally guards the
degenerate case: if y has fewer than 2 distinct levels (so the underlying
fit is empty), this function returns list(b = NA, ssq_b_2 = NA) instead of
an empty list.
Usage
fast_ordinal_cauchit_regression_with_var_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
Value
A list with components b, alpha, params, n_params,
neg_loglik, converged, iterations,
observed_information/fisher_information/information (all the same
observed-information matrix), information_type ("observed"), vcov,
and ssq_b_2 (the variance of b[1], i.e. X's first-column coefficient
— conventionally the treatment effect). vcov is omitted (and ssq_b_2 is
NA) if the fit did not converge; list(b = NA, ssq_b_2 = NA) if y has
fewer than 2 distinct levels.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise, when
smart_cold_start = TRUE (the default), starting values come from an
OLS-based heuristic, and when FALSE, thresholds start at
\tan(\pi(k/K - 1/2)) (an inverse-cauchit spacing of the empirical
marginal category proportions) with \beta at zero. Optimization runs
via optimization_alg (default "lbfgs") for up to maxit
iterations or until the parameter/gradient change falls below tol;
warm_start_fisher_info, if supplied, seeds the first iteration's
curvature estimate.
See Also
fast_ordinal_cauchit_regression_cpp for the estimate-only-capable
variant and the full model documentation.
Fast Cumulative Ordinal Regression with a Complementary Log-Log Link (C++)
Description
Fits a cumulative-link ordinal regression model with the complementary
log-log ("cloglog") link via direct maximum likelihood, jointly optimizing the
category thresholds and regression coefficients. y's distinct values (in
sorted order, whatever their original coding) are treated as K ordered
categories; for observation i in category k (k = 0, \ldots, K-1),
\Pr(Y_i \le k \mid x_i) = F(\alpha_k + x_i^\top \beta), \qquad
F(z) = 1 - \exp(-\exp(z)),
with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category
thresholds and \beta the regression coefficients on X (no separate
intercept column is needed — the thresholds serve that role); the category
probability is the corresponding CDF difference,
\Pr(Y_i = k \mid x_i) = F(\alpha_k + x_i^\top\beta) - F(\alpha_{k-1} + x_i^\top\beta)
(with F(\alpha_{-1} + \cdot) := 0 and F(\alpha_{K-1} + \cdot) := 1 at the
boundaries), each clamped below at 10^{-12} before taking logs for numerical
safety. Unlike the symmetric logit/probit/cauchit links, cloglog is
asymmetric: it is the natural link for a grouped/discretized
proportional-hazards (continuation-ratio-free) survival model, appropriate when
category probabilities are skewed toward the lower categories. If y has
fewer than 2 distinct levels, an empty result list is returned (the model is
degenerate).
Usage
fast_ordinal_cloglog_regression_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-06,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see Details). |
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
estimate_only |
If |
Value
A list with components b (the \beta coefficients), alpha
(the K-1 category thresholds), params (the concatenated
[\alpha, \beta] vector), n_params, converged, and iterations;
when estimate_only = FALSE (the default) and the fit converged, additionally
neg_loglik, observed_information/fisher_information/information
(all the same observed-information matrix), information_type (always
"observed"), vcov (the parameter covariance matrix), and ssq_b_j
(the variance of b[1], i.e. the coefficient on X's first column —
conventionally the treatment effect, since X carries no separate intercept
column here; the thresholds alpha play that role). Empty if y has
fewer than 2 distinct levels.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise, when
smart_cold_start = TRUE (the default), starting values come from an
OLS-based heuristic, and when FALSE, thresholds start evenly spaced on
(-1, 1) at -1 + 2(k+1)/K with \beta at zero. Optimization runs
via optimization_alg (default "lbfgs") for up to maxit
iterations or until the parameter/gradient change falls below tol;
warm_start_fisher_info, if supplied, seeds the first iteration's
curvature estimate.
See Also
fast_ordinal_cloglog_regression_with_var_cpp, which always computes
the variance quantities (equivalent to calling this function with
estimate_only = FALSE) and additionally guards against the degenerate/non-converged
case by returning NA placeholders instead of an empty list.
Fast Cumulative Ordinal Regression with a Complementary Log-Log Link, with Variance (C++)
Description
As fast_ordinal_cloglog_regression_cpp (see that page for the full
cumulative cloglog-link model, \Pr(Y_i \le k \mid x_i) = F(\alpha_k +
x_i^\top\beta), F(z) = 1 - \exp(-\exp(z))), but always fits with
estimate_only = FALSE (equivalent to calling that function with its
default), so the observed information matrix and variance-covariance matrix are
always computed. It additionally guards the degenerate case: if y has
fewer than 2 distinct levels (so the underlying fit is empty), this function
returns list(b = NA, ssq_b_2 = NA) instead of an empty list.
Usage
fast_ordinal_cloglog_regression_with_var_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
Value
A list with components b, alpha, params, n_params,
neg_loglik, converged, iterations,
observed_information/fisher_information/information (all the same
observed-information matrix), information_type ("observed"), vcov,
and ssq_b_2 (the variance of b[1], i.e. X's first-column coefficient
— conventionally the treatment effect). vcov is omitted (and ssq_b_2 is
NA) if the fit did not converge; list(b = NA, ssq_b_2 = NA) if y has
fewer than 2 distinct levels.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise, when
smart_cold_start = TRUE (the default), starting values come from an
OLS-based heuristic, and when FALSE, thresholds start evenly spaced on
(-1, 1) at -1 + 2(k+1)/K with \beta at zero. Optimization runs
via optimization_alg (default "lbfgs") for up to maxit
iterations or until the parameter/gradient change falls below tol;
warm_start_fisher_info, if supplied, seeds the first iteration's
curvature estimate.
See Also
fast_ordinal_cloglog_regression_cpp for the estimate-only-capable
variant and the full model documentation.
Fast Ordinal Cumulative-Logit Random-Intercept GLMM via Gauss-Hermite Quadrature (C++)
Description
Fits a cumulative-logit ordinal mixed model with a single Gaussian random intercept per group (e.g. a matched pair or singleton from a KK-style matched design):
\mathrm{logit}\,\Pr(Y_{ij} \le k \mid u_i) = \alpha_k - x_{ij}^\top\beta - u_i,
\qquad u_i \sim N(0, \sigma^2),
for group i, member j, ordinal outcome y_{ij} \in \{1, \ldots, K\}, and
increasing cutpoints \alpha_1 < \cdots < \alpha_{K-1} (X carries no separate
intercept column — the cutpoints serve that role). The marginal likelihood for group
i integrates the random intercept out,
L_i(\theta) = \int \prod_{j \in i} \Pr(Y_{ij} = y_{ij} \mid u_i) \, \phi(u_i / \sigma) \, du_i,
approximated by n_gh-point Gauss-Hermite quadrature (substituting
u = \sqrt{2}\,\sigma\, z for quadrature node z), and optimized on the
log scale by directly maximizing \sum_i \log L_i(\theta) (via
optimize_fixed_likelihood, default optimization_alg = "lbfgs") over the
reparameterized vector [\alpha_1, \log(\alpha_2-\alpha_1), \ldots,
\log(\alpha_{K-1}-\alpha_{K-2}), \beta, \log\sigma] — cutpoints are recovered as
successive partial sums of \alpha_1 and the exponentiated log-differences, which
enforces \alpha_1 < \cdots < \alpha_{K-1} by construction rather than as a fitting
constraint. Rows are stably sorted by group_id inside the kernel, so
matched-group membership is invariant to input row order. Optimization uses
supplied/cold, moderate-variance, and near-zero-variance starts, retains the
smallest finite negative log-likelihood, and polishes that solution before
applying a projected-gradient convergence check. If finite multistart
L-BFGS stops on function decrease while its projected score remains above
max(1e-5, eps_g), the kernel performs a local damped-Newton polish
using its numerical Hessian. A Newton trial is retained only when its
parameters, objective, and gradient are finite and its objective does not
exceed the L-BFGS objective. At a valid lower log_sigma boundary, the
KKT-satisfied variance coordinate is excluded from the Newton system, so
fixed-effect convergence can be established without rejecting a
near-zero random-effect variance. log_sigma is evaluated
within [-\code{max\_abs\_log\_sigma},
\code{max\_abs\_log\_sigma}] with a quadratic penalty on excursions beyond
that interval whose analytic gradient matches the bounded objective;
variance_boundary_hit in the
return value flags whether the fitted log_sigma landed at that boundary (a sign
the random-intercept variance is being driven to (near-)zero or is unbounded, and that
ssq_b_T/fisher_information should be treated with caution). The Hessian
used for inference is a numerical (central finite-difference, step 10^{-4},
symmetrized) second derivative of the analytic gradient, not a closed-form expression.
At the valid near-zero variance boundary, treatment variance is computed
conditional on that boundary by excluding the nonregular variance-parameter
row and column.
Usage
fast_ordinal_glmm_cpp(
X,
y,
group_id,
K,
j_T,
smart_cold_start = TRUE,
estimate_only = FALSE,
n_gh = 20L,
max_abs_log_sigma = 8,
maxit = 300L,
eps_g = 1e-06,
warm_start_params = NULL,
warm_start_beta = NULL,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors, one row per observation (member-level, not group-level); no intercept column (see Details). |
y |
Integer vector of 1-indexed ordinal outcomes ( |
group_id |
Integer vector of group (e.g. matched-pair) identifiers, one per row of |
K |
The number of ordinal levels. |
j_T |
0-based column index of |
smart_cold_start |
Logical. If |
estimate_only |
If |
n_gh |
Number of Gauss-Hermite quadrature nodes used to integrate out the random intercept. |
max_abs_log_sigma |
Symmetric clamp bound for |
maxit |
Maximum number of optimizer iterations. |
eps_g |
Gradient-norm convergence tolerance. |
warm_start_params |
Optional starting values for the full reparameterized vector
|
warm_start_beta |
Optional starting values either for the full parameter vector (same
length as |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into the reparameterized parameter vector) to hold fixed. |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature matrix for the first optimizer iteration. |
Value
A list with components b (\hat\beta), alpha (the K-1
cutpoints, recovered from the reparameterization), params (the full fitted
reparameterized vector), log_sigma, ssq_b_T (variance of b[j_T],
NA unless estimate_only = FALSE and the fit converged and the resulting
information matrix inverts successfully), converged, neg_loglik,
fisher_information (the numerical Hessian, always returned), score
(the log-likelihood score at the returned parameters), gradient_norm,
newton_polish_attempted, newton_polish_accepted, and
newton_polish_iterations (diagnostics for the conditional
damped-Newton fallback), and variance_boundary_hit (TRUE/FALSE, or NA if the optimizer
itself threw an exception, in which case converged = FALSE and all other quantities
besides b/alpha/log_sigma are NA).
References
Pinheiro, J. C., and Bates, D. M. (1995). "Approximations to
the Log-Likelihood Function in the Nonlinear Mixed-Effects Model."
Journal of Computational and Graphical Statistics, 4(1), 12-35,
doi:10.1080/10618600.1995.10474663, for Gauss-Hermite quadrature as
an approximation to the random-effect marginal likelihood integral used
throughout this package's GLMM backends (fast_poisson_glmm_cpp,
fast_logistic_glmm_cpp, fast_weibull_frailty_cpp, and this
function).
Fast Cumulative Ordinal Regression with a Probit Link (C++)
Description
Fits a cumulative-link ordinal regression model with the probit link
(the standard normal CDF) via direct maximum likelihood, jointly optimizing the
category thresholds and regression coefficients. y's distinct values (in
sorted order, whatever their original coding) are treated as K ordered
categories; for observation i in category k (k = 0, \ldots, K-1),
\Pr(Y_i \le k \mid x_i) = \Phi(\alpha_k - x_i^\top \beta),
with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category
thresholds and \beta the regression coefficients on X (no separate
intercept column is needed — the thresholds serve that role); the category
probability is the corresponding CDF difference,
\Pr(Y_i = k \mid x_i) = \Phi(\alpha_k - x_i^\top\beta) - \Phi(\alpha_{k-1} - x_i^\top\beta)
(with \Phi(\alpha_{-1} - \cdot) := 0 and \Phi(\alpha_{K-1} - \cdot) := 1
at the boundaries), each clamped below at 10^{-12} before taking logs for
numerical safety. This is the ordinal generalization of probit regression, and
the thin-tailed counterpart to
fast_ordinal_cauchit_regression_cpp and the standard-logit ordinal
model. If y has fewer than 2 distinct levels, an empty result list is
returned (the model is degenerate).
Usage
fast_ordinal_probit_regression_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-06,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see Details). |
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
estimate_only |
If |
Value
A list with components b (the \beta coefficients), alpha
(the K-1 category thresholds), params (the concatenated
[\alpha, \beta] vector), n_params, converged, and iterations;
when estimate_only = FALSE (the default) and the fit converged, additionally
neg_loglik, observed_information/fisher_information/information
(all the same observed-information matrix), information_type (always
"observed"), vcov (the parameter covariance matrix), and ssq_b_j
(the variance of b[1], i.e. the coefficient on X's first column —
conventionally the treatment effect, since X carries no separate intercept
column here; the thresholds alpha play that role — set to NA if it comes
out non-finite or non-positive). Empty if y has fewer than 2 distinct levels.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise, when
smart_cold_start = TRUE (the default), starting values come from an
OLS-based heuristic, and when FALSE, thresholds start at
\Phi^{-1}(k/K) (inverse-normal spacing of the empirical marginal category
proportions) with \beta at zero. Optimization runs via
optimization_alg (default "lbfgs") for up to maxit
iterations or until the parameter/gradient change falls below tol;
warm_start_fisher_info, if supplied, seeds the first iteration's
curvature estimate.
See Also
fast_ordinal_probit_regression_with_var_cpp, which always computes
the variance quantities (equivalent to calling this function with
estimate_only = FALSE) and additionally guards against the degenerate/non-converged
case by returning NA placeholders instead of an empty list.
Fast Cumulative Ordinal Regression with a Probit Link, with Variance (C++)
Description
As fast_ordinal_probit_regression_cpp (see that page for the full
cumulative probit-link model, \Pr(Y_i \le k \mid x_i) = \Phi(\alpha_k -
x_i^\top\beta)), but always fits with estimate_only = FALSE (equivalent
to calling that function with its default), so the observed information matrix
and variance-covariance matrix are always computed. It additionally guards the
degenerate case: if y has fewer than 2 distinct levels (so the underlying
fit is empty), this function returns list(b = NA, ssq_b_2 = NA) instead of
an empty list.
Usage
fast_ordinal_probit_regression_with_var_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
optimization_alg = "lbfgs",
fixed_idx = NULL,
fixed_values = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
optimization_alg |
Optimization algorithm (default |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
Value
A list with components b, alpha, params, n_params,
neg_loglik, converged, iterations,
observed_information/fisher_information/information (all the same
observed-information matrix), information_type ("observed"), vcov,
and ssq_b_2 (the variance of b[1], i.e. X's first-column coefficient
— conventionally the treatment effect; NA if it comes out non-finite or
non-positive). vcov is omitted (and ssq_b_2 is NA) if the fit did
not converge; list(b = NA, ssq_b_2 = NA) if y has fewer than 2 distinct levels.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise, when
smart_cold_start = TRUE (the default), starting values come from an
OLS-based heuristic, and when FALSE, thresholds start at
\Phi^{-1}(k/K) (inverse-normal spacing of the empirical marginal category
proportions) with \beta at zero. Optimization runs via
optimization_alg (default "lbfgs") for up to maxit
iterations or until the parameter/gradient change falls below tol;
warm_start_fisher_info, if supplied, seeds the first iteration's
curvature estimate.
See Also
fast_ordinal_probit_regression_cpp for the estimate-only-capable
variant and the full model documentation.
Fast Cumulative Ordinal Regression with a Logit Link, i.e. Proportional-Odds Regression (C++)
Description
Fits the classical proportional-odds ordinal regression model (a cumulative-link
model with the logit link) via direct maximum likelihood, jointly
optimizing the category thresholds and regression coefficients. y's
distinct values (in sorted order, whatever their original coding) are treated as
K ordered categories; for observation i in category k
(k = 0, \ldots, K-1),
\mathrm{logit}\,\Pr(Y_i \le k \mid x_i) = \alpha_k - x_i^\top \beta,
with \alpha_0 < \alpha_1 < \cdots < \alpha_{K-2} the (increasing) category
thresholds and \beta the regression coefficients on X (no separate
intercept column is needed — the thresholds serve that role); the category
probability is the corresponding CDF difference, each clamped below at
10^{-12} before taking logs for numerical safety. The "proportional odds"
name reflects that \beta does not depend on k: the odds ratio
\exp(-\beta_j) for a unit increase in covariate j is the same across
every cumulative cutpoint. If y has fewer than 2 distinct levels, fitting
still proceeds with K = 1, n_alpha = 0 (degenerate: no thresholds
to estimate); unlike the cauchit/probit/cloglog variants in this package, this
function does not special-case that as an early return.
Usage
fast_ordinal_regression_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see Details). |
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
estimate_only |
If |
Value
A list with components b (the \beta coefficients), alpha
(the K-1 category thresholds), params (the concatenated
[\alpha, \beta] vector), n_params, converged, and iterations;
when estimate_only = FALSE (the default), additionally neg_loglik,
observed_information/fisher_information/information (all the same
observed-information matrix), information_type (always "observed"),
ssq_b_j (the variance of b[1], i.e. the coefficient on X's first
column — conventionally the treatment effect, since X carries no separate
intercept column here), and vcov (the full parameter covariance matrix) — the
latter two are NA/omitted (vcov becomes NULL) if the free-parameter
information matrix is not invertible.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise thresholds always start
evenly spaced on (-1, 1) at -1 + 2(k+1)/K, and \beta starts at
either zero, or (when smart_cold_start = TRUE, the default) an OLS fit of
the rank-rescaled response (y - 1)/(K - 1) on X — falling back
silently to zero if that OLS solve is not well-posed. When
smart_cold_start = TRUE and no warm_start_fisher_info is supplied,
the Hessian at the starting values is additionally used to seed the optimizer's
first-iteration curvature estimate. Optimization runs via optimization_alg
(default "lbfgs") for up to maxit iterations or until the
parameter/gradient change falls below tol.
See Also
fast_ordinal_regression_weighted_cpp for the observation-weighted
variant; fast_ordinal_regression_with_var_cpp, which additionally guards
the non-invertible case with explicit NA placeholders.
Fast Cumulative Ordinal Regression with a Logit Link, Weighted (C++)
Description
As fast_ordinal_regression_cpp (see that page for the full
proportional-odds model), but each observation's log-likelihood contribution is
multiplied by a nonnegative weights[i] (negative weights are clamped to
zero internally by the underlying weighted log-likelihood). Always fits with
estimate_only = FALSE (equivalent to that function's default), so the
observed information and, when invertible, the full variance-covariance matrix
are always computed.
Usage
fast_ordinal_regression_weighted_cpp(
X,
y,
weights,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
weights |
A numeric vector of observation weights, length |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
Value
A list with components b, alpha, params, n_params,
converged, iterations, neg_loglik,
observed_information/fisher_information/information (all the same
observed-information matrix), information_type ("observed"),
ssq_b_j (the variance of b[1]), and vcov — the latter two are
NA/omitted if the free-parameter information matrix is not invertible.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise thresholds always start
evenly spaced on (-1, 1) at -1 + 2(k+1)/K, and \beta starts at
either zero, or (when smart_cold_start = TRUE, the default) an OLS fit of
the rank-rescaled response (y - 1)/(K - 1) on X — falling back
silently to zero if that OLS solve is not well-posed. When
smart_cold_start = TRUE and no warm_start_fisher_info is supplied,
the Hessian at the starting values is additionally used to seed the optimizer's
first-iteration curvature estimate. Optimization runs via optimization_alg
(default "lbfgs") for up to maxit iterations or until the
parameter/gradient change falls below tol.
See Also
fast_ordinal_regression_cpp for the unweighted variant and the
full model documentation.
Fast Cumulative Ordinal Regression with a Logit Link, with Variance (C++)
Description
Identical to fast_ordinal_regression_cpp (see that page for the full
proportional-odds model) called with estimate_only = FALSE — this is simply
a convenience export that hardcodes that default rather than exposing the flag, so
the observed information and, when invertible, the full variance-covariance matrix
are always computed. Unlike the cauchit/probit/cloglog families' _with_var
variants, this function does not add any extra degenerate-case guarding
beyond what fast_ordinal_regression_cpp already does.
Usage
fast_ordinal_regression_with_var_cpp(
X,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-06,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
warm_start_params |
Optional starting values for |
smart_cold_start |
Logical. If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional 1-indexed positions (into |
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
Value
A list with components b, alpha, params, n_params,
converged, iterations, neg_loglik,
observed_information/fisher_information/information (all the same
observed-information matrix), information_type ("observed"),
ssq_b_j (the variance of b[1], i.e. the coefficient on X's first
column — conventionally the treatment effect), and vcov — the latter two are
NA/omitted if the free-parameter information matrix is not invertible.
Fixed parameters, warm starts, and optimization
fixed_idx (1-indexed into the combined [\alpha, \beta] parameter
vector, thresholds first) and fixed_values optionally hold a subset of
parameters at caller-supplied constant values rather than estimating them.
warm_start_params supplies starting values for [\alpha, \beta]
directly (skipping smart_cold_start); otherwise thresholds always start
evenly spaced on (-1, 1) at -1 + 2(k+1)/K, and \beta starts at
either zero, or (when smart_cold_start = TRUE, the default) an OLS fit of
the rank-rescaled response (y - 1)/(K - 1) on X — falling back
silently to zero if that OLS solve is not well-posed. When
smart_cold_start = TRUE and no warm_start_fisher_info is supplied,
the Hessian at the starting values is additionally used to seed the optimizer's
first-iteration curvature estimate. Optimization runs via optimization_alg
(default "lbfgs") for up to maxit iterations or until the
parameter/gradient change falls below tol.
See Also
fast_ordinal_regression_cpp for the estimate-only-capable variant
and the full model documentation; fast_ordinal_regression_weighted_cpp
for the observation-weighted variant.
Fast Poisson Regression, Estimate-Only (C++ Backend)
Description
Fits a Poisson regression with the canonical log link, Y_i \sim
\mathrm{Poisson}(\mu_i), \mu_i = \exp(x_i^\top\beta) (with \eta_i =
x_i^\top\beta clamped above at 700 before exponentiating, to avoid overflow),
by maximum likelihood. By default (optimization_alg = "irls"), fitting
uses iteratively reweighted least squares: at each iteration the Fisher-scoring
(Poisson canonical-link, so Fisher = observed) system X^\top W X \, \delta =
X^\top(y - \mu) is solved via Eigen::LDLT, with a backtracking
step-halving line search (up to 10 halvings) accepting the step only if it does
not increase the negative log-likelihood; convergence is declared when the score
norm falls below tol. Passing optimization_alg = "lbfgs" or
"newton_raphson" instead routes through the generic likelihood optimizer
(.normalize_optimizer_algorithm) on the raw (non-IRLS)
negative log-likelihood/gradient/Hessian.
Usage
fast_poisson_regression_cpp(
X,
y,
warm_start_beta = NULL,
smart_cold_start = FALSE,
maxit = 100L,
tol = 1e-8,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be nonnegative-integer counts. |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations. Defaults to 100. |
tol |
Convergence tolerance. Defaults to 1e-8. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by |
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Accepted but unused; see Details. |
warm_start_fisher_info |
Optional initial curvature (information) matrix to warm-start the first iteration. |
estimate_only |
If |
Value
A list containing the following components:
- b
A numeric vector of the estimated Poisson regression coefficients
\hat\beta.- mu
(omitted if
estimate_only = TRUE) The fitted means\hat\mu_i.- XtWX, fisher_information
(omitted if
estimate_only = TRUE) Two aliases for the same curvature matrixX^\top W X(W = \mathrm{diag}(\hat\mu_i)) at the fitted coefficients — the Fisher information, which for the canonical log link coincides with the observed information.- score
(omitted if
estimate_only = TRUE) The score vectorX^\top(y - \hat\mu)at the fitted coefficients.- neg_ll
(omitted if
estimate_only = TRUE) The negative log-likelihood at the fitted coefficients.- converged
A logical value indicating whether the final gradient norm was below
tol(gradient_norm < tol); uniform across the"irls"/"lbfgs"/"newton_raphson"optimizers.- num_iter
The number of iterations performed.
- hit_iteration_cap
A logical value, mutually exclusive with
converged:TRUEiff the optimizer exhaustedmaxititerations without meeting the gradient-norm convergence criterion.- gradient_norm
The norm of the score vector at the returned coefficients.
Fixed parameters, warm starts
fixed_idx and fixed_values optionally hold a subset of coefficients
fixed at caller-supplied constant values (their contribution is folded into the
linear predictor as an offset) rather than estimated. warm_start_beta
supplies starting coefficients directly; otherwise, if smart_cold_start =
TRUE, a Poisson-specific heuristic start is used, and if
warm_start_fisher_info is also supplied (or, absent that, when
smart_cold_start = TRUE), it seeds the curvature estimate used for the
very first IRLS step (or first quasi-Newton step, for the non-IRLS algorithms).
warm_start_weights is accepted for interface parity with sibling
functions but is not consulted anywhere in this function's fitting
logic.
See Also
fast_poisson_regression_weighted_cpp for the observation-weighted
variant; fast_poisson_regression_with_var_cpp for the variance-computing
variant; fast_quasipoisson_regression_with_var_cpp for the
overdispersion-corrected variant.
Fast Weighted Poisson Regression (C++ Backend)
Description
Fits the same Poisson log-link model as fast_poisson_regression_cpp
(see that page for the full model and optimizer contract), with each
observation's contribution to the log-likelihood, score, and IRLS working
weights multiplied by a row weight weights[i]. Setting all weights to 1
recovers fast_poisson_regression_cpp exactly. Always fits with
estimate_only = FALSE (there is no flag to skip the post-fit mu/
XtWX/score computation for this variant).
Usage
fast_poisson_regression_weighted_cpp(
X,
y,
weights,
warm_start_beta = NULL,
smart_cold_start = FALSE,
maxit = 100L,
tol = 1e-8,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be nonnegative-integer counts. |
weights |
A numeric vector of nonnegative weights, one per observation. |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations. Defaults to 100. |
tol |
Convergence tolerance. Defaults to 1e-8. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by |
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Accepted but unused; see |
warm_start_fisher_info |
Optional initial curvature (information) matrix to warm-start the first iteration. |
Value
A list containing the following components:
- b
A numeric vector of the estimated Poisson regression coefficients
\hat\beta.- mu
The fitted means
\hat\mu_i.- XtWX, fisher_information
Two aliases for the same (weighted) curvature matrix
X^\top W X,W = \mathrm{diag}(\code{weights}_i \hat\mu_i), at the fitted coefficients.- score
The weighted score vector at the fitted coefficients.
- neg_ll
The weighted negative log-likelihood at the fitted coefficients.
- converged
A logical value indicating whether the algorithm converged.
- iterations
The number of iterations performed.
- gradient_norm
The norm of the score vector at convergence.
See Also
fast_poisson_regression_cpp for the unweighted variant and full
model documentation.
Fast Poisson Regression with Variance Calculation (C++ Backend)
Description
Fits the same Poisson log-link model as fast_poisson_regression_cpp
(see that page for the full model and optimizer contract; always with
estimate_only = FALSE), and additionally inverts the fitted Fisher
information matrix (Eigen::LDLT on the free-coefficient submatrix, via
compute_diagonal_inverse_entry) to report the variance of two
coefficients.
Usage
fast_poisson_regression_with_var_cpp(
X,
y,
j = 2L,
warm_start_beta = NULL,
smart_cold_start = FALSE,
maxit = 100L,
tol = 1e-8,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be nonnegative-integer counts. |
j |
The 1-indexed coefficient whose variance to compute in |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations. Defaults to 100. |
tol |
Convergence tolerance. Defaults to 1e-8. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by |
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Accepted but unused; see |
warm_start_fisher_info |
Optional initial curvature (information) matrix to warm-start the first iteration. |
Value
A list containing the following components:
- b, params
A numeric vector of the estimated Poisson regression coefficients
\hat\beta(both aliases of the same vector).- ssq_b_j
The variance of the
j-th coefficient estimator;NAifjindexes a fixed coefficient.- ssq_b_2
The variance of the second coefficient estimator specifically (typically the treatment effect), regardless of what
jis; equal tossq_b_jwhenj = 2.NAif the second column is a fixed coefficient.- mu
The fitted means
\hat\mu_i.- converged
A logical value indicating whether the final gradient norm was below
tol(gradient_norm < tol); uniform across the"irls"/"lbfgs"/"newton_raphson"optimizers.- num_iter
The number of iterations performed.
- score
The score vector at the fitted coefficients.
- observed_information, fisher_information, information
Three aliases for the same
X^\top W Xcurvature matrix (information_typeis always"fisher").- hessian
The negative of that same matrix (the actual Hessian of the log-likelihood).
- neg_loglik, neg_ll
The negative log-likelihood at the fitted coefficients (two aliases).
- loglik
The log-likelihood (
-neg_ll), orNAifneg_llis non-finite.- hit_iteration_cap
A logical value, mutually exclusive with
converged:TRUEiff the optimizer exhaustedmaxititerations without meeting the gradient-norm convergence criterion.- gradient_norm
The norm of the score vector at the returned coefficients.
See Also
fast_poisson_regression_cpp for the estimate-only variant and full
model documentation; fast_quasipoisson_regression_with_var_cpp for the
overdispersion-corrected variant.
Fast Probit Regression, Estimate-Capable (C++)
Description
Fits binary probit regression, Y_i \sim \mathrm{Bernoulli}(\Phi(\eta_i)),
\eta_i = x_i^\top\beta, by maximum likelihood, using a numerically stable
log-scale evaluation of \Phi (log_pnorm_lower/log_pnorm_upper,
matching pnorm's log.p = TRUE for |\eta| < 6 via a
erfc-based identity, and falling back to a wider-range series approximation
beyond that). By default (optimization_alg = "irls"), fitting uses
iteratively reweighted least squares with working weights w_i =
\phi(\eta_i)^2 / \max(\Phi(\eta_i)(1-\Phi(\eta_i)), 10^{-15}) (the standard probit
Fisher-scoring weight) and generalized residual r_i = y_i\,\phi(\eta_i)/\Phi(\eta_i)
- (1-y_i)\,\phi(\eta_i)/(1-\Phi(\eta_i)): each iteration solves X^\top W X\,
\delta = X^\top r via Eigen::LDLT and takes the full Newton step (no
step-halving line search), declaring convergence when either the score norm or the
step norm falls below tol. Any optimization_alg value other than
"lbfgs" runs this IRLS path; optimization_alg = "lbfgs" instead
minimizes the exact negative log-likelihood directly via a bespoke L-BFGS driver
with backtracking strong-Wolfe line search (mirroring RcppNumerical's
optim_lbfgs defaults), bypassing IRLS entirely — in that path,
warm_start_fisher_info is not consulted.
Usage
fast_probit_regression_cpp(
X,
y,
warm_start_beta = NULL,
smart_cold_start = TRUE,
maxit = 100L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (including an intercept column, if desired). |
y |
A numeric vector of binary responses (0/1). |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations (IRLS path only). |
tol |
Convergence tolerance. |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
optimization_alg |
Optimization algorithm: any value other than
|
warm_start_weights |
Accepted but unused; see Details. |
warm_start_fisher_info |
Optional initial curvature matrix for the first IRLS iteration (IRLS path only). |
estimate_only |
If |
Value
If estimate_only = TRUE: a list with b, converged,
iterations. Otherwise: a list additionally containing w (the final
probit IRLS working weights, evaluated at the fitted \hat\beta regardless of
which optimization_alg was used), fisher_information (X^\top W X
at those weights), score, and neg_ll (the negative log-likelihood).
Fixed parameters, warm starts
fixed_idx and fixed_values optionally hold a subset of coefficients
fixed at caller-supplied constant values (folded into the linear predictor as an
offset) rather than estimated. warm_start_beta supplies starting
coefficients directly; otherwise, if smart_cold_start = TRUE (the default),
an OLS fit of the probit-transformed response \Phi^{-1}((y+0.5)/2) on
X seeds the start. warm_start_fisher_info, if supplied, seeds the
curvature matrix used for the IRLS path's very first iteration only (see above for
the "lbfgs" exception). warm_start_weights is accepted for interface
parity with sibling functions but is not consulted anywhere in this
function's fitting logic.
See Also
fast_probit_regression_weighted_cpp() for the observation-weighted
variant; fast_probit_regression_with_var_cpp for the variance-computing
variant.
Export of C++ function fast_probit_regression_with_var_cpp
Description
Fits the same probit model as fast_probit_regression_cpp (see that
page for the full model and optimizer contract; always with estimate_only =
FALSE, and maxit/tol hardcoded to 100/10^{-8} — no override
arguments here), and additionally inverts the fitted Fisher information matrix
(compute_diagonal_inverse_entry on the free-coefficient submatrix) to
report the variance of two coefficients.
Usage
fast_probit_regression_with_var_cpp(
X,
y,
j = 2L,
warm_start_beta = NULL,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors (including an intercept column, if desired). |
y |
A numeric vector of binary responses (0/1). |
j |
The 1-indexed coefficient whose variance to compute in |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
optimization_alg |
Optimization algorithm: any value other than
|
warm_start_weights |
Accepted but unused; see |
warm_start_fisher_info |
Optional initial curvature matrix for the first IRLS iteration (IRLS path only). |
Value
A list with components b, params (the fitted coefficients \hat\beta,
two aliases of the same vector), ssq_b_j (the variance of the j-th
coefficient, NA if j indexes a fixed coefficient), ssq_b_2 (the
variance of the second coefficient specifically, regardless of j; NA if
the second column is fixed), score, observed_information /
fisher_information / information (three aliases for the same X^\top
W X curvature matrix; information_type is always "fisher"),
hessian (the negative of that same matrix), neg_loglik/neg_ll
(two aliases for the negative log-likelihood), loglik (-neg_ll, or
NA if non-finite), converged, and iterations.
See Also
fast_probit_regression_cpp for the estimate-only-capable variant
and full model documentation.
Fast Standard Normal Quantile Function, Vectorized (C++ Backend)
Description
Computes \Phi^{-1}(p), the standard normal quantile (inverse CDF),
elementwise over p, via Peter Acklam's rational (minimax) approximation,
accurate to within roughly 1.2 \times 10^{-9} relative error over the
representable range of p — faster than base R's
qnorm while matching it to that approximation precision.
Used as the cold-start heuristic in several of this package's ordinal- and
binary-response regression fitters (e.g. probit-family threshold
initialization) wherever an approximate normal quantile is needed on a hot
path, and exported standalone for the same reason as
fast_digamma_vec_cpp and friends: to let performance-sensitive R
or Python callers bypass qnorm's per-call dispatch overhead when
evaluating many quantiles at once. Benchmarked at roughly 2.33x the
speed of base R's vectorized qnorm() on this package's benchmark suite;
see the "Utility
/ Math Kernel Performance" section of the benchmark report for the full
methodology and current measured multiple.
Usage
fast_qnorm_vec_cpp(p)
Arguments
p |
Numeric vector of probabilities in |
Value
A numeric vector of standard normal quantiles, the same length as p.
See Also
fast_log_pnorm_vec_cpp and fast_log_dnorm_vec_cpp
for the corresponding forward (CDF/density) kernels.
Fast Quasi-Poisson Regression with Variance Calculation (C++ Backend)
Description
Fits the same Poisson log-link mean model as fast_poisson_regression_cpp
(see that page for the full model and optimizer contract; point estimates \hat\beta
are identical to what that function would return), but instead of the plain Poisson
Fisher-information-based variance, scales it by an estimated overdispersion
parameter to obtain quasi-likelihood standard errors that are robust to
variance-mean deviations from the strict Poisson assumption
\mathrm{Var}(Y_i) = \mu_i. The dispersion is the Pearson-statistic-based
moment estimator,
\hat\phi = \frac{1}{n-p}\sum_{i=1}^n \frac{(y_i - \hat\mu_i)^2}{\hat\mu_i},
computed only when the residual degrees of freedom n - p > 0; the reported
coefficient variances are then \widehat{\mathrm{Var}}(\hat\beta_k) = \hat\phi
\, [(X^\top \hat W X)^{-1}]_{kk} (the ordinary Poisson Fisher information scaled by
\hat\phi), matching the standard quasi-Poisson GLM correction (as in
stats::glm(family = quasipoisson())). If n \le p, or \hat\phi
comes out non-finite or non-positive, ssq_b_j/ssq_b_2/dispersion
are left at their NA defaults (point estimates b and mu are
still returned).
Usage
fast_quasipoisson_regression_with_var_cpp(
X,
y,
j = 2L,
warm_start_beta = NULL,
smart_cold_start = FALSE,
maxit = 100L,
tol = 1e-8,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "irls",
warm_start_weights = NULL,
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
y |
A numeric vector of the response variable, expected to be nonnegative-integer counts. |
j |
The 1-indexed coefficient whose variance to compute in |
warm_start_beta |
Optional starting values for coefficients |
smart_cold_start |
Logical. If |
maxit |
Maximum number of iterations. Defaults to 100. |
tol |
Convergence tolerance. Defaults to 1e-8. |
fixed_idx |
Optional integer indices of coefficients to hold fixed rather than estimate. |
fixed_values |
Optional values to fix the parameters named by |
optimization_alg |
Optimization algorithm: |
warm_start_weights |
Accepted but unused; see |
warm_start_fisher_info |
Optional initial curvature (information) matrix to warm-start the first iteration. |
Value
A list containing the following components:
- b
A numeric vector of the estimated Poisson regression coefficients
\hat\beta(point estimates, unaffected by the dispersion correction).- ssq_b_j
The dispersion-scaled variance of the
j-th coefficient estimator;NAifjindexes a fixed coefficient or the dispersion estimate is unusable.- ssq_b_2
The dispersion-scaled variance of the second coefficient estimator specifically (typically the treatment effect), regardless of what
jis; equal tossq_b_jwhenj = 2.- dispersion
The estimated Pearson-based overdispersion parameter
\hat\phi, orNAif the residual degrees of freedom are not positive.- mu
The fitted means
\hat\mu_i.- converged
A logical value indicating whether the algorithm converged.
- iterations
The number of iterations performed.
- gradient_norm
The norm of the score vector at convergence.
See Also
fast_poisson_regression_with_var_cpp for the plain
(non-overdispersion-corrected) variance variant; fast_poisson_regression_cpp
for the estimate-only variant and full mean-model documentation.
Fast Robust (M/MM-Estimator) Linear Regression (C++)
Description
Fits a robust linear regression by iteratively reweighted least squares (IRLS),
minimizing \sum_i \rho(r_i / \hat\sigma) for residuals r_i = y_i -
x_i^\top\beta and a fixed robustness scale \hat\sigma, rather than
ordinary least squares' \sum_i r_i^2. method = "M" uses Huber's
weight function w(u) = 1 for |u| \le c and w(u) = c/|u|
otherwise (c, default 1.345, tuned for 95% efficiency under normality);
any other value of method (including the default, "MM") uses
Tukey's bisquare weight w(u) = (1 - (u/c_b)^2)^2 for |u| \le c_b
and 0 otherwise, with c_b = 4.685 hardcoded (not settable through
this exported wrapper, though the internal fitter accepts it). The scale
\hat\sigma is fixed once at the start as the normalized median absolute
deviation of the OLS residuals, \hat\sigma = \mathrm{median}(|r_i|) /
0.6745 (the internal fitter also accepts a caller-supplied fixed scale, but
this wrapper always estimates it). Each IRLS iteration re-weights and re-solves
the weighted normal equations X^\top W X\, \beta = X^\top W y via
Eigen::LDLT; convergence is declared when the relative change in
\beta falls below tol. Column 1 of a real design typically holds
the intercept, but no columns are treated specially except via fixed_idx.
Usage
fast_robust_regression_cpp(
X,
y,
warm_start_beta = NULL,
smart_cold_start = TRUE,
method = "MM",
j = 2L,
c = 1.345,
maxit = 50L,
tol = 1e-07,
fixed_idx = NULL,
fixed_values = NULL,
warm_start_weights = NULL,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of responses. |
warm_start_beta |
Optional starting values for coefficients. If provided, |
smart_cold_start |
Logical. If |
method |
Robust estimation method: |
j |
1-based index of the coefficient whose asymptotic variance to return in |
c |
Huber tuning constant (default 1.345; only used when |
maxit |
Maximum number of IRLS iterations. |
tol |
Relative parameter-change convergence tolerance. |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
warm_start_weights |
Optional initial working weights for the first IRLS iteration. |
warm_start_fisher_info |
Optional initial curvature ( |
estimate_only |
If |
Value
If estimate_only = TRUE: a list with coefficients, scale
(the fixed MAD-based robustness scale \hat\sigma), converged, iterations.
Otherwise, additionally: ssq_b_j and fisher_information (the final IRLS
X^\top W X curvature matrix). ssq_b_j is computed only if the fit converged
or ran the full maxit iterations, as the standard M-estimator asymptotic variance
\widehat{\mathrm{Var}}(\hat\beta_j) = \left(\frac{n}{n-p}\right) \frac{\sum_i
\psi(r_i)^2}{n\,\bar\psi'^2} \, [(X^\top X)^{-1}]_{jj}, where \psi is the
derivative of \rho (i.e. \psi(r) = w(r/\hat\sigma)\,r) and \bar\psi' is
the mean of \psi' across observations, matching the classical Huber (1981)
sandwich-free M-estimator variance formula; NA if j indexes a fixed
coefficient or the fit neither converged nor exhausted maxit.
Fixed parameters, warm starts
fixed_idx and fixed_values optionally hold a subset of coefficients
fixed at caller-supplied constant values (subtracted out of y as an
offset) rather than estimated. warm_start_beta supplies a starting
coefficient vector directly; otherwise, if smart_cold_start = TRUE (the
default), an ordinary QR least-squares fit seeds the start (and, when variance
will later be requested via j, also caches the QR-based [(X^\top
X)^{-1}]_{jj} entry for reuse in the variance formula below). warm_start_weights
seeds the IRLS weights for the first iteration only (skipping that iteration's
Huber/bisquare weight computation); warm_start_fisher_info similarly seeds
the first iteration's X^\top W X curvature matrix.
Fast Stereotype (Reduced-Rank Multinomial) Logistic Regression (C++)
Description
Fits Anderson's stereotype logit model — a reduced-rank multinomial
logit for a categorical (nominal or ordinal) response with K distinct
observed levels, using a single linear predictor \eta_i =
x_i^\top\beta scaled by a category-specific "score" \phi_k \in [0, 1]:
\Pr(Y_i = k \mid x_i) = \frac{\exp(\alpha_k + \phi_k \eta_i)}
{\sum_{l=1}^K \exp(\alpha_l + \phi_l \eta_i)}, \qquad \alpha_1 := 0,\ \phi_1 := 0,\ \phi_K := 1,
with free intercepts \alpha_2, \ldots, \alpha_K and free interior scores
\phi_2, \ldots, \phi_{K-1} reparameterized via unconstrained
\gamma_1, \ldots, \gamma_{K-2} as cumulative softmax-style partial sums,
\phi_{j} = \left(\sum_{r \le j-2} e^{\gamma_r}\right) \big/ \left(1 +
\sum_r e^{\gamma_r}\right) for j = 2, \ldots, K-1, which guarantees
0 = \phi_1 \le \phi_2 \le \cdots \le \phi_{K-1} \le \phi_K = 1 without an
explicit constraint. A single \hat\beta therefore governs the covariate
effect for every category, with the fitted \hat\phi_k determining how much
of that effect applies to category k — collapsing categories with similar
fitted scores are "stereotyped" together, which is the model's namesake use case
(a parsimony-inducing alternative to full multinomial or ordinal cumulative-link
models when categories are not clearly ordered but the covariate effect is
plausibly one-dimensional). K = 2 reduces exactly to ordinary binary
logistic regression (\phi_2 = 1 by construction, no \gamma
parameters). At least 2 distinct observed outcome categories are required; fewer
throws an error. Fitting optimizes the joint parameter vector [\alpha_2,
\ldots, \alpha_K, \beta, \gamma_1, \ldots, \gamma_{K-2}] via
optimization_alg (default "newton_raphson"), using this model's
analytic gradient and Hessian.
Usage
fast_stereotype_logit_cpp(
X,
y,
maxit = 100L,
tol = 1e-08,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "newton_raphson",
warm_start_fisher_info = NULL,
warm_start_params = NULL,
warm_start_beta = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; the
category intercepts |
y |
A numeric vector of categorical (nominal or ordinal) responses; only
the set of distinct values matters (mapped to |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
smart_cold_start |
Present for interface parity with sibling functions but
currently has no effect: starting values always come from
|
fixed_idx |
Optional 1-indexed positions (into the joint parameter vector,
|
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature matrix for the first optimizer iteration. |
warm_start_params |
Optional starting values for the full joint parameter vector.
Takes precedence over |
warm_start_beta |
Optional starting values for |
estimate_only |
If |
Value
A list with components b (\hat\beta), alpha (the K-1
free intercepts), scores_raw (the raw \hat\gamma reparameterization
parameters, length \max(0, K-2); recover \hat\phi from these via the
cumulative-softmax formula above), params (the full fitted joint parameter
vector), neg_loglik, converged, and, unless estimate_only = TRUE,
fisher_information (the negative Hessian of the log-likelihood at the fitted
parameters — despite the name, this is the observed, not expected,
information).
See Also
fast_stereotype_logit_with_var_cpp for the variance-computing
variant.
Fast Stereotype (Reduced-Rank Multinomial) Logistic Regression with Variance (C++)
Description
As fast_stereotype_logit_cpp (see that page for the full stereotype
logit model), but always computes the observed information and the variance of
\hat\beta_1 (the first, and typically only meaningfully identified,
regression coefficient — conventionally the treatment effect). The primary
variance estimate is [(-H)^{-1}]_{\beta_1\beta_1} from the observed
information (negative Hessian) at the fitted parameters. If \beta_1 is not
held fixed (via fixed_idx) and that entry comes out non-finite (e.g. the
information matrix is singular), this function falls back to a
profile-likelihood variance: it re-optimizes all nuisance parameters
(everything except \beta_1) at \hat\beta_1, \hat\beta_1 \pm h
(h = \max(10^{-4}, 10^{-3}(|\hat\beta_1| + 1))), takes the central
second-difference of the resulting profile log-likelihood to approximate the
profile information I(\hat\beta_1), and reports 1/I(\hat\beta_1) if
that comes out finite and positive (otherwise the variance remains NA).
vcov is never populated (always NULL/missing) — only the single
\beta_1 variance is available from this function.
Usage
fast_stereotype_logit_with_var_cpp(
X,
y,
maxit = 100L,
tol = 1e-08,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "newton_raphson",
warm_start_fisher_info = NULL,
warm_start_params = NULL,
warm_start_beta = NULL,
estimate_only = FALSE
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters, not their numeric coding or order. |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
smart_cold_start |
Present for interface parity but currently has no effect;
see |
fixed_idx |
Optional 1-indexed positions (into the joint parameter vector,
|
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature matrix for the first optimizer iteration. |
warm_start_params |
Optional starting values for the full joint parameter vector.
Takes precedence over |
warm_start_beta |
Optional starting values for |
estimate_only |
Accepted for interface parity but ignored: this function
always computes the observed information and |
Value
A list with components b (\hat\beta), alpha (the K-1
free intercepts), params (the full fitted joint parameter vector),
ssq_b_1 and ssq_b_j (identical aliases for the variance of
\hat\beta_1, computed as described above; NA if unavailable),
vcov (always missing/NULL), converged, and
fisher_information (the observed information, i.e. negative Hessian, at the
fitted parameters).
See Also
fast_stereotype_logit_cpp for the estimate-only-capable variant
and the full model documentation.
Stereotype Logit Profile Log-Likelihood for a Fixed Treatment Coefficient (C++)
Description
Computes the profile log-likelihood of the
fast_stereotype_logit_cpp stereotype logit model (see that page for
the full model) at a caller-fixed value of \beta_1 (the first regression
coefficient, conventionally the treatment effect): all other parameters
(intercepts \alpha, the remaining \beta columns if p > 1, and
the score reparameterization \gamma) are re-optimized by damped Newton's
method (up to 50 iterations, gradient-norm tolerance 10^{-8}, backtracking
step-halving line search — both hardcoded and not controlled by the
maxit/tol arguments below, which are accepted but currently
unused) to maximize the log-likelihood conditional on \beta_1, and the
resulting maximized log-likelihood is returned. This is the building block used
by fast_stereotype_logit_with_var_cpp's profile-likelihood
variance fallback (finite-differencing this function's output in
\beta_1), and is exported standalone for constructing profile-likelihood
confidence intervals or diagnostic profile plots directly.
Usage
fast_stereotype_profile_loglik_cpp(
X,
y,
beta_fixed,
maxit = 100L,
tol = 1e-08,
warm_start_params = NULL,
warm_start_beta = NULL
)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters, not their numeric coding or order. |
beta_fixed |
The fixed value at which to profile |
maxit |
Accepted but currently unused; see Details. |
tol |
Accepted but currently unused; see Details. |
warm_start_params |
Optional starting values for the full joint parameter vector
(used to initialize the nuisance-parameter optimization). Takes precedence over
|
warm_start_beta |
Optional starting values for |
Value
The maximized profile log-likelihood at beta_fixed, a single number.
See Also
fast_stereotype_logit_with_var_cpp, whose profile-likelihood
variance fallback calls this function three times per fallback invocation.
Fast Trigamma Function, Vectorized (C++ Backend)
Description
Computes \psi'(x), the trigamma function (the second derivative of
\log\Gamma(x), i.e. the derivative of the digamma function
fast_digamma_vec_cpp), elementwise over x, via an
asymptotic series expansion combined with the recurrence relation \psi'(x)
= \psi'(x+1) + 1/x^2 (shifting small arguments up into the expansion's
accurate range before applying it) — faster than base R's
trigamma while matching it to within the approximation's own
precision. Used wherever this package's likelihood kernels need the variance of
a log-Gamma-based sufficient statistic or a Fisher-information second
derivative involving \log\Gamma (e.g. negative-binomial dispersion-parameter
curvature), and exported standalone for the same reason as
fast_digamma_vec_cpp and friends. Benchmarked at roughly
19.3x the speed of base R's vectorized trigamma() on this
package's benchmark suite; see the
"Utility
/ Math Kernel Performance" section of the benchmark report for the full
methodology and current measured multiple.
Usage
fast_trigamma_vec_cpp(x)
Arguments
x |
Numeric vector of arguments (per the trigamma function's domain, should
not be a non-positive integer, where |
Value
A numeric vector of \psi'(x) values, the same length as x.
References
Trigamma
function for orientation. Analogous Python API:
SciPy
polygamma(1, x).
See Also
fast_digamma_vec_cpp for the corresponding first-derivative
kernel this function's recurrence builds on.
Fast Weibull AFT Regression (R Wrapper: Rcpp Backend or survival)
Description
Fits the Weibull accelerated failure time model documented in full at
fast_weibull_regression_general_cpp (\log T_i = \eta_i +
\sigma W_i, \eta_i = x_i^\top\beta, right-censoring only), via either
that C++ backend (use_rcpp = TRUE, the default) or
survreg with dist = "weibull"
(use_rcpp = FALSE) as a fallback/cross-check implementation.
Usage
fast_weibull_regression(
y,
dead,
X,
use_rcpp = TRUE,
estimate_only = FALSE,
optimization_alg = "lbfgs",
warm_start_params = NULL,
warm_start_fisher_info = NULL
)
Arguments
y |
Observed survival/censoring times (must be positive). |
dead |
Event indicator: 1 for an exactly observed event, 0 for
right-censored (survival known only to exceed |
X |
A numeric matrix of predictor variables. It is assumed that an intercept column
(e.g., a column of ones) is already included in |
use_rcpp |
Logical. If |
estimate_only |
Logical. If |
optimization_alg |
Optimization algorithm: |
warm_start_params |
Optional starting values for |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed
information) matrix. Only has an effect when |
Details
When use_rcpp = TRUE, an intercept column is prepended to
X automatically if not already present (detected as a first column
of all 1s), fitting always starts from a zero cold start
(smart_cold_start = FALSE is hardcoded, regardless of whether
warm_start_params is supplied), and a non-converged C++ fit is
escalated to an R-level stop() rather than returned silently.
When use_rcpp = FALSE, any existing intercept column is stripped and
survreg is left to add its own; remaining covariate columns are first
passed through drop_linearly_dependent_cols to remove
collinear columns before fitting (silently — no error or warning is raised
for dropped columns). estimate_only, optimization_alg,
warm_start_params, and warm_start_fisher_info have no
effect on this path — survreg always computes the full
variance-covariance matrix, and std_errs (from sqrt(diag(vcov)))
is included only in this path's return value, not the Rcpp path's. Both
non-finite coefficients and (unless estimate_only = TRUE, which is
ignored on this path regardless) non-finite variance-covariance entries from
survreg are escalated to an R-level stop().
Value
A list containing the following components:
- coefficients
A numeric vector of the estimated Weibull regression coefficients
\hat\beta, including the intercept.- log_sigma
The logarithm of the fitted scale parameter
\hat\sigmaof the Weibull AFT distribution.- vcov
The variance-covariance matrix of the estimated coefficients, or
NULLifestimate_only = TRUE(Rcpp path only).- neg_log_lik
(Rcpp path only) The negative log-likelihood at the fitted parameters.
- fisher_information
(Rcpp path only) The observed information matrix, or
NULLifestimate_only = TRUE.- std_errs
(survival path only) The coefficient standard errors,
sqrt(diag(vcov)).
See Also
fast_weibull_regression_general_cpp for the full model documentation
and Rcpp backend contract.
Examples
X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
fast_weibull_regression(y, dead, X)
Fast Weibull AFT Regression (C++)
Description
Weibull Accelerated Failure Time model fitting, exact/ right-censored responses only. See fast_weibull_regression_general_cpp for the left-/interval-censored extension.
Usage
fast_weibull_regression_cpp(
X,
y,
dead,
warm_start_params = NULL,
smart_cold_start = TRUE,
estimate_only = FALSE,
maxit = 100L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of survival times. |
dead |
A numeric vector of event indicators (1=event, 0=censored). |
warm_start_params |
Optional starting values for coefficients. |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess. |
estimate_only |
Logical. If TRUE, do not compute variance-covariance. |
maxit |
Maximum number of iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
optimization_alg |
Optimization algorithm. |
warm_start_fisher_info |
Optional initial Fisher Information matrix. |
Value
A list containing coefficients, log_sigma, and convergence status.
Fast Weibull Regression with General Censoring (C++ Backend)
Description
Weibull Accelerated Failure Time model fitting extended to
left-, right-, and interval-censored responses (TODO-3 in
interval_censored_survival_response.md). Zero-regression by
construction: exact/right-censored-only input uses the same likelihood
contributions as the corresponding survival::Surv() response.
Usage
fast_weibull_regression_general_cpp(
X,
y,
y_L,
y_R,
warm_start_params = NULL,
smart_cold_start = TRUE,
estimate_only = FALSE,
maxit = 100L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
A numeric matrix of predictors. |
y |
Exact survival times, |
y_L |
Censored-interval lower bounds, |
y_R |
Censored-interval upper bounds, |
warm_start_params |
Optional starting values for coefficients. |
smart_cold_start |
Logical. If TRUE, use an initial OLS-based guess. |
estimate_only |
Logical. If TRUE, do not compute variance-covariance. |
maxit |
Maximum number of iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
optimization_alg |
Optimization algorithm. |
warm_start_fisher_info |
Optional initial Fisher Information matrix. |
Value
A list containing the following components:
- coefficients
A numeric vector of the estimated Weibull regression coefficients, including the intercept.
- log_sigma
The logarithm of the scale parameter from the Weibull distribution.
- vcov
The variance-covariance matrix of the estimated coefficients.
- neg_ll
The negative log-likelihood at the final iteration.
- converged
A logical value indicating whether the algorithm converged.
A list containing coefficients, log_sigma, and convergence status.
Fast Zero-Inflated or Hurdle Poisson Regression (C++)
Description
Fits, by direct maximum likelihood, a two-component count model with a
Poisson log-link count component (\lambda_i = \exp(x_i^\top\beta_{\mathrm{cond}}),
X) and a logit-link binary component
(\pi_i = \mathrm{logit}^{-1}(x_{\mathrm{zi},i}^\top\beta_{\mathrm{zi}}),
Xzi). The two supported models differ in how \pi_i enters the
likelihood:
-
Zero-inflated Poisson (
is_hurdle = FALSE):\pi_iis the probability of an always-zero latent class, mixed with a Poisson count that can itself produce zeros,\Pr(Y_i = 0) = \pi_i + (1-\pi_i) e^{-\lambda_i}and\Pr(Y_i = y \mid y > 0) = (1-\pi_i)\,\mathrm{Poisson}(y; \lambda_i). -
Hurdle Poisson (
is_hurdle = TRUE):\pi_i = \Pr(Y_i = 0)directly, via a simple binary (zero vs. positive) logistic sub-model, and positive counts follow a zero-truncated Poisson,\Pr(Y_i = y \mid y > 0) = (1-\pi_i)\,\lambda_i^y e^{-\lambda_i} / \left(y!\,(1 - e^{-\lambda_i})\right).
Both branches share one likelihood/gradient/(analytic and expected) Hessian
implementation, switched at each observation by is_hurdle. Optimizes the
joint parameter vector [\beta_{\mathrm{cond}}, \beta_{\mathrm{zi}}] via
optimization_alg (default "lbfgs"). If the optimizer throws an
exception internally, this function does not propagate an R error:
it returns a diagnostic list with converged = FALSE, evaluated at the
optimizer's starting values and including the caught exception_message
(see Value), so callers must check converged before using the estimates.
Usage
fast_zero_augmented_poisson_cpp(
X,
y,
Xzi,
is_hurdle,
warm_start_params = NULL,
smart_cold_start = TRUE,
estimate_only = FALSE,
maxit = 1000L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL
)
Arguments
X |
Matrix of predictors for the conditional (Poisson count) component. |
y |
Vector of nonnegative-integer count responses. |
Xzi |
Matrix of predictors for the zero-inflation/hurdle (logistic) component. |
is_hurdle |
If |
warm_start_params |
Optional starting values for the joint
|
smart_cold_start |
Logical. If |
estimate_only |
If |
maxit |
Maximum number of optimizer iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional 1-indexed positions of parameters to hold fixed. |
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
Value
On optimizer failure (an internal exception, never an R error), and
regardless of estimate_only: a list with converged = FALSE,
num_iter = 0, hit_iteration_cap = FALSE, params (the
starting values the optimizer began from, after applying
fixed_idx/fixed_values), neg_ll/neg_loglik,
observed_information/fisher_information/information
(all evaluated at those starting values), information_type (always
"observed"), hessian, gradient_norm (NA if not
finite), min_eigenvalue_information (NA), params_origin
(a message that terminal parameters are unavailable after the exception) and
exception_message (the caught exception text).
Otherwise, if estimate_only = TRUE: a list with params (the joint
fitted [\hat\beta_{\mathrm{cond}}, \hat\beta_{\mathrm{zi}}] vector),
converged, neg_ll/neg_loglik (two aliases), and
gradient_norm. Otherwise, additionally: vcov (the joint
variance-covariance matrix), observed_information/fisher_information/
information (three aliases for the same observed-information matrix;
information_type is always "observed"), hessian (the negative of
that same matrix), and coefficients (a list with cond and zi
sub-vectors splitting params back into its two components).
Fixed parameters, warm starts
fixed_idx (1-indexed into the joint parameter vector, X's
coefficients first) and fixed_values optionally hold a subset of
parameters fixed at caller-supplied constant values rather than estimated.
warm_start_params supplies the full starting vector directly; otherwise,
if smart_cold_start = TRUE (the default), a model-specific heuristic
start is used, and if FALSE, all parameters start at zero except the
first conditional-model coefficient, initialized to \log(\bar y) (if
\bar y > 0). warm_start_fisher_info, if supplied, seeds the
curvature estimate used for the optimizer's first iteration.
Fast Zero/One-Inflated Beta Regression (C++)
Description
Fits, by direct maximum likelihood, a three-component mixture model for a
response Y_i \in [0, 1] (e.g. a bounded proportion with excess exact 0s
and 1s): with probability \pi_{0,i} the response is exactly 0, with
probability \pi_{1,i} it is exactly 1, and with probability
\pi_{b,i} = 1 - \pi_{0,i} - \pi_{1,i} it falls strictly inside
(0, 1) and follows a mean-precision Beta distribution. The three
category probabilities come from a multinomial-logit-style softmax over two
linear predictors on X_zero_one, with the "interior" (Beta) category as
the implicit zero baseline:
\pi_{0,i} = \frac{e^{\eta_{0,i}}}{e^{\eta_{0,i}} + e^{\eta_{1,i}} + 1},
\quad \pi_{1,i} = \frac{e^{\eta_{1,i}}}{e^{\eta_{0,i}} + e^{\eta_{1,i}} + 1},
\quad \eta_{0,i} = x_{\mathrm{zo},i}^\top\gamma_0,\ \ \eta_{1,i} = x_{\mathrm{zo},i}^\top\gamma_1,
and, conditional on 0 < Y_i < 1, Y_i \sim \mathrm{Beta}(\mu_i\phi,
(1-\mu_i)\phi) with \mu_i = \mathrm{logit}^{-1}(x_i^\top\beta) (clamped
to [10^{-8}, 1-10^{-8}]) and precision \phi = e^{\log\phi} —
matching the mean-precision Beta regression parameterization documented at
fast_beta_regression_cpp. Optimizes the joint parameter vector
[\beta, \log\phi, \gamma_0, \gamma_1] via optimization_alg
(default "lbfgs"), for up to 1500 iterations at gradient-norm tolerance
10^{-6} — both hardcoded, with no maxit/tol arguments
exposed by this function.
Usage
fast_zero_one_inflated_beta_cpp(
X,
X_zero_one,
y,
warm_start_params = NULL,
smart_cold_start = TRUE,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
Matrix of predictors for the interior Beta (mean) component. |
X_zero_one |
Matrix of predictors for the zero- and one-inflation
(mixture-probability) components; shared between |
y |
Vector of responses in |
warm_start_params |
Optional starting values for the joint |
smart_cold_start |
Logical. If |
fixed_idx |
Optional 1-indexed positions of parameters to hold fixed. |
fixed_values |
Optional values, parallel to |
optimization_alg |
Optimization algorithm (default |
warm_start_fisher_info |
Optional initial curvature (Fisher/observed information) matrix. |
estimate_only |
If |
Value
A list with components b (\hat\beta), log_phi,
zero_one_b0/zero_one_b1 (\hat\gamma_0/\hat\gamma_1),
params (the full fitted joint vector), neg_loglik, converged;
unless estimate_only = TRUE, additionally vcov (the joint
variance-covariance matrix — an all-NA matrix if the observed information
is non-finite or its free-parameter submatrix is not invertible, rather than an
error), observed_information/fisher_information/information
(three aliases for the same observed-information matrix; information_type
is always "observed"), and hessian (the negative of that same matrix).
Fixed parameters, warm starts
fixed_idx (1-indexed into [\beta, \log\phi, \gamma_0, \gamma_1])
and fixed_values optionally hold a subset of parameters fixed at
caller-supplied constant values rather than estimated. warm_start_params
supplies the full starting vector directly; otherwise, if smart_cold_start
= TRUE (the default): \beta starts from an OLS fit of
\mathrm{logit}(y_i) on X restricted to interior observations
(0 < y_i < 1), falling back to zero if fewer such observations than
ncol(X) are available; \log\phi starts at 2; and \gamma_0/
\gamma_1 each start from a separate OLS fit of X_zero_one on the
corresponding 0/1 indicator (\mathbb{1}[y_i = 0], \mathbb{1}[y_i =
1]). If smart_cold_start = FALSE, all parameters start at zero except
\log\phi, which still starts at 2. warm_start_fisher_info, if
supplied, seeds the curvature estimate used for the optimizer's first
iteration.
Fast Zero-Inflated Negative Binomial Regression (C++)
Description
High-performance zero-inflated negative binomial model fitting via L-BFGS.
Usage
fast_zinb_cpp(
X,
Xzi,
y,
warm_start_params = NULL,
maxit = 1000L,
tol = 1e-08,
fixed_idx = NULL,
fixed_values = NULL,
optimization_alg = "lbfgs",
smart_cold_start = TRUE,
warm_start_fisher_info = NULL,
estimate_only = FALSE
)
Arguments
X |
Numeric matrix of predictors for the count component (including intercept). |
Xzi |
Numeric matrix of predictors for the zero-inflation component (including intercept). |
y |
Numeric vector of non-negative integer count responses. |
warm_start_params |
Optional starting values for all parameters. |
maxit |
Maximum number of iterations. |
tol |
Convergence tolerance. |
fixed_idx |
Optional indices of fixed parameters. |
fixed_values |
Optional values for fixed parameters. |
optimization_alg |
Optimization algorithm (default "lbfgs"). |
smart_cold_start |
Logical. If TRUE, use a heuristic initial guess. |
warm_start_fisher_info |
Optional initial Fisher Information matrix. |
estimate_only |
Logical. If TRUE, skip variance computation and return only coefficients. |
Value
A list containing coefficients and convergence status.
Fast G-Computation (Standardization) Point Estimate for a Logit-Link Model (C++)
Description
Computes the G-computation (regression standardization) point estimate
of the marginal treatment effect under a fitted logit-link model — logistic
regression for a binary outcome, or the algebraically identical fractional
logit / quasi-binomial model for a proportion outcome in [0, 1], since
the standardization formula depends only on the fitted linear predictor and
link function, not the response's distributional assumptions. For each
subject i in the fitted sample, the fitted linear predictor is
decomposed into its treatment-free baseline \eta_{\mathrm{base},i} =
x_i^\top\hat\beta - \hat\beta_{j_{\mathrm{treat}}}\, x_{i,j_{\mathrm{treat}}}
and the two counterfactual predictions
\widehat{\Pr}(Y_i = 1 \mid \mathrm{do}(T=1)) = \mathrm{logit}^{-1}(\eta_{\mathrm{base},i}
+ \hat\beta_{j_{\mathrm{treat}}}) and \widehat{\Pr}(Y_i = 1 \mid \mathrm{do}(T=0)) =
\mathrm{logit}^{-1}(\eta_{\mathrm{base},i}) are computed by setting every
subject's treatment column to 1 (respectively 0) while leaving all other
covariates at their observed values — this is standard G-computation /
standardization: average the model-implied outcome over the empirical
covariate distribution under each counterfactual treatment assignment. The
two averages (mean1, mean0) and their difference (md, the
standardized average treatment effect on the risk-difference scale) are
returned.
Usage
gcomp_fractional_logit_point_estimate_cpp(X_fit, coef_hat, j_treat)
Arguments
X_fit |
Numeric matrix of predictors used to fit the model, including an intercept column if the model has one. |
coef_hat |
Numeric vector of fitted model coefficients |
j_treat |
1-based column index of the treatment indicator in |
Value
A list with elements mean1 (standardized mean outcome under
T=1 for everyone), mean0 (standardized mean outcome under T=0
for everyone), and md (mean1 - mean0, the standardized risk
difference).
See Also
gcomp_logistic_point_estimate_cpp, which computes the
identical quantity (it delegates directly to this function) under the
"logistic regression" framing.
Fast G-Computation (Standardization) Point Estimate for Logistic Regression (C++)
Description
Computes the standardized (G-computation) marginal risk difference under a
fitted logistic regression model. This is a thin alias: it delegates directly
to gcomp_fractional_logit_point_estimate_cpp (see that page for
the full standardization formula and counterfactual-averaging methodology,
which is identical for logistic and fractional-logit/quasi-binomial models),
passing its arguments through unchanged.
Usage
gcomp_logistic_point_estimate_cpp(X_fit, coef_hat, j_treat)
Arguments
X_fit |
Numeric matrix of predictors used to fit the model, including an intercept column if the model has one. |
coef_hat |
Numeric vector of fitted logistic regression coefficients
|
j_treat |
1-based column index of the treatment indicator in |
Value
A list with elements mean1 (standardized mean risk under
T=1 for everyone), mean0 (standardized mean risk under T=0
for everyone), and md (mean1 - mean0, the standardized risk
difference).
See Also
gcomp_fractional_logit_point_estimate_cpp for the full
documentation of the underlying computation.
Export of C++ function gcomp_logistic_post_fit_cpp
Description
Given an already-fitted logistic regression, computes a Huber-White/Eicker
sandwich (heteroskedasticity-robust, "HC0") coefficient covariance matrix and,
by the delta method, standard errors for the G-computed standardized risk
difference and log risk ratio (see
gcomp_logistic_point_estimate_cpp for the standardization point
estimates this builds inference around). The sandwich covariance is
\widehat{\mathrm{Var}}(\hat\beta) = B\,M\,B, with "bread"
B = (X^\top W X)^{-1} (W = \mathrm{diag}(\hat\mu_i(1-\hat\mu_i)), the
model-based Fisher information weights) and "meat" M = X^\top
\mathrm{diag}((y_i - \hat\mu_i)^2) X (the empirical score outer-product,
making this robust to model misspecification, not just relying on the
working Bernoulli variance). The risk-difference standard error is obtained
by propagating this sandwich covariance through the standardized risks'
gradients with respect to \beta (\nabla_\beta \bar{\mathrm{risk}}_1 -
\nabla_\beta \bar{\mathrm{risk}}_0, each a population-averaged logistic-derivative
weighted design-matrix sum, with the treatment column's gradient entry replaced
by the sum of the standardized-risk derivative directly since every subject's
treatment indicator is held fixed at 1 or 0 in the counterfactual averages);
the log risk ratio's standard error is obtained the same way via the gradient
of \log(\overline{\mathrm{risk}}_1) - \log(\overline{\mathrm{risk}}_0), and is only
computed (non-NA) when both standardized risks are strictly positive.
Aborts with an R error (rather than returning NAs) if mu_hat
contains non-finite or boundary (0 or 1) values, if the weighted design
crossproduct X^\top W X is not invertible, if the sandwich covariance
comes out non-finite anywhere, or if the treatment coefficient's variance is
non-positive.
Usage
gcomp_logistic_post_fit_cpp(X_fit, y, coef_hat, mu_hat, j_treat)
Arguments
X_fit |
Numeric matrix of predictors used to fit the model, including an intercept column if the model has one. |
y |
The observed binary (0/1) response used to fit the model. |
coef_hat |
Numeric vector of fitted logistic regression coefficients
|
mu_hat |
Numeric vector of fitted probabilities |
j_treat |
1-based column index of the treatment indicator in |
Value
A list with components vcov (the p \times p sandwich
covariance matrix), std_err and z_vals (per-coefficient standard
errors and Wald z-statistics, NA for any coefficient with non-finite or
non-positive variance), risk1/risk0 (the standardized mean risks
under T=1/T=0 for everyone), rd (risk1 - risk0) and
se_rd (its delta-method standard error), and log_rr/rr/
se_log_rr (the log risk ratio, risk ratio, and the log risk ratio's
delta-method standard error — all NA if either standardized risk is
not strictly positive).
See Also
gcomp_logistic_point_estimate_cpp for the point-estimate
computation this function's variances are built around;
gcomp_fractional_logit_post_fit_cpp() for the analogous
fractional-logit/quasi-binomial post-fit inference.
Export of C++ function gcomp_ordinal_proportional_odds_post_fit_cpp
Description
Computes the standardized (G-computation) marginal difference in
expected ordinal category score under a fitted proportional-odds
(cumulative logit) model for a K-category ordinal outcome coded
1, \ldots, K: \mathrm{logit}\,\Pr(Y_i \le k \mid x_i) = \alpha_k -
x_i^\top\beta, k = 1, \ldots, K-1. For each subject i, the fitted
linear predictor is decomposed into its treatment-free baseline
\eta_{\mathrm{base},i} = x_i^\top\hat\beta - \hat\beta_{j_{\mathrm{treat}}}\,
x_{i,j_{\mathrm{treat}}} and the counterfactual predictors \eta_{1,i} =
\eta_{\mathrm{base},i} + \hat\beta_{j_{\mathrm{treat}}} (treatment column set
to 1 for everyone) and \eta_{0,i} = \eta_{\mathrm{base},i} (set to 0 for
everyone). The subject-level expected category score under each counterfactual
is recovered from the fitted cumulative probabilities via the identity
E[Y_i] = \sum_{k=1}^K k\,\Pr(Y_i=k) = 1 + \sum_{k=1}^{K-1} \Pr(Y_i > k) = 1 +
\sum_{k=1}^{K-1} \left(1 - \mathrm{logit}^{-1}(\hat\alpha_k - \eta_i)\right),
and the two population-averaged expected scores (mean1, mean0)
and their difference (md) are returned — the ordinal analogue of
gcomp_logistic_point_estimate_cpp's risk difference. Unlike
gcomp_logistic_post_fit_cpp, this function computes
point estimates only; no sandwich covariance, standard errors, or
inferential quantities are returned despite the _post_fit name.
Usage
gcomp_ordinal_proportional_odds_post_fit_cpp(
X_fit,
coef_hat,
alpha_hat,
j_treat
)
Arguments
X_fit |
Numeric matrix of predictors used to fit the model, including an
intercept column if the model has one (conventionally absorbed into the
thresholds |
coef_hat |
Numeric vector of fitted proportional-odds regression coefficients
|
alpha_hat |
Numeric vector of the |
j_treat |
1-based column index of the treatment indicator in |
Value
A list with elements mean1 (standardized expected category score
under T=1 for everyone), mean0 (standardized expected category score
under T=0 for everyone), and md (mean1 - mean0).
See Also
gcomp_logistic_point_estimate_cpp for the binary-outcome
analogue; gcomp_logistic_post_fit_cpp for an example of the
sandwich-variance inference this function does not provide.
Generate Synthetic Simulation Covariates and Continuous Response
Description
A helper function to generate synthetic covariates and a latent continuous response
identical to the logic used within SimulationFramework. Covariates may be
supplied directly via X_mat or drawn randomly via cov_draw_method;
exactly one of the two must be non-NULL.
Usage
generate_covariate_dataset(
n,
p,
cond_exp_func_model = c("linear", "nonlinear"),
norm_sq_beta_vec = 1,
X_mat = NULL,
cov_draw_method = stats::rnorm,
cov_draw_method_args = list(mean = 0, sd = 1)
)
Arguments
n |
Integer. Sample size (number of rows). |
p |
Integer. Number of covariates (number of columns). |
cond_exp_func_model |
Character scalar. Either |
norm_sq_beta_vec |
Positive numeric scalar. The desired squared Euclidean norm
of the coefficient vector, i.e. |
X_mat |
Numeric matrix of dimensions |
cov_draw_method |
A function used to draw |
cov_draw_method_args |
Named list of additional arguments forwarded to
|
Details
Two conditional-expectation models are supported, both rescaled so that
the latent response y_cont has a controlled signal magnitude:
"linear"y_i = x_i^\top \beta, with\betaa fixed evenly spaced sequence from1to-1across thepcovariates (seq(1, -1, length.out = p)) — i.e. the first covariate gets the strongest positive effect, the last the strongest negative effect, and (for evenp) covariates near the middle get effects near zero — rescaled so that\|\beta\|_2^2 = \code{norm\_sq\_beta\_vec}."nonlinear"The Friedman (1991) MARS benchmark function applied to the first five covariates (requires
p \ge 5; covariates beyond the fifth do not enter the response at all),y_i = c \left(10 \sin(\pi x_{i1} x_{i2}) + 20 (x_{i3} - 0.5)^2 + 10 x_{i4} + 5 x_{i5}\right),with the overall scale
cchosen so thatc^2 \sum_k \beta_{\mathrm{friedman},k}^2 = \code{norm\_sq\_beta\_vec}for the nominal term coefficients\beta_{\mathrm{friedman}} = (10, 20, 10, 5)(the\sinterm's coefficient is not part of this nominal vector, so the realized\|c \cdot (10,20,10,5)\|_2^2may not exactly equalnorm_sq_beta_veconce the\sinterm's own variance is accounted for —norm_sq_beta_veccalibrates the linear-in-covariate part of the scale, not the full response variance). This function assumes each covariate lies roughly in[-1, 1], matching Friedman's original specification; covariates drawn or supplied on a very different scale will not reproduce the intended signal shape.
Value
A list with two elements: X (a data frame of covariates) and
y_cont (a numeric vector of the latent continuous response).
References
Friedman, J. H. (1991). "Multivariate Adaptive Regression Splines."
The Annals of Statistics, 19(1), 1-67, doi:10.1214/aos/1176347963, for
the nonlinear benchmark function used when cond_exp_func_model =
"nonlinear".
Examples
generate_covariate_dataset(n = 10, p = 5)
Generate Atkinson optimal-design randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_atkinson_cpp(X_sexp, n, p_raw, prob_T, nsim)
Arguments
X_sexp |
Numeric matrix: the design's covariate model matrix for the first
|
n |
Number of subjects. |
p_raw |
Number of raw covariate columns (before model-matrix expansion);
the first |
prob_T |
Probability of assignment to treatment. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate Bernoulli randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_bernoulli_cpp(n, nsim, prob_T)
Arguments
n |
Number of subjects. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
prob_T |
Probability of assignment to treatment. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate blocked randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_blocking_cpp(n, nsim, prob_T, strata_indices)
Arguments
n |
Number of subjects. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
prob_T |
Probability of assignment to treatment. |
strata_indices |
List of integer vectors, one per stratum, holding the (1-based) indices of the subjects in that stratum. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate cluster randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_cluster_cpp(n, nsim, prob_T, cluster_indices)
Arguments
n |
Number of subjects. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
prob_T |
Probability of assignment to treatment. |
cluster_indices |
List of integer vectors, one per cluster, holding the (1-based) indices of the subjects in that cluster; each cluster is assigned to one arm as a whole. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate Efron biased-coin randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_efron_cpp(n, nsim, prob_T, weighted_coin_prob)
Arguments
n |
Number of subjects. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
prob_T |
Probability of assignment to treatment. |
weighted_coin_prob |
Efron's biased-coin probability: the chance of assigning the currently under-represented arm. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate IBCRD randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_ibcrd_cpp(n, nsim, prob_T)
Arguments
n |
Number of subjects. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
prob_T |
Probability of assignment to treatment. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate matched-pair randomization permutations
Description
Every generate_permutations_*_cpp function in this file is seeded from one R::unif_rand() draw into edi_rng::RRng (RNG.h), a portable re-implementation of R's own Mersenne-Twister generator – a given seed therefore produces identical draws in R and in any future binding (e.g. Python) using the same core and the same seed.
Usage
generate_permutations_matching_cpp(m_vec, nsim, prob_T)
Arguments
m_vec |
Integer vector of match ids, one per subject: subjects sharing a positive id form a matched pair (randomized within the pair); 0 marks an unmatched (reservoir) subject, assigned by an independent coin flip. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
prob_T |
Probability of assignment to treatment. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate Pocock-Simon minimization randomization permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_pocock_simon_cpp(
x_levels_matrix,
num_levels_total,
weights,
p_best,
prob_T,
nsim
)
Arguments
x_levels_matrix |
Integer matrix with one row per subject and one column
per stratification covariate; each entry is the subject's level for that
covariate as a (1-based) index into 1.. |
num_levels_total |
Total number of levels across all stratification covariates. |
weights |
Numeric vector of per-covariate imbalance weights (one per column
of |
p_best |
Probability of assigning the arm that minimizes the weighted imbalance. |
prob_T |
Probability of assignment to treatment. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Generate stratified permuted-block randomization (SPBR) permutations
Description
See generate_permutations_matching_cpp for the reproducibility note that applies to every function in this file.
Usage
generate_permutations_spbr_cpp(strata_keys, block_size, prob_T, nsim)
Arguments
strata_keys |
Character vector giving each subject's stratum label, in arrival order. |
block_size |
Size of the permuted blocks used within each stratum. |
prob_T |
Probability of assignment to treatment. |
nsim |
Number of randomization draws (columns of the returned matrix) to generate. |
Value
A list with w_mat, an n x nsim integer matrix whose
columns are independent 0/1 treatment-assignment draws, and m_mat
(always NULL).
Beta Regression Hessian, Standalone (C++)
Description
Computes the Hessian matrix (second derivatives with respect to
[\beta, \log\phi]) of the log-likelihood of the mean-precision Beta
regression model documented in full at fast_beta_regression_cpp,
at arbitrary caller-supplied parameters params (not necessarily the
MLE). Exported standalone — independent of any optimizer run — for direct
numerical diagnostics (e.g. checking curvature or building a custom variance
estimate at a specific parameter value) and for use by
get_beta_regression_score_cpp's sibling relationship in
optimizer/inference code that needs both quantities at the same point.
Usage
get_beta_regression_hessian_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors, as used to fit the model. |
y |
A numeric vector of responses in |
params |
A numeric vector |
Value
The (p+1) \times (p+1) Hessian matrix of the log-likelihood
(i.e. the negative of the observed information) at params.
See Also
get_beta_regression_score_cpp for the corresponding
gradient at the same point; fast_beta_regression_cpp for the
full mean-precision Beta regression model documentation.
Compute Beta Regression Score
Description
Calculates the score vector (gradient of the log-likelihood) for a beta regression model.
Usage
get_beta_regression_score_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of responses (in (0, 1)). |
params |
A numeric vector of parameters [beta, log_phi]. |
Value
A numeric vector representing the score.
Get the default bootstrap dispatch policy
Description
Returns EDI's built-in policy table, consulted by the internal (non-exported)
dispatcher edi_bootstrap_dispatch_policy(), for choosing which bootstrap
confidence-interval type — "bca" (bias-corrected and accelerated) or
"percentile" — an inference class uses by default.
Usage
get_bootstrap_dispatch_policy()
Details
The dispatcher resolves a type for a given inference-class name (and, if available, the fitted inference object) in this precedence order, returning the first match:
If the object's experimental design class matches a key of
design_class_overrides(viais), and the inference class name matches one of that design's named regular-expression patterns, use the associated type.Otherwise, if the inference class name matches one of
inference_class_overrides's named regular-expression patterns (checked in list order, first match wins), use the associated type.Otherwise, fall back to
default_type("bca").
Whichever type is resolved by that process, a final safety check applies: if the
resolved type is "bca" and the fitted object reports (via its private
jackknife_block_size_gt_one_unsupported() method) that BCa's required
jackknife computation is unsupported for its current data (e.g. a block size
greater than 1), the type is silently downgraded to "percentile" instead.
This override table exists because BCa is the generally preferred default (it
corrects for both bias and skewness in the bootstrap distribution), but is
empirically unreliable or computationally unsupported for specific inference/
design class combinations — the "percentile" overrides listed here were
added as those cases were identified, not derived from a general rule.
Value
A named list describing the default bootstrap type configuration, with
components default_type (the fallback type, "bca"),
inference_class_overrides (a named character vector: regular-expression
pattern names to bootstrap-type values, matched against the inference class
name), and design_class_overrides (a named list keyed by experimental
design class name, each value itself a named character vector of
pattern-to-type overrides scoped to that design).
See Also
get_parallel_dispatch_policy for the analogous policy
controlling forced-serial dispatch; get_optimization_dispatch_policy
for the analogous policy controlling default optimizer algorithm choice.
Examples
get_bootstrap_dispatch_policy()
Get the default cold-start dispatch policy
Description
Returns EDI's built-in policy table, consulted by the internal (non-exported)
dispatcher edi_cold_start_dispatch_policy(), for the smart_cold_start
default used by each inference class's C++ model-fitting backend. A TRUE
entry means the solver initializes via an OLS (or otherwise model-appropriate
heuristic) warm-up before iterating; FALSE means a plain zero-vector cold
start. Benchmarks show the OLS warm-up is net-negative for logistic and Poisson
IRLS at typical trial sizes (the one extra OLS solve costs more than the IRLS
iterations it saves), so those families — along with several G-computation-based
incidence/proportion inference classes — default to FALSE here.
Usage
get_cold_start_dispatch_policy()
Details
The dispatcher checks the inference class name against
inference_class_overrides's named regular-expression patterns in list
order, returning the associated logical value at the first match; if none
match, it falls back to default (TRUE). Unlike
get_bootstrap_dispatch_policy, there is no separate
design-class-scoped override table here — only a single flat pattern list.
Value
A named list with default (logical, the fallback when no
override pattern matches; TRUE in the built-in policy) and
inference_class_overrides (a named logical vector: regular-expression
pattern names to TRUE/FALSE values, matched against the
inference class name).
These TRUE/FALSE defaults are empirical performance
judgments computed on the maintainer's machine, not correctness facts —
the same heuristic can be net-positive or net-negative depending on
your hardware's core count, cache sizes, and BLAS backend. Run
tune_EDI_for_this_machine to re-measure this axis on your
own machine and persist any better setting it finds.
See Also
get_bootstrap_dispatch_policy and
get_optimization_dispatch_policy for the analogous policies
controlling bootstrap CI type and default optimizer algorithm;
set_cold_start_dispatch_policy to override this policy at
runtime; tune_EDI_for_this_machine to re-benchmark it on
your own hardware.
Examples
get_cold_start_dispatch_policy()
Combined Conditional-Poisson/Poisson Hessian, Standalone (C++)
Description
Computes the Hessian matrix of the log-likelihood of the combined KK
matched-pair conditional-Poisson (conditional-Binomial) plus reservoir
marginal-Poisson model documented in full at
fast_cpoisson_combined_with_var_cpp, at arbitrary
caller-supplied params_r (not necessarily the MLE). Internally reuses
the same score-and-information computation as
get_cpoisson_combined_score_cpp (a single shared routine
computes both at once) and returns the negative of the resulting information
matrix, i.e. the actual Hessian of the log-likelihood. Exported standalone —
independent of any optimizer run — for direct numerical diagnostics at a
specific parameter value.
Usage
get_cpoisson_combined_hessian_cpp(
yT_v_r,
n_k_v_r,
X_diff_v_r,
y_r_r,
w_r_r,
X_r_r,
params_r
)
Arguments
yT_v_r |
Treated-subject outcome count per matched pair. |
n_k_v_r |
Total (treated + control) outcome count per matched pair. |
X_diff_v_r |
Covariate differences (treated minus control) between the members of each matched pair. |
y_r_r |
Reservoir (unmatched) subjects' outcomes. |
w_r_r |
Reservoir subjects' treatment indicators. |
X_r_r |
Reservoir subjects' covariates. |
params_r |
A numeric vector of model parameters at which to evaluate the Hessian. |
Value
The Hessian matrix of the log-likelihood (the negative of the
information matrix) at params_r.
See Also
get_cpoisson_combined_score_cpp for the corresponding
gradient at the same point; fast_cpoisson_combined_with_var_cpp
for the full model documentation.
Combined Conditional-Poisson/Poisson Score, Standalone (C++)
Description
Computes the score vector (gradient of the log-likelihood) of the combined KK
matched-pair conditional-Poisson (conditional-Binomial) plus reservoir
marginal-Poisson model documented in full at
fast_cpoisson_combined_with_var_cpp, at arbitrary
caller-supplied params_r (not necessarily the MLE). Exported standalone
— independent of any optimizer run — for direct numerical diagnostics (e.g.
verifying convergence, or building a custom estimating-equation solver) at a
specific parameter value.
Usage
get_cpoisson_combined_score_cpp(
yT_v_r,
n_k_v_r,
X_diff_v_r,
y_r_r,
w_r_r,
X_r_r,
params_r
)
Arguments
yT_v_r |
Treated-subject outcome count per matched pair. |
n_k_v_r |
Total (treated + control) outcome count per matched pair. |
X_diff_v_r |
Covariate differences (treated minus control) between the members of each matched pair. |
y_r_r |
Reservoir (unmatched) subjects' outcomes. |
w_r_r |
Reservoir subjects' treatment indicators. |
X_r_r |
Reservoir subjects' covariates. |
params_r |
A numeric vector of model parameters at which to evaluate the score. |
Value
The score vector (gradient of the log-likelihood) at params_r.
See Also
get_cpoisson_combined_hessian_cpp for the corresponding
Hessian at the same point; fast_cpoisson_combined_with_var_cpp
for the full model documentation.
Effective capabilities for an inference class or instance
Description
Effective capabilities for an inference class or instance
Usage
get_effective_capabilities(name, des_obj = NULL, live_obj = NULL)
Arguments
name |
Either a class name (character), or an already-constructed inference object. Passing an object additionally refines the static answer with that object's own live-checkable capability gates (see EDI_INFERENCE_LIVE_CAPABILITY_GATES) – some supports_*() private methods are conditional on constructor arguments (e.g. InferenceSurvivalCoxPHRegr's use_rcpp), which a class-name-only, cached lookup can never reflect. The name-only path's cost and cached result are unaffected by this – a live object is never written into EDI_INFERENCE_EFFECTIVE_CAPABILITIES_CACHE, since the answer for one specific instance must not be silently handed to every future name-only caller for that class. |
des_obj |
Optional design object. Capabilities the design rules out
(see |
live_obj |
Optional: an already-constructed inference object to use
for the live-gate refinement, separate from |
Identity-Link (Risk-Difference) Binomial Regression Hessian, Standalone (C++)
Description
Computes a numerical (central finite-difference 4-point stencil, step
h = 10^{-4}) approximation of the Hessian matrix of the log-likelihood of
the constrained identity-link binomial regression model documented in full at
fast_identity_binomial_regression_cpp, at arbitrary
caller-supplied beta (not necessarily the MLE) — not an analytic
second derivative. Exported standalone — independent of any optimizer run —
for direct numerical diagnostics at a specific parameter value.
Usage
get_identity_binomial_regression_hessian_cpp(X, y_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated Hessian matrix of the log-likelihood at beta.
See Also
get_identity_binomial_regression_score_cpp for the
corresponding (also finite-difference) gradient at the same point;
fast_identity_binomial_regression_cpp for the full model
documentation, including the probability-boundary constraint this Hessian is
evaluated without enforcing.
Identity-Link (Risk-Difference) Binomial Regression Score, Standalone (C++)
Description
Computes a numerical (central finite-difference, step h =
10^{-6}) approximation of the score vector (gradient of the log-likelihood)
of the constrained identity-link binomial regression model documented in full
at fast_identity_binomial_regression_cpp, at arbitrary
caller-supplied beta (not necessarily the MLE) — not an analytic
derivative. Exported standalone — independent of any optimizer run — for
direct numerical diagnostics (e.g. verifying convergence, or cross-checking an
analytic gradient elsewhere) at a specific parameter value.
Usage
get_identity_binomial_regression_score_cpp(X, y_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated score vector at beta.
See Also
get_identity_binomial_regression_hessian_cpp for the
corresponding (also finite-difference) Hessian at the same point;
fast_identity_binomial_regression_cpp for the full model
documentation.
Weighted Identity-Link (Risk-Difference) Binomial Regression Hessian, Standalone (C++)
Description
Computes the observation-weighted Hessian matrix of the weighted log-likelihood
of the constrained identity-link binomial regression model documented in full
at fast_identity_binomial_regression_cpp, at arbitrary
caller-supplied beta (not necessarily the MLE), with each observation's
contribution multiplied by weights_r[i], via a numerical
(central finite-difference 4-point stencil, step h = 10^{-4})
approximation — not an analytic second derivative. Exported standalone —
independent of any optimizer run — for direct numerical diagnostics at a
specific parameter value.
Usage
get_identity_binomial_regression_weighted_hessian_cpp(X, y_r, weights_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
weights_r |
A nonnegative numeric vector of observation weights. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated weighted Hessian matrix at beta.
See Also
get_identity_binomial_regression_weighted_score_cpp for
the corresponding weighted gradient at the same point;
get_identity_binomial_regression_hessian_cpp for the unweighted
version; fast_identity_binomial_regression_cpp for the full model
documentation.
Weighted Identity-Link (Risk-Difference) Binomial Regression Score, Standalone (C++)
Description
Computes the observation-weighted score vector (gradient of the weighted
log-likelihood) of the constrained identity-link binomial regression model
documented in full at fast_identity_binomial_regression_cpp, at
arbitrary caller-supplied beta (not necessarily the MLE), with each
observation's contribution multiplied by weights_r[i], via a
numerical (central finite-difference, step h = 10^{-6})
approximation — not an analytic derivative. Exported standalone — independent
of any optimizer run — for direct numerical diagnostics at a specific
parameter value.
Usage
get_identity_binomial_regression_weighted_score_cpp(X, y_r, weights_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
weights_r |
A nonnegative numeric vector of observation weights. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated weighted score vector at beta.
See Also
get_identity_binomial_regression_weighted_hessian_cpp for
the corresponding weighted Hessian at the same point;
get_identity_binomial_regression_score_cpp for the unweighted
version; fast_identity_binomial_regression_cpp for the full model
documentation.
Show this machine's saved EDI tuning, if any
Description
Reads the per-user config file written by
tune_EDI_for_this_machine and returns it as an
EDILocalMachineTuning object (whose print method shows when and
how it was produced, the hardware fingerprint it was measured on, and
every policy deviation it stores). Does not apply anything –
application happens inside tune_EDI_for_this_machine() itself and
at package load.
Usage
get_local_EDI_optimization()
Value
Invisibly, the saved EDILocalMachineTuning object, or
NULL (with a message) if no valid saved tuning exists.
See Also
tune_EDI_for_this_machine,
clear_local_EDI_optimization.
Examples
get_local_EDI_optimization()
Log-Link (Relative-Risk) Binomial Regression Hessian, Standalone (C++)
Description
Computes a numerical (central finite-difference 4-point stencil, step
h = 10^{-4}) approximation of the Hessian matrix of the log-likelihood of
the constrained log-link binomial regression model documented in full at
fast_log_binomial_regression_cpp, at arbitrary caller-supplied
beta (not necessarily the MLE) — not an analytic second derivative.
Exported standalone — independent of any optimizer run — for direct numerical
diagnostics at a specific parameter value.
Usage
get_log_binomial_regression_hessian_cpp(X, y_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated Hessian matrix of the log-likelihood at beta.
See Also
get_log_binomial_regression_score_cpp for the
corresponding (also finite-difference) gradient at the same point;
fast_log_binomial_regression_cpp for the full model
documentation, including the probability-boundary constraint this Hessian is
evaluated without enforcing.
Log-Link (Relative-Risk) Binomial Regression Score, Standalone (C++)
Description
Computes a numerical (central finite-difference, step h =
10^{-6}) approximation of the score vector (gradient of the log-likelihood)
of the constrained log-link binomial regression model documented in full at
fast_log_binomial_regression_cpp, at arbitrary caller-supplied
beta (not necessarily the MLE) — not an analytic derivative. Exported
standalone — independent of any optimizer run — for direct numerical
diagnostics (e.g. verifying convergence) at a specific parameter value.
Usage
get_log_binomial_regression_score_cpp(X, y_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated score vector at beta.
See Also
get_log_binomial_regression_hessian_cpp for the
corresponding (also finite-difference) Hessian at the same point;
fast_log_binomial_regression_cpp for the full model
documentation.
Weighted Log-Link (Relative-Risk) Binomial Regression Hessian, Standalone (C++)
Description
Computes the observation-weighted Hessian matrix of the weighted log-likelihood
of the constrained log-link binomial regression model documented in full at
fast_log_binomial_regression_cpp, at arbitrary caller-supplied
beta (not necessarily the MLE), with each observation's contribution
multiplied by weights_r[i], via a numerical (central
finite-difference 4-point stencil, step h = 10^{-4}) approximation —
not an analytic second derivative. Exported standalone — independent of any
optimizer run — for direct numerical diagnostics at a specific parameter value.
Usage
get_log_binomial_regression_weighted_hessian_cpp(X, y_r, weights_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
weights_r |
A nonnegative numeric vector of observation weights. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated weighted Hessian matrix at beta.
A numeric matrix representing the weighted Hessian.
See Also
get_log_binomial_regression_weighted_score_cpp for
the corresponding weighted gradient at the same point;
get_log_binomial_regression_hessian_cpp for the unweighted
version; fast_log_binomial_regression_cpp for the full model
documentation.
Weighted Log-Link (Relative-Risk) Binomial Regression Score, Standalone (C++)
Description
Computes the observation-weighted score vector (gradient of the weighted
log-likelihood) of the constrained log-link binomial regression model
documented in full at fast_log_binomial_regression_cpp, at
arbitrary caller-supplied beta (not necessarily the MLE), with each
observation's contribution multiplied by weights_r[i], via a
numerical (central finite-difference, step h = 10^{-6})
approximation — not an analytic derivative. Exported standalone — independent
of any optimizer run — for direct numerical diagnostics at a specific
parameter value.
Usage
get_log_binomial_regression_weighted_score_cpp(X, y_r, weights_r, beta)
Arguments
X |
A numeric matrix of predictors. |
y_r |
A binary (0/1) numeric vector of responses. |
weights_r |
A nonnegative numeric vector of observation weights. |
beta |
A numeric vector of coefficients |
Value
The finite-difference-approximated weighted score vector at beta.
See Also
get_log_binomial_regression_weighted_hessian_cpp for
the corresponding weighted Hessian at the same point;
get_log_binomial_regression_score_cpp for the unweighted
version; fast_log_binomial_regression_cpp for the full model
documentation.
Negative Binomial Regression Hessian, Standalone (C++)
Description
Computes the (analytic) Hessian matrix of the log-likelihood of the
mean/dispersion-parameterized negative binomial regression model documented
in full at fast_neg_bin_cpp (see also
fast_dnbinom_mu_vec_cpp for the underlying density), at
arbitrary caller-supplied params (not necessarily the MLE). Exported
standalone — independent of any optimizer run — for direct numerical
diagnostics at a specific parameter value.
Usage
get_negbin_regression_hessian_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors, as used to fit the model. |
y |
A numeric vector of nonnegative-integer count responses. |
params |
A numeric vector |
Value
The (p+1) \times (p+1) Hessian matrix of the log-likelihood at params.
See Also
get_negbin_regression_score_cpp for the corresponding
gradient at the same point; fast_neg_bin_cpp for the full
model documentation.
Compute Negative Binomial Regression Score
Description
Calculates the score vector (gradient of the log-likelihood) for a negative binomial regression model.
Usage
get_negbin_regression_score_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of responses (non-negative integers). |
params |
A numeric vector of parameters [beta, log_theta]. |
Value
A numeric vector representing the score.
Get the maximum number of threads for OpenMP
Description
Get the maximum number of threads for OpenMP
Usage
get_omp_max_threads_cpp()
Value
Integer.
Get the default optimization dispatch policy
Description
Returns EDI's built-in policy table, consulted by the internal (non-exported)
dispatcher edi_optimization_dispatch_policy(), for choosing which
optimization algorithm ("newton_raphson", "lbfgs", or
"irls") an inference class's C++ model-fitting backend uses by default.
Usage
get_optimization_dispatch_policy()
Details
The dispatcher checks the inference class name against
inference_class_overrides's named regular-expression patterns in list
order, returning the associated algorithm string at the first match; if none
match, it falls back to default_alg ("newton_raphson"). Unlike
get_bootstrap_dispatch_policy, there is no separate
design-class-scoped override table here — only a single flat pattern list.
The built-in overrides are empirical, chosen per model family based on which
algorithm converges fastest/most reliably for that likelihood surface in
practice — e.g. plain-vanilla generalized linear models with a canonical or
near-canonical link (Poisson, quasi-Poisson, robust Poisson, various incidence
models) default to "irls", most non-canonical-link and ordinal/survival
models default to "lbfgs", and stratified Cox PH and most matched
(KK*GLMM-adjacent) models default to "newton_raphson".
Value
A named list with components default_alg (the fallback
algorithm, "newton_raphson" in the built-in policy) and
inference_class_overrides (a named character vector: regular-expression
pattern names to algorithm-name values, matched against the inference class
name).
Which algorithm converges fastest/most reliably per family is partly a
hardware fact (relative cost of Hessian solves vs. L-BFGS iterations
depends on BLAS and cache), so these defaults, computed on the
maintainer's machine, are not necessarily optimal on yours. Run
tune_EDI_for_this_machine to re-measure this axis on your
own machine — it will only switch a family's algorithm when the
candidate converges on every benchmark replicate, never trading speed
for a convergence failure.
See Also
get_bootstrap_dispatch_policy and
get_cold_start_dispatch_policy for the analogous policies
controlling bootstrap CI type and cold-start behavior;
.normalize_optimizer_algorithm for how a resolved algorithm
string is validated/normalized before being passed to a C++ backend;
tune_EDI_for_this_machine to re-benchmark this policy on
your own hardware.
Examples
get_optimization_dispatch_policy()
Proportional-Odds Ordinal Regression Hessian, Standalone (C++)
Description
Computes the (analytic) Hessian matrix of the log-likelihood of the logit-link
cumulative (proportional-odds) ordinal regression model documented in full at
fast_ordinal_regression_cpp, at arbitrary caller-supplied
params (not necessarily the MLE). Exported standalone — independent of
any optimizer run — for direct numerical diagnostics at a specific parameter
value.
Usage
get_ordinal_regression_hessian_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
params |
A numeric vector |
Value
The Hessian matrix of the log-likelihood at params.
See Also
get_ordinal_regression_score_cpp for the corresponding
gradient at the same point; fast_ordinal_regression_cpp for
the full model documentation.
Proportional-Odds Ordinal Regression Score, Standalone (C++)
Description
Computes the (analytic) score vector (gradient of the log-likelihood) of the
logit-link cumulative (proportional-odds) ordinal regression model documented
in full at fast_ordinal_regression_cpp, at arbitrary
caller-supplied params (not necessarily the MLE). Exported standalone
— independent of any optimizer run — for direct numerical diagnostics (e.g.
verifying convergence) at a specific parameter value.
Usage
get_ordinal_regression_score_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of ordinal responses; only the rank order of distinct values matters, not their numeric coding. |
params |
A numeric vector |
Value
The score vector (gradient of the log-likelihood) at params.
See Also
get_ordinal_regression_hessian_cpp for the
corresponding Hessian at the same point; fast_ordinal_regression_cpp
for the full model documentation.
Get the default parallel dispatch policy
Description
Returns EDI's built-in blocklist-first policy table, consulted by an
internal (non-exported) dispatcher, for deciding whether a given
"bootstrap" or "rand_ci" (randomization confidence interval)
operation is forced to run serially rather than in parallel, for a
given inference class and response type. "Blocklist-first" means every
operation is parallel-eligible by default; only combinations explicitly
matched below are forced serial.
Usage
get_parallel_dispatch_policy()
Details
For a given operation ("bootstrap" or
"rand_ci"), the dispatcher looks up that operation's sub-list (e.g.
bootstrap) and forces serial execution if either the inference class
name matches any of serial_inference_class_patterns (regular
expressions, PCRE-flavored — serial_inference_class_patterns supports
lookahead, e.g. the built-in
"^InferenceSurvival(?!.*KK)" pattern matches non-KK survival inference
classes but excludes matched-design (KK) survival variants) or the response
type matches any of serial_response_types exactly. In the built-in
policy, incidence-response inference generally runs serially for both
operations (its bootstrap/randomization resampling is not currently
parallel-safe or not worth parallelizing at typical trial sizes), and
non-KK survival inference and InferenceAllKKWilcoxIVWC are
additionally forced serial for bootstrapping specifically.
Value
A named list with two components, bootstrap and rand_ci,
each itself a list with serial_inference_class_patterns (a character
vector of PCRE-flavored regular expressions matched against the inference
class name) and serial_response_types (a character vector of exact
response-type strings) — any match in either forces that operation to run
serially rather than in parallel.
Unlike get_cold_start_dispatch_policy,
get_warm_start_dispatch_policy, and
get_optimization_dispatch_policy, this table is a
correctness fact, not a performance one — every entry exists
because that operation is not currently parallel-safe for that class or
response type, not because it happens to be slower in parallel.
tune_EDI_for_this_machine benchmarks a related but
separate question — at what sample size parallel execution starts to
beat serial, and which core count wins — and will never propose
un-serializing anything named here.
See Also
get_bootstrap_dispatch_policy,
get_cold_start_dispatch_policy, and
get_optimization_dispatch_policy for the analogous policies
controlling bootstrap CI type, cold-start behavior, and default optimizer
algorithm (none of which control parallel-vs-serial dispatch);
tune_EDI_for_this_machine for the machine-dependent
parallel-crossover/core-count benchmark this blocklist constrains but
is never itself a target of.
Examples
get_parallel_dispatch_policy()
Calculates the standard error of the difference in restricted mean survival times
Description
Calculates the standard error of the difference in restricted mean survival times
Usage
get_restricted_mean_se_diff(y, dead, w)
Arguments
y |
Numeric vector of survival times. |
dead |
Integer vector of event indicators (1=event, 0=censored). |
w |
Integer vector of treatment assignments (1=treatment, 0=control). |
Value
The standard error of the difference.
Calculates standard variance using the formula from Uno et al
Description
Var(RMST) = \sum_j A(t_j)^2 d_j / (n_j (n_j - d_j))
where A(t_j) = \int_{t_j}^{\tau} S(u) du is the remaining area under the KM
curve from event time t_j to the last observation \tau.
Here d_j is the number of events at t_j, and n_j
is the number at risk just before t_j.
Terms where n_j == d_j are omitted: S drops to 0 there, so A(t_j) = 0 and the
contribution is 0 in the limit regardless of the undefined Greenwood denominator.
Usage
get_restricted_mean_se_for_group(y, dead)
Arguments
y |
Numeric vector of survival times. |
dead |
Integer vector of event indicators (1=event, 0=censored). |
Value
The standard error of the restricted mean.
Stereotype Logit Regression Hessian, Standalone (C++)
Description
Computes the (analytic) Hessian matrix of the log-likelihood of the
stereotype (reduced-rank multinomial) logistic regression model documented
in full at fast_stereotype_logit_cpp, at arbitrary
caller-supplied params (not necessarily the MLE). Exported standalone
— independent of any optimizer run — for direct numerical diagnostics at a
specific parameter value.
Usage
get_stereotype_logit_hessian_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors (no intercept column needed; see
|
y |
A numeric vector of categorical (nominal or ordinal) responses; only the set of distinct values matters, not their numeric coding or order. |
params |
A numeric vector of the full joint parameter vector |
Value
The Hessian matrix of the log-likelihood at params.
See Also
get_stereotype_logit_score_cpp for the corresponding
gradient at the same point; fast_stereotype_logit_cpp for the
full model documentation.
Compute Stereotype Logit Score
Description
Calculates the score vector (gradient of the log-likelihood) for a stereotype logit model.
Usage
get_stereotype_logit_score_cpp(X, y, params)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of responses. |
params |
A numeric vector of parameters. |
Value
A numeric vector representing the score.
Calculates the difference in a survival statistic (median or restricted mean) between two groups (treatment vs control)
Description
Calculates the difference in a survival statistic (median or restricted mean) between two groups (treatment vs control)
Usage
get_survival_stat_diff(y, dead, w, requested_stat)
Arguments
y |
Numeric vector of survival times. |
dead |
Integer vector of event indicators (1=event, 0=censored). |
w |
Integer vector of treatment assignments (1=treatment, 0=control). |
requested_stat |
A string, either "median" or "restricted_mean". |
Value
The difference in the statistic (treatment - control).
Calculates the median or restricted mean survival time for a single group
Description
Calculates the median or restricted mean survival time for a single group
Usage
get_survival_stat_for_group(y, dead, requested_stat)
Arguments
y |
Numeric vector of survival times. |
dead |
Integer vector of event indicators (1=event, 0=censored). |
requested_stat |
A string, either "median" or "restricted_mean". |
Value
The calculated statistic.
Get the default warm-start dispatch policy
Description
Returns EDI's built-in policy table for choosing whether warm starts
(reusing a previous fit's parameters/curvature to seed the next fit, e.g.
across bootstrap or randomization replicates) are enabled for a given
inference class during a given resampling or simulation operation
(one of "jackknife", "non_param_boot", "bayesian_boot",
"param_boot", or "rand").
Usage
get_warm_start_dispatch_policy()
Details
The internal (non-exported) dispatcher,
edi_warm_start_dispatch_policy(inference_class, operation, n),
consults this table's default plus, per operation, two override
layers: inference_class_overrides (a sample-size-independent
pattern table, as before) and n_conditioned_overrides (a list of
list(pattern, value, n_min, n_max) rules — each pattern only
applies when the current sample size n falls in
[n_min, n_max) — encoding empirical findings like "the extra
bookkeeping only pays off once resampling is expensive enough per
replicate, so disable below n=200/500/1000 for these families"). Both
layers are returned by this function and are both reachable via
set_warm_start_dispatch_policy.
Value
A named list with default (logical, TRUE in the
built-in policy) and one component per operation (jackknife,
non_param_boot, bayesian_boot, param_boot,
rand), each itself a list containing inference_class_overrides
(a named logical vector: regular-expression pattern names to
TRUE/FALSE values, matched against the inference class name,
applied regardless of sample size) and n_conditioned_overrides (a
list of list(pattern, value, n_min, n_max) rules, applied only when
n is supplied and falls in [n_min, n_max)); see Details.
Like the cold-start table, every TRUE/FALSE entry here
(including the n-thresholds) is an empirical performance judgment
computed on the maintainer's machine, not a correctness fact. Run
tune_EDI_for_this_machine to re-measure this axis,
per resampling operation and sample size, on your own machine.
See Also
get_cold_start_dispatch_policy for the analogous
(simpler, single-layer) policy governing the initial cold-start heuristic
rather than cross-replicate warm-starting; set_warm_start_dispatch_policy
to override this table at runtime; tune_EDI_for_this_machine
to re-benchmark it on your own hardware.
Examples
get_warm_start_dispatch_policy()
Compute Weibull Regression Hessian (General Censoring)
Description
Hessian matrix for the Weibull AFT log-likelihood extended
to left-, right-, and interval-censored responses. See
get_weibull_regression_general_score_cpp() for the input
convention.
Usage
get_weibull_regression_general_hessian_cpp(X, y, y_L, y_R, params)
Arguments
X |
A numeric matrix of predictors. |
y |
Exact survival times, |
y_L |
Censored-interval lower bounds, |
y_R |
Censored-interval upper bounds, |
params |
A numeric vector of parameters [beta, log_sigma]. |
Value
A numeric matrix representing the Hessian.
Compute Weibull Regression Score (General Censoring)
Description
Score vector for the Weibull AFT log-likelihood extended to
left-, right-, and interval-censored responses (see TODO-3 in
interval_censored_survival_response.md). Exactly one of y[i] or
(y_L[i], y_R[i]) must be finite per subject: NA in
the unused slot(s).
Usage
get_weibull_regression_general_score_cpp(X, y, y_L, y_R, params)
Arguments
X |
A numeric matrix of predictors. |
y |
Exact survival times, |
y_L |
Censored-interval lower bounds, |
y_R |
Censored-interval upper bounds, |
params |
A numeric vector of parameters [beta, log_sigma]. |
Value
A numeric vector representing the score.
Compute Weibull Regression Hessian
Description
Calculates the Hessian matrix (second derivatives of the log-likelihood) for a Weibull AFT regression model.
Usage
get_weibull_regression_hessian_cpp(X, y, dead, params)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of survival times. |
dead |
A numeric vector of event indicators. |
params |
A numeric vector of parameters [beta, log_sigma]. |
Value
A numeric matrix representing the Hessian.
Compute Weibull Regression Score
Description
Calculates the score vector (gradient of the log-likelihood) for a Weibull AFT regression model.
Usage
get_weibull_regression_score_cpp(X, y, dead, params)
Arguments
X |
A numeric matrix of predictors. |
y |
A numeric vector of survival times. |
dead |
A numeric vector of event indicators. |
params |
A numeric vector of parameters [beta, log_sigma]. |
Value
A numeric vector representing the score.
Generic aliased overrides for the incidence g-computation component
Description
Uses the shared randomization two-sided p-value contract; see
InferenceRand.
Uses the shared Wald testing-type contract; see
InferenceAsymp. 'Wald' is not composed
directly (its private 'get_standard_error' clashes with
‘IncidenceGComputation'’s public 'get_standard_error', an R6-forbidden
same-name-both-slots collision), so only this piece is aliased.
Usage
incidence_gcomp_generic_alias_overrides
Details
Shared self$-aliased public overrides for methods of the IncidenceGComputation component whose bodies call super$...; composed by InferenceIncidGCompRiskDiff and InferenceIncidGCompRiskRatio.
Short display label for an inference class
Description
Short display label for an inference class
Usage
inference_class_short_label(name, tau = NA_real_)
Arguments
name |
Inference class name (character), e.g. |
tau |
Quantile level for a quantile-regression class's display label (e.g. '"InferenceContinQuantileRegr"', '"InferenceContinKKQuantileRegrOneLik"' – every concrete class whose wordified label contains the bare word "Quantile" is tagged the '"quantile_regression_effect"' estimand, so this is safe as an unconditional word-level substitution rather than a per-class allowlist). 'NA_real_' (default) or '0.5' renders "Quantile" as "Median" (e.g. "KK Quantile Regr" -> "KK Median Regr"); any other value renders it as "Quantile (<tau*100> so the actual tested quantile is visible when more than one might be compared – mirrors ‘estimand_short_label()'’s own tau handling (per user request, 2026-08-27: same rewording, now applied to the inference *class* name, not just the estimand label). |
Inverse Logit (Logistic) Function
Description
Computes the inverse logit (standard logistic sigmoid) function,
\mathrm{logit}^{-1}(x) = 1/(1 + e^{-x}), the canonical mean function
for binomial/logistic-family models throughout this package (mapping a
linear predictor on the log-odds scale back to a probability). The result is
clamped to [\code{zero\_one\_logit\_clamp}, 1 -
\code{zero\_one\_logit\_clamp}] before being returned, so an extreme
x (e.g. from a poorly identified or diverging fit) cannot produce an
exact 0 or 1 probability that would later cause a -Inf/
NaN when log-transformed downstream (e.g. in a log-likelihood).
Usage
inv_logit(x, zero_one_logit_clamp = .Machine$double.eps)
Arguments
x |
Any real number (or vector), typically a fitted linear predictor
|
zero_one_logit_clamp |
The clamping distance from the 0/1 boundaries
applied to the result. Default |
Value
The inverse-logit-transformed value(s), in (0, 1), the same
length as x.
See Also
logit for the forward transform.
Examples
inv_logit(0)
Logit (Log-Odds) Transform
Description
Computes the logit (log-odds) function \mathrm{logit}(p) = \log\left(p /
(1-p)\right), the canonical link function for binomial/logistic-family
models throughout this package. p is first clamped to
[\code{zero\_one\_logit\_clamp}, 1 - \code{zero\_one\_logit\_clamp}]
before transforming, so exact 0 or 1 inputs (which would
otherwise map to -\infty/\infty) instead return a large but
finite value; this is what lets proportion/fractional responses with mass
exactly at the boundary be used as pseudo-continuous inputs to logit-scale
machinery elsewhere in the package (e.g. fast_ols_cpp-backed
logit-transform-then-OLS shortcuts) without producing non-finite values.
Usage
logit(p, zero_one_logit_clamp = .Machine$double.eps)
Arguments
p |
The value(s) to transform, nominally in |
zero_one_logit_clamp |
The clamping distance from the 0/1 boundaries
applied to |
Value
The logit-transformed value(s) as a real number (or vector), the same
length as p.
See Also
inv_logit for the inverse transform.
Examples
logit(0.25)
LRT confidence interval by Newton-Raphson + bisection (Rcpp implementation)
Description
Implements the bracket search + NR+bisection loop entirely in C++, calling
back into R only for fit_null_fn, neg_loglik_fn, and
score_fn. The derivative dp/d\delta = 2 f_{\chi^2}(T) \cdot
\mathrm{score}[j] follows from the envelope theorem.
Usage
lrt_ci_nr_cpp(
fit_null_fn,
neg_loglik_fn,
score_fn,
est,
full_negloglik,
alpha,
step,
lower_seed,
upper_seed,
j,
max_bracket = 60L,
max_nr_iter = 25L,
tol_p = 1e-07,
tol_bracket = 1e-08
)
Arguments
fit_null_fn |
R function |
neg_loglik_fn |
R function |
score_fn |
R function |
est |
Point estimate of the treatment effect |
full_negloglik |
Negative log-likelihood of the unrestricted model |
alpha |
Significance level (e.g. 0.05) |
step |
Initial step size for exponential bracket search |
lower_seed |
Initial lower-bound candidate, typically the Wald lower CI |
upper_seed |
Initial upper-bound candidate, typically the Wald upper CI |
j |
1-indexed position of the treatment coefficient in the score vector |
max_bracket |
Maximum exponential bracket search iterations (default 60) |
max_nr_iter |
Maximum NR+bisection iterations per bound (default 25) |
tol_p |
P-value convergence tolerance (default 1e-7) |
tol_bracket |
Bracket-width convergence tolerance (default 1e-8) |
Value
Unnamed numeric vector of length 2: [lower_bound, upper_bound]
Fast Mean Calculation
Description
Calculates the mean of a numeric vector using Rcpp for speed.
Usage
mean_cpp(x)
Arguments
x |
A numeric vector. |
Value
The mean of the vector.
Miettinen-Nurminen Confidence Interval for Risk Difference
Description
Computes an approximate Miettinen-Nurminen confidence interval for the risk difference by inverting the score test with a bisection search.
Usage
mn_ci_cpp(x_t, n_t, x_c, n_c, p_t_obs, p_c_obs, alpha, pval_epsilon)
Arguments
x_t |
Number of events in treatment. |
n_t |
Number of subjects in treatment. |
x_c |
Number of events in control. |
n_c |
Number of subjects in control. |
p_t_obs |
Observed treatment-arm risk. |
p_c_obs |
Observed control-arm risk. |
alpha |
The confidence level is |
pval_epsilon |
Bisection tolerance in p-value space. |
Value
A length-2 numeric vector containing the lower and upper CI bounds.
Constrained MLE for Risk Difference (Miettinen-Nurminen)
Description
Solves the likelihood equations for p_C subject to p_T - p_C = delta. Uses bisection on the score function (derivative of log-likelihood) which is monotonic and well-behaved.
Usage
mn_constrained_mle_pc_cpp(x_t, n_t, x_c, n_c, delta)
Export of C++ function mn_pvalue_cpp
Description
Evaluates the two-sided asymptotic p-value for the Miettinen-Nurminen
restricted-maximum-likelihood score test of H_0: p_T - p_C = \delta in
two independent binomial samples (see
InferenceIncidMiettinenNurminenRiskDiff
for the class that consumes this function). Internally: mn_z_statistic_cpp
computes the score z statistic using the constrained MLEs
\tilde p_C, \tilde p_T = \tilde p_C + \delta
(mn_constrained_mle_pc_cpp, found by bisecting the constrained
score equation to zero) in place of the unconstrained sample proportions in
the variance formula, with a small-sample correction factor (n_T+n_C)/(n_T+n_C-1)
applied to the naive binomial variance; this function then returns
2\,\Phi(-|z|), the two-sided normal-tail p-value. Returns NA if
either arm is empty, \delta is outside (-1, 1), or the resulting
z is not finite (e.g. the constrained variance estimate is
non-positive).
Usage
mn_pvalue_cpp(x_t, n_t, x_c, n_c, delta, p_t_obs, p_c_obs)
Arguments
x_t |
Number of events in treatment. |
n_t |
Number of subjects in treatment. |
x_c |
Number of events in control. |
n_c |
Number of subjects in control. |
delta |
Null risk difference. |
p_t_obs |
Observed treatment-arm risk. |
p_c_obs |
Observed control-arm risk. |
Value
The two-sided p-value.
Miettinen-Nurminen Score Z Statistic for Risk Difference
Description
Computes the score-style z statistic for testing the null risk difference
p_T - p_C = delta in two independent binomial samples, using the
Miettinen-Nurminen constrained nuisance estimates.
Usage
mn_z_statistic_cpp(x_t, n_t, x_c, n_c, delta, p_t_obs, p_c_obs)
Arguments
x_t |
Number of events in treatment. |
n_t |
Number of subjects in treatment. |
x_c |
Number of events in control. |
n_c |
Number of subjects in control. |
delta |
Null risk difference. |
p_t_obs |
Observed treatment-arm risk. |
p_c_obs |
Observed control-arm risk. |
Value
The asymptotic z statistic.
Export of C++ function newcombe_independent_ci_cpp
Description
Computes Newcombe's "Method 10" hybrid confidence interval for the
difference between two independent proportions p_1 - p_2
(see
InferenceIncidNewcombeRiskDiff
for the class that consumes this function). Separate Wilson score intervals
[\ell_1, u_1] and [\ell_2, u_2] are computed for each proportion
individually (via wilson_score_interval_cpp), then combined as
\left[\,(p_1-p_2) - \sqrt{(p_1-\ell_1)^2 + (u_2-p_2)^2},\ \ (p_1-p_2) +
\sqrt{(u_1-p_1)^2 + (p_2-\ell_2)^2}\,\right],
clamped to [-1, 1]. This avoids the boundary/coverage problems of the
naive normal-approximation (Wald) interval on a risk difference while
remaining closed-form (no iterative score-test inversion). Returns
c(NA, NA) if either sample size is non-positive.
Usage
newcombe_independent_ci_cpp(x1, n1, x2, n2, alpha)
Arguments
x1 |
Number of events in group 1. |
n1 |
Number of subjects in group 1. |
x2 |
Number of events in group 2. |
n2 |
Number of subjects in group 2. |
alpha |
The confidence level is |
Value
A length-2 numeric vector containing the lower and upper CI bounds
for p_1 - p_2.
References
Newcombe, R. G. (1998). "Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods." Statistics in Medicine, 17(8), 873-890, doi:10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I.
See Also
newcombe_paired_ci_cpp for the matched-pair
generalization of this same hybrid-score method.
Newcombe Hybrid Score Interval for Paired Proportions
Description
Newcombe Hybrid Score Interval for Paired Proportions
Usage
newcombe_paired_ci_cpp(n11, n10, n01, n00, alpha)
Export of C++ function ols_hc2_post_fit_cpp
Description
Given an already-fitted ordinary least squares model, computes the
HC2 heteroskedasticity-consistent sandwich covariance matrix
(MacKinnon-White) for the coefficients: \widehat{\mathrm{Var}}(\hat\beta)
= B\,M\,B, with "bread" B = (X^\top X)^{-1} and "meat" M =
X^\top \mathrm{diag}(\omega_i) X, where \omega_i = r_i^2 / (1 - h_{ii})
— the squared OLS residual r_i leverage-corrected by dividing
by 1 - h_{ii} (h_{ii} the i-th diagonal of the OLS hat
matrix X(X^\top X)^{-1}X^\top), unlike the plain HC0 sandwich
(gcomp_logistic_post_fit_cpp's logistic analogue, or this
function's own uncorrected r_i^2 meat) which does not correct for
leverage. HC2 is unbiased under homoskedasticity for balanced designs and
generally has better small-sample properties than HC0/HC1 when leverage is
uneven. Internally, this function first computes the (design-only) "setup"
quantities bread/hat via ols_hc2_setup_cpp, then calls
ols_hc2_post_fit_precomputed_cpp; callers who already have those
precomputed (e.g. across repeated resampling on the same fixed design) can
call the precomputed variant directly instead to skip recomputing the
(X^\top X)^{-1} bread and leverage each time.
Usage
ols_hc2_post_fit_cpp(X_fit, y, coef_hat, j_treat)
Arguments
X_fit |
A numeric matrix of predictors, as used to fit the model. |
y |
A numeric vector of responses. |
coef_hat |
A numeric vector of fitted OLS coefficients |
j_treat |
1-based column index of the treatment indicator in |
Value
A list with components beta_hat (\hat\beta_{j_{\mathrm{treat}}}),
ssq_hat (its HC2 variance), se (its HC2 standard error),
vcov (the full p \times p HC2 covariance matrix), std_err
(per-coefficient HC2 standard errors), and z_vals (per-coefficient
Wald z-statistics, \hat\beta_j / \widehat{\mathrm{SE}}(\hat\beta_j)).
References
MacKinnon, J. G., and White, H. (1985). "Some Heteroskedasticity-Consistent Covariance Matrix Estimators with Improved Finite Sample Properties." Journal of Econometrics, 29(3), 305-325, doi:10.1016/0304-4076(85)90158-7, for the HC2 estimator used here.
Fast G-Computation (Standardization) Point Estimate and Model-Based Inference for a Proportional-Odds Ordinal Model (C++)
Description
Computes the same standardized (G-computation) marginal difference in
expected ordinal category score as
gcomp_ordinal_proportional_odds_post_fit_cpp — under a fitted
proportional-odds (cumulative logit) model \mathrm{logit}\,\Pr(Y_i \le k
\mid x_i) = \alpha_k - x_i^\top\beta, k = 1, \ldots, K-1 — but
additionally supplies model-based inferential quantities that the
other function omits. The fitted linear predictor for each subject is
recomputed with the treatment column (j_treat) forced to 1
(\eta_{1,i}) and to 0 (\eta_{0,i}), the standardized expected
scores mean1, mean0, and their difference md are formed
exactly as in gcomp_ordinal_proportional_odds_post_fit_cpp, and
then:
An ordinal-regression log-likelihood Hessian is evaluated at
[\hat\alpha, \hat\beta]and inverted (with a symmetrization step and a finiteness check) to give the model-based (non-sandwich) variance-covariance matrix of the full parameter vector;vcov/std_err/z_valsreport the\hat\beta-block of this covariance.The delta-method standard error of
md,se_md, is obtained by numerically differentiatingmd(central differences, step10^{-5}) with respect to every element of[\alpha,\beta]to get a gradientg, then computing\sqrt{g^\top V g}whereVis the full parameter covariance above;NAif any perturbedmdevaluation is non-finite (e.g. a perturbed\alphaviolates monotonicity) or the resulting variance is negative/non-finite.
Unlike the sandwich-based post-fit helpers elsewhere in this package (e.g.
gcomp_logistic_post_fit_cpp), the covariance here comes purely
from the model's observed information and does not attempt to be robust to
misspecification.
Usage
ordinal_gcomp_post_fit_cpp(X_fit, y, coef_hat, alpha_hat, j_treat)
Arguments
X_fit |
Numeric matrix of predictors used to fit the model. |
y |
Numeric vector of the ordinal responses used to fit the model
(needed to reconstruct the |
coef_hat |
Numeric vector of fitted proportional-odds regression
coefficients |
alpha_hat |
Numeric vector of the |
j_treat |
1-based column index of the treatment indicator in |
Value
A list with elements vcov (model-based covariance of
\hat\beta), std_err, z_vals (both length p, one
per column of X_fit), mean1, mean0, md
(mean1 - mean0), and se_md (delta-method SE of md).
Errors (via stop()) if j_treat is out of bounds, the
dimensions of y/coef_hat are inconsistent with X_fit,
the Hessian is not invertible, or the resulting covariance has any
non-finite entry.
See Also
gcomp_ordinal_proportional_odds_post_fit_cpp for the
point-estimate-only variant (no y, no Hessian, no inference) that
this function's point estimates match; fast_ordinal_regression_cpp
for the fitting routine that produces coef_hat/alpha_hat.
Pocock-Simon Covariate-Adaptive Minimization: Assign and Update Counts (C++)
Description
The stateful wrapper around pocock_simon_assign_cpp actually
used to drive a one-subject-at-a-time Pocock-Simon minimization design: it
computes the next assignment using the identical imbalance-score/biased-coin
logic documented on pocock_simon_assign_cpp (see that page for
the full model), then mutates counts in place, incrementing,
for every covariate in subject_levels_idx, the count of the assigned
arm at that covariate's level — so the running covariate-by-arm counts stay
correct for the next subject's assignment.
Usage
pocock_simon_assign_and_update_cpp(
counts,
subject_levels_idx,
weights,
p_best,
prob_T
)
Arguments
counts |
A numeric matrix, one row per stratification-covariate level and one column per treatment arm (2 columns); modified in place to record the new assignment. |
subject_levels_idx |
An integer vector (1-indexed) giving, for the
subject being assigned, which row of |
weights |
A numeric vector, parallel to |
p_best |
The probability of assigning the arm that minimizes the combined imbalance score (the biased coin). |
prob_T |
The Bernoulli probability used to break an exact imbalance tie. |
Value
The assigned treatment arm, 0 or 1. As a side effect,
counts is incremented at row subject_levels_idx[j], column
w (the returned assignment) for every covariate j.
Note
Reproducibility notes: see pocock_simon_assign_cpp.
References
Pocock, S. J. and Simon, R. (1975). "Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial." Biometrics, 31(1), 103-115.
See Also
pocock_simon_assign_cpp for the underlying
non-mutating decision logic and the full imbalance-score model.
Pocock-Simon Covariate-Adaptive Minimization: Assignment Decision (C++)
Description
Decides the next subject's treatment assignment under Pocock and Simon's
(1975) covariate-adaptive minimization algorithm, without modifying any
state (see pocock_simon_assign_and_update_cpp for the
state-updating wrapper actually used by the stepwise design). counts
stacks, for every level of every stratification covariate, the number of
previously-assigned subjects at that level currently in each treatment arm
(one row per covariate level, one column per arm — the design supports
exactly 2 arms). The subject to be assigned belongs to one level per
covariate, given by subject_levels_idx (1-indexed rows into
counts).
Usage
pocock_simon_assign_cpp(counts, subject_levels_idx, weights, p_best, prob_T)
Arguments
counts |
A numeric matrix, one row per stratification-covariate level (across all covariates, stacked) and one column per treatment arm (2 columns), of counts of subjects previously assigned to that level/arm combination. |
subject_levels_idx |
An integer vector (1-indexed, length = number of
stratification covariates) giving, for the subject being assigned, which
row of |
weights |
A numeric vector, parallel to |
p_best |
The probability of assigning the arm that minimizes |
prob_T |
The Bernoulli probability used to break an exact tie in
|
Details
For each candidate arm k \in \{0,1\}, a marginal imbalance score is
computed as the weighted sum, over the subject's covariates j, of the
variance across arms of the level's counts if the subject
were assigned to arm k:
G_k = \sum_j w_j \, \mathrm{Var}_t\big(n_{j,t} + \mathbb{1}\{t=k\}\big),
where n_{j,t} is the current count for the subject's level of
covariate j in arm t, and the variance is taken over the (here,
2) arms with the usual n-1 divisor — so G_k is smallest for the
arm that would leave covariate margins most balanced. The arm with the
smaller G_k (best_trt) is then assigned with a biased-coin
probability p_best (and the other arm with probability
1 - p_best); if G_0 = G_1 exactly (perfect tie), the assignment
is instead a single Bernoulli(prob_T) draw for arm 1. Setting
p_best = 1 recovers Taves' deterministic minimization; p_best
strictly between 0.5 and 1 (Pocock and Simon's recommendation, e.g. 0.75-0.85)
retains most of minimization's balancing power while preserving some
unpredictability of the next assignment.
Value
The assigned treatment arm, 0 or 1.
Note
Seeded from one R::unif_rand() draw into edi_rng::RRng (RNG.h), a portable re-implementation of R's own Mersenne-Twister generator – a given seed therefore produces identical draws in R and in any future binding (e.g. Python) using the same core and the same seed, even though this call does not continue R's own live session stream bit- for-bit (see pocock_simon_redraw_w_cpp for the one function in this file where that distinction matters and is handled).
References
Pocock, S. J. and Simon, R. (1975). "Sequential Treatment Assignment with Balancing for Prognostic Factors in the Controlled Clinical Trial." Biometrics, 31(1), 103-115.
See Also
pocock_simon_assign_and_update_cpp for the
assign-and-mutate-counts wrapper; pocock_simon_redraw_w_cpp
for the batch re-derivation of a whole assignment sequence from scratch.
Pocock-Simon Covariate-Adaptive Minimization: Batch Redraw of a Whole Assignment Sequence (C++)
Description
Replays the same Pocock-Simon minimization decision logic as
pocock_simon_assign_cpp (see that page for the imbalance-score
and biased-coin model) over an entire sequence of n subjects in one
call, starting from empty covariate-by-arm counts (all zero) and updating
them internally after each subject, rather than being called once per
subject with externally-maintained counts. This is used to
re-derive (“redraw”) a full sequence of assignments in a
single vectorized pass — e.g. for simulation, or for reconstructing what a
one-at-a-time run would have produced — while consuming R's uniform random
stream in exactly the same order a subject-by-subject loop calling
pocock_simon_assign_cpp would have.
Usage
pocock_simon_redraw_w_cpp(
x_levels_matrix,
num_levels_total,
weights,
p_best,
prob_T
)
Arguments
x_levels_matrix |
An integer matrix with one row per subject and one
column per stratification covariate; entry |
num_levels_total |
The total number of distinct covariate levels across all covariates (i.e. the number of rows the internal counts table has). |
weights |
A numeric vector, one per covariate column of
|
p_best |
The probability of assigning the arm that minimizes the combined imbalance score (the biased coin). |
prob_T |
The Bernoulli probability used to break an exact imbalance tie. |
Value
An integer vector of length n, the treatment arm (0 or
1) assigned to each subject in row order of x_levels_matrix.
Note
Continues R's actual live .Random.seed stream via edi_rng::RRng (RNG.h) rather than seeding independently – output is bit-identical to what calling R's own unif_rand() directly, in this same loop, would have produced (verified against a pure-R reference implementation in test-pocock-simon-redraw-buffers.R). Requires RNGkind ("Mersenne-Twister", "Inversion"), R's default.
Prints the results table from an InferenceSuite
run_all_inference() call – the same table screen = TRUE prints during
the call itself, so a user who assigned the return value and later types its name
(or calls print() on it) sees a readable table rather than a raw nested list
dump. The table itself is rendered by
run_all_inference_format_pretty_table(): rows sorted by
estimand, with a double rule under the header and a single rule
between estimand groups and at the bottom, class names and
estimand values shortened for display (never the underlying
results_table values), and a cov_model letter-key legend
appended when applicable – see that function's own documentation for
the exact column-by-column rendering rules.
Description
Prints the results table from an InferenceSuite
run_all_inference() call – the same table screen = TRUE prints during
the call itself, so a user who assigned the return value and later types its name
(or calls print() on it) sees a readable table rather than a raw nested list
dump. The table itself is rendered by
run_all_inference_format_pretty_table(): rows sorted by
estimand, with a double rule under the header and a single rule
between estimand groups and at the bottom, class names and
estimand values shortened for display (never the underlying
results_table values), and a cov_model letter-key legend
appended when applicable – see that function's own documentation for
the exact column-by-column rendering rules.
Usage
## S3 method for class 'EDIInferenceSuiteResults'
print(x, ...)
Arguments
x |
An |
... |
Ignored; present for S3 consistency with the generic. |
Value
x, invisibly.
CI inversion by p-value bracket search + bisection (Rcpp implementation)
Description
Used by invert_test_pval_confidence_interval (score CI and any other
CI that inverts a scalar p-value function with no derivative available).
The Wald bounds are tried first as bracket candidates before falling back
to exponential search. Bisection then polishes to tol in delta-space.
Usage
pval_invert_ci_cpp(
pval_fn,
est,
alpha,
step,
lower_seed,
upper_seed,
max_bracket = 60L,
max_bisect = 60L,
tol = 1e-06
)
Arguments
pval_fn |
R function |
est |
Point estimate of the treatment effect |
alpha |
Significance level (e.g. 0.05) |
step |
Initial step size for exponential bracket search |
lower_seed |
Wald lower CI bound (used as first bracket candidate;
pass |
upper_seed |
Wald upper CI bound (same) |
max_bracket |
Maximum exponential bracket search iterations (default 60) |
max_bisect |
Maximum bisection iterations (default 60) |
tol |
Bracket-width convergence tolerance in delta-space (default 1e-6) |
Value
Unnamed numeric vector of length 2: [lower_bound, upper_bound]
Robust Negative Binomial Regression with Backward Column-Dropping Fallback
Description
Fits a negative-binomial GLM via glm.nb (log link, joint
ML estimation of the regression coefficients and the dispersion parameter
\theta), falling back to a smaller model when the fit throws an error
(typically non-convergence of \theta, or a singular design). On each
failure, the last column of data_obj is dropped and the fit is
retried against the same form_obj (which must resolve to y ~ .
or similar so that its right-hand side tracks the shrinking column set); this
repeats until a fit succeeds or every predictor column has been removed, at
which point NA is returned. Because columns are dropped strictly from
the right, callers should order data_obj's columns from most to least
important a priori, or accept that this is a best-effort robustness
measure rather than a principled model-selection procedure.
Usage
robust_negbinreg(form_obj, data_obj)
Arguments
form_obj |
The model formula, typically |
data_obj |
The data frame to run negative-binomial regression on; its last column is dropped on each retry, in order, until a fit converges or no columns remain. |
Value
The fitted glm.nb model object, or NA if no column subset
(down to and including the response alone) produced a successful fit.
Examples
dat = data.frame(y = rpois(10, 2), x1 = rnorm(10), x2 = rnorm(10))
robust_negbinreg(y ~ ., dat)
Robust Parametric Survival Regression from Response/Censoring Vectors
Description
Convenience wrapper around robust_survreg_with_surv_object that
builds the Surv object from separate response and
censoring vectors first. See that function for the full description of the
warm-start-then-random-restart fitting strategy used to make
survreg converge reliably even from poor or
near-singular starting points.
Usage
robust_survreg(
y,
dead,
cov_matrix_or_vector,
dist = "weibull",
num_max_iter = 50
)
Arguments
y |
The (possibly right-censored) response vector (event/censoring time). |
dead |
The event indicator (1 if the event was observed/uncensored, 0 if
right-censored at |
cov_matrix_or_vector |
The design matrix (or a single covariate vector) of predictors,
excluding the intercept (one is added by the internal |
dist |
The parametric AFT distribution family passed to
|
num_max_iter |
Maximum number of random-restart attempts if the direct fit fails
or does not converge (default 50); see |
Value
The fitted survreg model object, or NULL if no
attempt converged to a fit with no NA coefficients within num_max_iter tries.
Examples
X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
robust_survreg(y, dead, X)
Robust Parametric Survival Regression (AFT) with Warm-Start and Random-Restart Fallback
Description
Fits a parametric accelerated-failure-time (AFT) survival regression via
survreg on surv_object ~ . over the columns of
cov_matrix_or_vector, with two layers of robustness against
survreg's well-known sensitivity to starting values and
near-collinear design matrices:
-
Preprocessing: near-collinear columns of the design matrix are dropped first via
drop_highly_correlated_colsthendrop_linearly_dependent_cols, before any fitting is attempted. -
Warm start (Weibull only): when
dist = "weibull", a fast closed-form-gradient Weibull fit (fast_weibull_regression) is attempted first; if it succeeds and returns a finite log-likelihood, its coefficients and\log(\hat\sigma)are passed tosurvregas theinitvector, which typically converges the true MLE in a singlesurvregcall. If this warm-started fit is unavailable, fails, or producesNAcoefficients, fitting falls through to the general random-restart loop below (for all otherdistvalues, this warm start is skipped entirely). -
Random-restart loop: starting from an all-zero
initvector,survregis called repeatedly (perturbinginitby an independent standard-normal jitter,init + rnorm(length(init)), after every failed attempt) until a fit with noNAcoefficients is obtained ornum_max_iterattempts are exhausted, at which pointNULLis returned.
survreg.control(maxiter = 100, rel.tolerance = 1e-9, outer.max = 10) is
used throughout (tighter than survreg's own defaults) to reduce the
chance of a spuriously "converged" fit at a poor optimum.
Usage
robust_survreg_with_surv_object(
surv_object,
cov_matrix_or_vector,
dist = "weibull",
num_max_iter = 50
)
Arguments
surv_object |
The survival object (built from the response vector
and censoring vector via |
cov_matrix_or_vector |
The design matrix (or a single covariate vector) of predictors,
excluding the intercept (one is added by the internal |
dist |
The parametric AFT distribution family passed to
|
num_max_iter |
Maximum number of random-restart attempts if the (possibly warm-started) direct fit fails or does not converge (default 50). |
Value
The fitted survreg model object, or NULL if no
attempt converged to a fit with no NA coefficients within num_max_iter tries.
Examples
X = matrix(rnorm(500), 100, 5)
y = runif(100)
dead = rbinom(100, 1, 0.5)
surv = survival::Surv(y, dead)
robust_survreg_with_surv_object(surv, X)
Sample Mode
Description
Thin R wrapper around sample_mode_cpp(), which returns the most
frequently occurring value in data. Integer, logical, double,
character, and factor vectors are all supported (dispatched internally on
TYPEOF(data); factors preserve their class/levels
attributes on the returned value). NA (and, for doubles, NaN
as a category distinct from NA) is counted like any other value and
can itself be returned as "the mode" if it is the most frequent entry.
Ties are broken by first occurrence: among values tied for the
highest count, the one that appears earliest in data is returned —
this is a positional, not a numeric/lexicographic, tie-break rule.
Usage
sample_mode(data)
Arguments
data |
A vector (integer, logical, double, character, or factor) to compute the mode of. |
Value
A length-1 vector (same type as data) holding the most frequently occurring
value, with ties broken in favor of whichever tied value occurs earliest in data.
Examples
sample_mode(c(1, 2, 2, 3))
Update the cold-start dispatch policy
Description
Overrides, queries, or resets the runtime policy consulted by
edi_cold_start_dispatch_policy() for whether an inference class's C++
fitting backend defaults to an OLS-based smart_cold_start or a plain
zero-vector start; see get_cold_start_dispatch_policy for the
built-in default table and the rationale behind it.
Usage
set_cold_start_dispatch_policy(policy = NULL, reset = FALSE)
Arguments
policy |
Either |
reset |
If |
Details
Call with no arguments (policy = NULL, reset = FALSE)
to retrieve the current configuration without changing it. Pass a named
list to policy to merge new/overriding entries into the current
configuration via modifyList (so default and/or
inference_class_overrides can each be supplied independently, and
entries not mentioned are left untouched — this cannot remove an
existing override pattern, only add or replace one). Pass reset =
TRUE to discard any accumulated overrides and restore the package's
built-in default policy exactly as returned by
get_cold_start_dispatch_policy.
Value
Invisible NULL when policy is supplied (a mutation), or
invisibly the current policy configuration list when called for its
side-effect-free query/reset value.
See Also
get_cold_start_dispatch_policy for the policy schema
and built-in defaults; set_warm_start_dispatch_policy and
set_optimization_dispatch_policy for the analogous setters
governing warm-starting and optimizer-algorithm choice;
tune_EDI_for_this_machine, which calls this setter with
machine-measured overrides rather than hand-picked ones.
Examples
set_cold_start_dispatch_policy(reset = TRUE)
Set the number of cores for parallelization
Description
This function initializes a persistent parallel cluster (either a fork cluster on Unix-like systems or a mirai cluster on others) to be used by all Design and Inference objects. This avoids the overhead of creating clusters repeatedly.
Usage
set_num_cores(num_cores, force_mirai = FALSE)
Arguments
num_cores |
Integer number of worker processes to make available. |
force_mirai |
If |
Details
set_num_cores() sets a global upper bound for parallel work. It does
not guarantee that every inference routine will use all requested workers.
EDI's inference dispatcher applies a blocklist-first heuristic informed by
package benchmarks: workloads that have shown consistent multicore slowdowns
are forced to run serially, while the remaining workloads are allowed to use
their method-specific warmup heuristics and native thread caps.
The default forced-serial blocklist covers incidence randomization confidence intervals, bootstrap for non-regression KK Wilcoxon inference, bootstrap for non-KK survival procedures, and bootstrap for incidence procedures. Do not expect a universal "more cores is faster" rule.
If you want to change the default policy, use
set_parallel_dispatch_policy().
Value
Invisible NULL.
Examples
set_num_cores(2)
unset_num_cores()
Set the number of threads for OpenMP, Eigen, and MKL
Description
Set the number of threads for OpenMP, Eigen, and MKL
Usage
set_omp_num_threads_cpp(n_threads)
Arguments
n_threads |
Integer. |
Update the optimization dispatch policy
Description
Overrides, queries, or resets the runtime policy consulted by
edi_optimization_dispatch_policy() for which optimization algorithm
("newton_raphson", "lbfgs", or "irls") an inference
class's C++ model-fitting backend uses by default; see
get_optimization_dispatch_policy for the built-in default
table and the empirical rationale behind its per-family choices.
Usage
set_optimization_dispatch_policy(policy = NULL, reset = FALSE)
Arguments
policy |
Either |
reset |
If |
Details
Call with no arguments (policy = NULL, reset = FALSE)
to retrieve the current configuration without changing it. Pass a named
list to policy to merge new/overriding entries into the current
configuration via modifyList (default_alg
and/or inference_class_overrides can each be supplied
independently). Pass reset = TRUE to discard any accumulated
overrides and restore the package's built-in default policy exactly as
returned by get_optimization_dispatch_policy. Note that this
only changes the default algorithm consulted when a fitting call
does not itself specify optimization_alg explicitly; an explicit
per-call argument still takes precedence.
Value
Invisible NULL when policy is supplied (a mutation), or
invisibly the current policy configuration list when called for its
side-effect-free query/reset value.
See Also
get_optimization_dispatch_policy for the policy
schema and built-in defaults; set_cold_start_dispatch_policy
and set_warm_start_dispatch_policy for the analogous setters
governing cold-start and warm-start behavior;
tune_EDI_for_this_machine, which calls this setter with
machine-measured, convergence-checked overrides rather than
hand-picked ones.
Examples
set_optimization_dispatch_policy(reset = TRUE)
Update the parallel dispatch policy
Description
EDI uses an empirical, blocklist-first dispatch policy to decide when an
inference routine's "bootstrap" or "rand_ci" resampling should
be forced serial even if multiple cores are available (see
get_parallel_dispatch_policy for the built-in table and the
PCRE pattern/response-type matching rules it encodes). This function lets
the user update that policy at runtime without editing package internals.
Usage
set_parallel_dispatch_policy(policy = NULL, reset = FALSE)
Arguments
policy |
Either |
reset |
If |
Details
The policy can be updated in two mutually exclusive ways:
Pass a named list to
policyto merge with the current policy configuration viamodifyList, one section at a time. Supported top-level keys arebootstrapandrand_ci, and each key's value is itself a list that may containserial_inference_class_patterns(character vector of PCRE regular expressions) and/orserial_response_types(character vector of exact response-type strings); an unrecognized top-level key raises an error rather than being silently ignored.Pass a custom function with signature
function(inference_class, response_type, operation)topolicyto replace the entire pattern-table dispatch logic with your own rule; it must return a list with at leastforce_serial(logical) andreason(character). Supplying a function clears any list-based overrides accumulated so far and is stored separately from the pattern-table configuration.
Use reset = TRUE to discard both the list-based overrides and any
custom function, restoring the package's built-in default policy exactly as
returned by get_parallel_dispatch_policy. Do not expect a
universal "more cores is faster" rule — this policy exists because several
resampling workloads have shown consistent multicore slowdowns in
package benchmarks.
Value
Invisible NULL when policy is supplied (a mutation), or
invisibly the current policy configuration list when called for its
side-effect-free query/reset value.
See Also
get_parallel_dispatch_policy for the policy schema and
built-in defaults (including why this table is a safety blocklist, not
a performance one); set_num_cores for setting the actual
worker-count upper bound this policy operates within;
tune_EDI_for_this_machine for the separate,
machine-dependent question of when parallel execution is
worthwhile (never applied here — this policy is only ever
overridden explicitly, by you).
Examples
set_parallel_dispatch_policy(reset = TRUE)
Update the warm-start dispatch policy
Description
Overrides, queries, or resets the runtime policy consulted by
edi_warm_start_dispatch_policy() for whether an inference class
reuses a previous fit's parameters/curvature to seed the next fit during
resampling; see get_warm_start_dispatch_policy for the
built-in default table and its jackknife/non_param_boot/
bayesian_boot/param_boot/rand operation schema
(each with an inference_class_overrides layer and an
n_conditioned_overrides layer).
Usage
set_warm_start_dispatch_policy(policy = NULL, reset = FALSE)
Arguments
policy |
Either |
reset |
If |
Details
Call with no arguments (policy = NULL, reset = FALSE)
to retrieve the current configuration without changing it. Pass a named
list to policy to merge new/overriding entries into the current
configuration via modifyList (per-operation
sub-lists, e.g. list(rand = list(inference_class_overrides = ...))
or list(rand = list(n_conditioned_overrides = ...)), are merged
rather than replaced wholesale — note n_conditioned_overrides is a
plain list of rules, so overriding it replaces the whole list for that
operation, not a per-rule merge). Pass reset = TRUE to discard any
accumulated overrides and restore the package's built-in default policy
exactly as returned by get_warm_start_dispatch_policy.
This function controls the full dispatch policy, including the
sample-size-conditioned n_conditioned_overrides layer.
Value
Invisible NULL when policy is supplied (a mutation), or
invisibly the current policy configuration list when called for its
side-effect-free query/reset value.
See Also
get_warm_start_dispatch_policy for the policy schema
and built-in defaults; set_cold_start_dispatch_policy for
the analogous, simpler single-layer setter governing the initial
cold-start heuristic; tune_EDI_for_this_machine, which
calls this setter with machine-measured overrides rather than
hand-picked ones.
Examples
set_warm_start_dispatch_policy(reset = TRUE)
Summarizes an InferenceSuite run_all_inference()
result: counts by status, the estimate range across status == "ok"
classes, and how many reject at alpha.
Description
Summarizes an InferenceSuite run_all_inference()
result: counts by status, the estimate range across status == "ok"
classes, and how many reject at alpha.
Usage
## S3 method for class 'EDIInferenceSuiteResults'
summary(object, ...)
## S3 method for class 'summary.EDIInferenceSuiteResults'
print(x, ...)
Arguments
object |
An |
... |
Ignored; present for S3 consistency with the generic. |
x |
A |
Value
An object of class summary.EDIInferenceSuiteResults, printable via its
own print method.
x, invisibly.
Lean GLM Summary (Skips Deviance Residual Quantiles)
Description
A drop-in replacement for summary.glm that produces the
identical coefficient table, dispersion estimate, and (optionally)
correlation matrix, but omits the five-number summary of the
deviance residuals (summary(object$deviance.resid)) that
summary.glm() always computes and stores in its deviance.resid
component. That residual summary is cheap for a single fit but adds up when
summarizing thousands of GLM fits in a resampling loop (e.g. bootstrap or
randomization replicates elsewhere in this package), so this function skips
it entirely; the returned object's deviance.resid component is simply
absent rather than populated, which will matter to code that calls
print.summary.glm() on the result or otherwise inspects that field.
Every other computation — dispersion estimation (Pearson X^2/\mathrm{df}
for Gaussian/Gamma/inverse-Gaussian families, fixed at 1 for
Poisson/binomial, unless dispersion is supplied explicitly), the
coefficient table (Wald z tests when dispersion is fixed/known,
t tests with df.residual degrees of freedom when dispersion is
estimated), and the optional correlation/symbolic.cor outputs,
is identical to summary.glm.
Usage
summary_glm_lean(
object,
dispersion = NULL,
correlation = FALSE,
symbolic.cor = FALSE,
...
)
Arguments
object |
A fitted |
dispersion |
The dispersion parameter for the fitting family; if |
correlation |
Logical; if |
symbolic.cor |
Logical; if |
... |
Currently unused; present only for signature compatibility with
|
Value
An object of class c("summary.glm") with the same components as
summary.glm's return value except deviance.resid,
which is not computed and is absent from the result.
See Also
summary.glm, of which this is a residual-summary-skipping variant.
Examples
fit = glm(rbinom(10, 1, 0.5) ~ rnorm(10), family = binomial)
summary_glm_lean(fit)
Toggle the execution of assertions throughout the package
Description
This function enables or disables the internal input validation checks (assertions)
for the rest of the R session. It does not modify options(); setting
options(edi.run_asserts = FALSE) yourself also disables them.
Disabling assertions can provide a significant performance boost in heavy
simulations (often 10x-20x speedup), but it removes the safety rails that
prevent invalid data from reaching the internal algorithms.
Warning: If assertions are disabled, passing malformed or invalid data to package functions may result in cryptic R errors, incorrect statistical results, or even hard system crashes (SEGFAULTs) at the C++ layer. Only disable assertions if you are certain your data is pre-validated and follows the package requirements exactly.
Usage
toggle_asserts(on = TRUE)
Arguments
on |
Logical scalar. If TRUE (default), assertions are executed. If FALSE, they are skipped. |
Transform continuous latent signal to the response type scale
Description
A helper function to transform a latent continuous signal to the scale
appropriate for a given response_type, identical to the logic used
within SimulationFramework but not used within SimulationFramework.
Usage
transform_cont_y_based_on_response_type(
y_cont,
response_type,
n_ordinal_levels = 4L,
proportion_epsilon = 1e-06,
survival_min_time = 0.1,
count_min_rate = 0L,
count_shift = 0
)
Arguments
y_cont |
Numeric vector. The latent continuous response signal. |
response_type |
Character scalar. One of |
n_ordinal_levels |
Positive integer. Number of ordinal categories when
|
proportion_epsilon |
Numeric scalar. Small value added to proportion to avoid 0 and 1. Default |
survival_min_time |
Numeric scalar. Minimum survival time and shift. Default |
count_min_rate |
Integer scalar. Minimum baseline rate for count response. Default |
count_shift |
Numeric scalar. Constant added to counts after zero-centering. Default |
Value
A numeric vector of transformed responses on the appropriate scale.
Examples
transform_cont_y_based_on_response_type(rnorm(10), 'incidence')
Benchmark this machine and tune EDI's performance-policy defaults to it
Description
Every performance-policy default shipped in EDI (whether an inference
class uses a smart cold start, whether resampling reuses warm starts and
at what sample sizes, which optimizer algorithm a family uses, and at what
sample size parallel bootstrapping starts to beat serial) was measured
empirically on the maintainer's machine. Yours differs – core count,
cache, BLAS, compiler flags – so a policy that is net-positive there can
be net-negative here, and vice versa. This function re-runs those
benchmarks on your hardware, decides the winning setting per
axis, saves the result to a per-user config file, and applies it
immediately; every later library(EDI) re-applies it. Hardware
changed? Re-run; it overwrites.
Usage
tune_EDI_for_this_machine(
effort = c("standard", "quick", "thorough"),
axes = NULL,
families = NULL,
n_grid = NULL,
reps = NULL,
num_cores_grid = NULL,
converged_fn = NULL,
quiet = FALSE,
dry_run = FALSE,
force = FALSE
)
Arguments
effort |
One of |
axes |
Which axes to tune: any subset of |
families |
Optional character vector of inference class names to
restrict every axis to; |
n_grid |
Optional integer vector of sample sizes overriding the effort tier's grid. |
reps |
Optional replicate count per timed cell overriding the effort tier's. |
num_cores_grid |
Optional integer vector of core counts (each
|
converged_fn |
|
quiet |
If |
dry_run |
If |
force |
If |
Details
What is tuned. Four axes, each against the corresponding
get_*_dispatch_policy() table: cold start
(get_cold_start_dispatch_policy), warm start per resampling
operation (get_warm_start_dispatch_policy), optimizer
algorithm (get_optimization_dispatch_policy), and the
parallel-vs-serial crossover sample size per family
(get_parallel_dispatch_policy). The bootstrap
confidence-interval type policy is a statistical-validity table,
not a performance one, and is never touched; nor are entries in the
parallel policy's serial blocklist that exist for parallel safety.
How a deviation is accepted. Per axis, per (family, sample
size) cell, both settings are timed on identical synthetic data –
interleaved (A/B/A/B) for the cold-start/warm-start/optimizer axes, and
blocked (all serial reps, then all parallel reps) for the parallel axis,
whose fork-cluster setup cost rules out per-replicate interleaving. A
candidate displaces the shipped default only if its median time is at
least 5% better and that improvement exceeds twice the candidate's
own interquartile spread; ties keep the shipped default. The optimizer
axis additionally requires the candidate to have converged on every
replicate of the cell – speed never trumps a convergence failure. Only
deviations from the shipped defaults are stored, in the exact shape
the matching set_*_dispatch_policy() setter accepts, and they are
merged into (never replacing) the shipped tables.
Progress. A single progress bar with a running
estimated-time-left is redrawn in place as each benchmark cell completes
– the same bar InferenceSuite$run_all_inference() shows.
Correctness gate. A timing win alone does not displace a shipped
default: every accepted deviation is re-fit once under both
settings on identical synthetic data and the outputs compared (point
estimates for cold start/optimizer/parallel; the resampling operation's
own output, RNG-matched, for warm start). A disagreement – or an
unverifiable comparison – discards the deviation, with a
warning() naming it; discarded deviations are available on the
returned object via attr(x, "discarded_by_correctness_gate") and
are never written to the config file or applied.
The optimizer axis and converged_fn. There is not yet a
generic, class-independent accessor on an inference object that reports
whether its fit converged, so the optimizer axis needs you to supply one
as converged_fn(inf). When axes is left NULL the
optimizer axis is included only if converged_fn is given; asking
for it explicitly without one is an error. An always-TRUE
converged_fn disables the convergence guard and is not
appropriate for a real tuning run.
The parallel axis. Runs only on Unix-alikes with at least two
logical cores (it benchmarks a real fork cluster). Its preferred core
count is recorded only – never applied at package load – you
still opt into parallelism with set_num_cores.
Idle machine. Run this on an otherwise idle machine; a tuning
run under contention measures the contention, not the hardware. Before
benchmarking, the function checks the 1-minute load average against the
core count and times a small fixed calibration operation for noise; if
either says the machine is busy it refuses to run (interactively, it
asks first) unless force = TRUE.
Value
Invisibly, an EDILocalMachineTuning object: the policy
diffs (policy_diffs), the raw per-axis deviations
(raw_deviations), the hardware fingerprint, the effort/grid/reps
used, timing, and the config file path – printable via
print().
See Also
get_local_EDI_optimization to see what is saved,
clear_local_EDI_optimization to return to shipped
defaults; the underlying tables:
get_cold_start_dispatch_policy,
get_warm_start_dispatch_policy,
get_optimization_dispatch_policy,
get_parallel_dispatch_policy.
Examples
# See what would change without writing anything. force = TRUE skips the
# idle-machine contention guard (see Details/`force` above) -- an example
# must not fail just because the machine running R CMD check happens to
# be busy (e.g. a shared CI runner); the guard itself has its own
# dedicated tests (test-local-machine-tuning-assembly.R). Scoped to one
# axis/family/n rather than the full live registry effort = "quick"
# otherwise walks (narrowing "quick" itself to a handful of
# high-effect-size families across every axis is still open, see
# edi_tuning_effort_presets()'s docs) -- this keeps the example a
# few-second sanity check instead of a multi-minute benchmark run.
res = tune_EDI_for_this_machine(effort = "quick", dry_run = TRUE, force = TRUE,
axes = "cold_start", families = "InferenceIncidLogRegr",
n_grid = 50L, reps = 1L)
print(res)
Unset the number of cores and stop parallel clusters
Description
This function stops any global fork or mirai clusters stored in the package environment and resets the core count to serial execution.
Usage
unset_num_cores()
Value
Invisible NULL.
Examples
set_num_cores(2)
unset_num_cores()
Fast Variance Calculation
Description
Calculates the variance of a numeric vector using Rcpp for speed.
Usage
var_cpp(x)
Arguments
x |
A numeric vector. |
Value
The variance of the vector.
Wilson Score Interval for a Single Proportion
Description
Wilson Score Interval for a Single Proportion
Usage
wilson_score_interval_cpp(x, n, alpha)
Zhang Exact Inference Helpers
Description
Internal method. Standalone functions to support Zhang (2026) exact test-inversion inference. These functions handle the bisection solver, p-value combination rules, and component-wise p-value logic.
Usage
zhang_combine_exact_pvals(p_M, p_R, m, nRT, nRC, method)
Arguments
p_M |
Matched p-value. |
p_R |
Reservoir p-value. |
m |
Number of matches. |
nRT |
Number of treated in reservoir. |
nRC |
Number of control in reservoir. |
method |
Combination method (Fisher or Stouffer). |