Package {datascan}


Title: Scan Data for Quick Structural Summaries and Checks
Version: 0.1.1
Description: Scans data for checking columns that are constant, one-to-one, missing or all unique. Users can also try to identify the columns that uniquely index the observational unit and whether some columns are nested or complete.
License: MIT + file LICENSE
Encoding: UTF-8
Suggests: testthat (≥ 3.0.0)
Config/testthat/edition: 3
Config/roxygen2/version: 8.1.0
Imports: cli, dplyr, glue, magrittr, pillar, rlang, tibble, tidyr, tidyselect
Depends: R (≥ 4.1.0)
LazyData: true
URL: https://github.com/emitanaka/datascan, https://emitanaka.org/datascan/
BugReports: https://github.com/emitanaka/datascan/issues
NeedsCompilation: no
Packaged: 2026-09-25 05:46:27 UTC; emitanaka
Author: Emi Tanaka ORCID iD [aut, cre, cph]
Maintainer: Emi Tanaka <dr.emi.tanaka@gmail.com>
Repository: CRAN
Date/Publication: 2026-10-06 07:40:28 UTC

datascan: Scan Data for Quick Structural Summaries and Checks

Description

Scans data for checking columns that are constant, one-to-one, missing or all unique. Users can also try to identify the columns that uniquely index the observational unit and whether some columns are nested or complete.

Author(s)

Maintainer: Emi Tanaka dr.emi.tanaka@gmail.com (ORCID) [copyright holder]

Authors:

See Also

Useful links:


Add a row ID column to a data frame

Description

This function adds a new column to a data frame containing a unique identifier for each row. The default column name is ".id", but this can be customized using the name parameter.

Usage

add_row_id(data, name = ".id")

Arguments

data

A data frame to which the row ID column will be added.

name

The name of the new row ID column. Default is ".id".

Value

The input data frame with an additional column containing row IDs.

Examples

df <- data.frame(x = c(1, 2, 3), y = c("a", "b", "c"))
add_row_id(df)

Coerce to numerical and identify non-numerical entries

Description

This function attempts to coerce a variable to numerical and identifies any non-numerical entries. It returns the coerced numerical vector and prints a message indicating any non-numerical entries found.

Usage

as_numerical(x)

Arguments

x

A vector to be coerced to numerical.

Details

Factors are coerced using their labels rather than their underlying integer codes.

Value

A numerical vector with non-numerical entries coerced to NA.

See Also

Other numerical: is_numerical()

Examples

as_numerical(c("1", "2", "A", "4", "B"))
as_numerical(factor(c("10", "20", "A")))

Identify columns with all unique values

Description

Find columns that have all unique values. If there are no identified columns, it will return NA.

Usage

cols_all_unique(data)

Arguments

data

The data frame

Value

A character vector of column names with all unique values.

See Also

Other quality checks: cols_bijective(), cols_constant(), cols_identify_all(), cols_missing(), cols_nested()

Examples

cols_all_unique(ChickWeight)


Identify bijective (one-to-one correspondence) columns

Description

Any values that have all unique values will not be included in the output. If there are no identified columns, it will return an emtpy list.

Usage

cols_bijective(data)

Arguments

data

The data frame

Value

A list of character vectors of bijective columns.

See Also

Other quality checks: cols_all_unique(), cols_constant(), cols_identify_all(), cols_missing(), cols_nested()

Examples

cols_bijective(OrchardSprays)


Identify columns with a specific number of unique values

Description

This function identifies columns in a data frame that have a specific number of unique values. It can be used to find constant, binary, ternary, or any other n-ary columns. If there are no identified columns, it will return NA.

Usage

cols_binary(data, na.rm = FALSE)

cols_ternary(data, na.rm = FALSE)

cols_multinary(data, na.rm = FALSE, n = 1)

Arguments

data

The data frame

na.rm

Logical, whether to remove NA values before counting unique values.

n

The specific number of unique values to look for.

Value

A character vector of column names that have exactly n unique values.

Examples

cols_binary(mtcars)

Identify constant columns

Description

Similar to janitor::remove_constant, except this function aims to identify the constant columns without removing them. If there are no constant columns, it will return NA.

Usage

cols_constant(data, na.rm = FALSE)

cols_unary(data, na.rm = FALSE)

Arguments

data

The data frame

na.rm

Should missing values be removed?

Value

A character vector of constant columns.

See Also

Other quality checks: cols_all_unique(), cols_bijective(), cols_identify_all(), cols_missing(), cols_nested()

Examples

df <- data.frame(x1 = 1, x2 = c("A", "B", "C"), x3 = "a")
cols_constant(df)

Identify all columns with potential issues

Description

This function identifies all columns with potential issues, including all unique values, bijective columns, constant columns, and columns with missing values above a certain cutoff. It returns a list of identified columns for each category and can print the results to the console.

Usage

cols_identify_all(data, cutoff = 1, na.rm = FALSE, print = TRUE)

Arguments

data

The data frame

cutoff

The minimum cutoff for the proportion of missing values.

na.rm

Should missing values be removed?

print

Print the output or not.

Value

A named list with one element per check.

See Also

Other quality checks: cols_all_unique(), cols_bijective(), cols_constant(), cols_missing(), cols_nested()

Examples

cols_identify_all(airquality)


Identify columns that have certain proportion of missing values

Description

Similar to janitor::remove_empty, except this function aims to identify the columns without removing them. If there are no identified columns, it will return NA.

Usage

cols_missing(data, cutoff = 1)

Arguments

data

The data frame

cutoff

The minimum cutoff for the proportion of missing values.

Value

A character vector of column names with missing values with at least a certain proportion.

See Also

Other quality checks: cols_all_unique(), cols_bijective(), cols_constant(), cols_identify_all(), cols_nested()

Examples

# find columns that have any missing values
cols_missing(airquality, 0)

Nestedness checks

Description

This function checks for nestedness in a data frame. It identifies columns that are nested within other columns, meaning that the values in one column are subsets of the values in another column. The function returns a list of identified nested columns and can print the results to the console.

Usage

cols_nested(data, na.rm = FALSE, ignore = c("constant", "unique", "bijective"))

Arguments

data

The data frame to be checked for nestedness.

na.rm

Remove NA values when checking for nestedness.

ignore

A character vector of trivial cases of nesting to ignore, i.e. not count as nested. Any of "constant" (x or y has a single value), "unique" (x has all unique values) and "bijective" (x and y have a one-to-one correspondence). Use NULL to count all cases.

Details

By default, trivial cases of nesting (constant columns, columns with all unique values and pairs of bijective columns) are not reported as nested. Use ignore to change this (see is_nested()).

Value

A list of identified nested columns. The first element of each list is the child variable, and the second element is the parent variable.

See Also

Other quality checks: cols_all_unique(), cols_bijective(), cols_constant(), cols_identify_all(), cols_missing()

Examples

cols_nested(CO2)

Get the concurrence matrix or table

Description

Find the number of common levels of first variable across the levels of the second variable. The diagonal shows the total number of levels at the corresponding level. The off-diagonal shows the number of common levels between the two levels. The matrix is symmetric.

Usage

concurrence_matrix(data, x, group, na.rm = FALSE)

concurrence_table(data, x, group, na.rm = FALSE)

Arguments

data

A data frame containing the MET data

x

Categorical variable to be used for the concurrence matrix.

group

Categorical grouping variable.

na.rm

Remove NA values when checking for concurrence.

Value

concurrence_matrix() returns a symmetric numeric matrix of class concurrence_mat with one row and one column per level of group. The diagonal entries give the number of distinct levels of x observed at each level of group, and the off-diagonal entries give the number of levels of x shared by the two corresponding levels of group. The names of x and group are stored in the .vars attribute.

concurrence_table() returns the same information in long format as a tibble of class concurrence_tbl, with one row per pair of group levels and the columns:

⁠<group>_1⁠, ⁠<group>_2⁠

Factors giving the pair of group levels.

concurrence

Number of levels of x shared by the pair.

prop_in_1

Proportion of the levels of x in ⁠<group>_1⁠ that are also in ⁠<group>_2⁠.

prop_in_2

Proportion of the levels of x in ⁠<group>_2⁠ that are also in ⁠<group>_1⁠.

Examples

df <- expand.grid(gen = paste0("G", 1:20), env = paste0("E", 1:15))
df <- df[sample(nrow(df), 100), ]
(mat <- concurrence_matrix(df, gen, env))
extract(mat, "E1", "E2")
extract(mat, "E1", c("E2", "E3"))
extract(mat, c("E5", "E6", "E7"), c("E2", "E3"))
# proportion of levels in the first variable
# at each level of the second variable
proportions(mat, 1)


Extract operator

Description

See magrittr::extract for details. This works well in conjunction with the concurrence_matrix function to extract specific rows and columns from the resulting matrix.

Usage

x |> extract(i, j)

Value

The subset of x selected by i and j, equivalent to x[i, j].


Is variable complete within another variable?

Description

This function checks if a variable is complete within another variable. If all levels of the first variable appear in all levels of the second variable, it returns TRUE; otherwise, it returns FALSE. The function can also handle NA values based on the na.rm parameter.

Usage

is_complete(x, y, na.rm = TRUE)

Arguments

x, y

Vectors of the same size where x (child variable) is checked for completeness within y (parent variable)

na.rm

Remove NA values when checking for completeness.

Value

TRUE if the values in vector x are complete

Examples

df <- data.frame(
  site = c("A", "A", "B", "B"),
  plot1 = c("A1", "A2", "B1", "B2"),
  plot2 = c("A1", "A2", "A1", "A2")
)
is_complete(df$plot1, df$site) # FALSE
is_complete(df$plot2, df$site) # TRUE

Is variable nested within another variable?

Description

This function checks if a variable is nested within another variable. If the levels in the first variable appear only within the levels of the second variable, it returns TRUE; otherwise, it returns FALSE. The function can also handle NA values based on the na.rm parameter.

Usage

is_nested(x, y, na.rm = TRUE, ignore = c("constant", "unique", "bijective"))

Arguments

x, y

Vectors of the same size where x (child variable) is nested in y (parent variable)

na.rm

Remove NA values when checking for nestedness.

ignore

A character vector of trivial cases of nesting to ignore, i.e. not count as nested. Any of "constant" (x or y has a single value), "unique" (x has all unique values) and "bijective" (x and y have a one-to-one correspondence). Use NULL to count all cases.

Details

By default, trivial cases of nesting are not counted and it returns FALSE when x or y is constant (has a single value), x has all unique values, or x and y are bijective (have a one-to-one correspondence). Use ignore to choose which of these cases are not counted.

Value

TRUE if the values in vector x are nested within vector y, FALSE otherwise. It is also FALSE for any trivial case listed in ignore.

Examples

df <- data.frame(
  site = rep(c("A", "B"), each = 4),
  plot1 = rep(c("A1", "A2", "B1", "B2"), each = 2),
  plot2 = rep(c("A1", "A2", "A1", "A2"), each = 2)
)
is_nested(df$plot1, df$site) # TRUE
is_nested(df$plot2, df$site) # FALSE since A1 and A2 appear in both sites
df$id <- seq_len(nrow(df))
is_nested(df$id, df$site) # FALSE since id has all unique values
is_nested(df$id, df$site, ignore = c("constant", "bijective")) # TRUE

Is the variable numerical?

Description

This function checks if a variable is numerical. It returns TRUE if the variable is numerical and FALSE otherwise. If non-numerical entries are found, it prints a message indicating the non-numerical entries.

Usage

is_numerical(x)

Arguments

x

A vector to be checked for numericality.

Value

TRUE if the variable is numerical, FALSE otherwise.

See Also

Other numerical: as_numerical()

Examples

is_numerical(c("1", "2", "3"))
is_numerical(c("1", "2", "3", "A"))

Is it the observational unit?

Description

An observational unit is the entity that is being measured in a study. For example, if blood pressure of a person is measured, then the person is the observational unit. If blood pressure is measured multiple times for a person, then the observational unit is the person and the time of measurement.

Usage

is_observational_unit(.data, ..., .check_redundancy = TRUE)

is_obs_unit(.data, ..., .check_redundancy = TRUE)

Arguments

.data

The data frame

...

The columns in the data frame that potentially index the observational unit.

.check_redundancy

Logical. If TRUE, checks if any of the columns are redundant.

Details

Supply the columns that potentially index the observational unit. The column cannot uniquely identify the row. If the number of rows in the data frame is equal to the number of distinct rows in the selected columns, then it returns TRUE, otherwise FALSE. If TRUE, it checks if any of the columns are redundant.

Value

A single logical value. TRUE if the selected columns uniquely identify each row of .data (and, when .check_redundancy = TRUE, no selected column is redundant), and FALSE otherwise. In an interactive session, a message with the result is also printed.

Examples

is_observational_unit(ChickWeight, Time, Chick)
is_observational_unit(ChickWeight, Time, Chick, Diet)


Identify rows that have certain proportion of missing values

Description

Similar to janitor::remove_empty, except this function aims to identify the columns without removing them. If there are no identified columns, it will return NA.

Usage

rows_missing(data, cutoff = 1)

Arguments

data

The data frame

cutoff

The minimum cutoff for the proportion of missing values.

Value

An integer vector of rows with missing values with at least a certain proportion.

Examples

rows_missing(airquality, 0.1)


Coordination of care by breeders and helpers in the cooperatively breeding long-tailed tit

Description

Usage

watch_tits

Format

An object of class tbl_df (inherits from tbl, data.frame) with 7950 rows and 21 columns.

Source

Chay et al. (2022) "Coordination of care by breeders and helpers in the cooperatively breeding long-tailed tit, Aegithalos caudatus". Data available at Dryad: https://datadryad.org/dataset/doi:10.5061/dryad.mkkwh712k


A field experiment assessing the roles of drought, herbivory, and local climate for white clovers

Description

This dataset contains data from a field experiment conducted by Albano et al. (2026) to assess the roles of drought, herbivory, and local climate on white clovers (Trifolium repens). The experiment was designed to investigate how these factors influence the growth, survival, and reproductive success of white clovers in different environmental conditions.

Usage

white_clover

Format

An object of class tbl_df (inherits from tbl, data.frame) with 768 rows and 27 columns.

Details

Missing values (NA) in the 2021 and 2022 measurements represent plants for which data could not be collected, or, for 2022, plants that had already died.

Source

Albano et al. (2026) "Data from: A field experiment assessing the roles of drought, herbivory, and local climate on cyanogenesis cline formation and local adaptation in Trifolium repens". Data available at Dryad: https://datadryad.org/dataset/doi%3A10.5061/dryad.4xgxd25n1