Geospatial data presents unique challenges for machine learning that are often overlooked in traditional data science workflows. This vignette explains why spatial cross-validation is essential for reliable model evaluation when working with geographic data.
“Everything is related to everything else, but near things are more related than distant things.”
This fundamental principle means that spatial observations are rarely independent. Two observations that are close in space tend to have similar characteristics due to:
Standard random k-fold cross-validation assumes that observations are independent and identically distributed (i.i.d.). When this assumption is violated by spatial dependence:
# Example: What happens with random CV on spatial data
set.seed(123)
n <- 100
x <- runif(n, 0, 100)
y <- runif(n, 0, 100)
z <- 10 + 0.5*x + 0.3*y + rnorm(n, 0, 2) # Spatially structured variable
# Random CV might put nearby points in both train and test
train_idx <- sample(1:n, 80)
test_idx <- setdiff(1:n, train_idx)
# Calculate minimum distance between train and test
distances <- numeric(length(test_idx))
for (i in seq_along(test_idx)) {
distances[i] <- min(sqrt((x[test_idx[i]] - x[train_idx])^2 +
(y[test_idx[i]] - y[train_idx])^2))
}
min(distances) # Often very small!## [1] 1.506625
When training and test observations are spatially close:
Spatial cross-validation methods explicitly control the separation between training and test observations to ensure:
You should use spatial cross-validation when:
Spatial cross-validation is particularly important in:
The spatialcvR package by Mamadou SOW offers:
library(spatialcvR)
# Load sample data
data(sample_spatial_data)
# Create spatial folds
folds <- spatial_folds(
data = sample_spatial_data,
x = "longitude",
y = "latitude",
k = 5,
method = "block"
)
# Examine the folds
print(folds)## Spatial Cross-Validation Folds
## ==============================
## Method: spatial_block
## Number of folds: 5
## Observations: 200
## CRS: Not defined
## Has duplicate coordinates: FALSE
##
## Fold sizes:
## Fold 1: 155 train, 45 test
## Fold 2: 150 train, 50 test
## Fold 3: 150 train, 50 test
## Fold 4: 168 train, 32 test
## Fold 5: 177 train, 23 test
# Detect spatial leakage
leakage <- detect_spatial_leakage(
data = sample_spatial_data,
folds = folds,
x = "longitude",
y = "latitude"
)
print(leakage)## Spatial Leakage Detection
## =========================
## Method: spatial_block
## Overall Risk Level: LOW
## Distance Threshold: 53.04
##
## Summary Statistics:
## Min distance: 11.24
## Mean distance: 521.24
## Median distance: 530.91
## Proportion below threshold: 0.1%
##
## Fold Analysis:
## Fold 1: LOW risk (0.1% below threshold)
## Fold 2: LOW risk (0.0% below threshold)
## Fold 3: LOW risk (0.1% below threshold)
## Fold 4: LOW risk (0.2% below threshold)
## Fold 5: LOW risk (0.1% below threshold)
##
## Recommendations:
## - Spatial separation appears adequate.
## - Current cross-validation setup should provide reliable performance estimates.
## - Consider increasing spatial separation if you need more conservative estimates.