← Blog/blog/spatial-validation-buffer-leakage

A spatial fold can still have next-door neighbours

01

A map can leak without sharing a single pixel

Old-growth forest is rare, spatially clustered, and expensive to label. Ratsakatika and colleagues map it across 211,893 hectares of Romania's Southern Carpathians. They compare ordinary Sentinel-1/2 satellite features with AlphaEarth and TESSERA geospatial foundation-model embeddings, using XGBoost to score individual 10 metre pixels before averaging predictions within forest parcels.

Their six validation folds are already spatial blocks. Adjacent parcels are kept together, tuning is nested inside the training folds, and the final metric is pooled out-of-fold precision-recall area under the curve (PR-AUC). Yet parcels on opposite sides of a fold boundary can still be neighbours. Smooth terrain, road access, labels, and embeddings can therefore make the test landscape look familiar.

02

The buffer is a geometric filter

For each held-out fold, remove every training parcel whose distance to any test parcel is less than the chosen radius. The comparison is strict about purpose: the same model is refitted after filtering, while a parcel-matched control removes the same number of training parcels at random. That separates proximity from the simpler penalty of having less data.

def buffered_indices(train, test, distance):
    if distance < 0:
        raise ValueError('distance must be non-negative')
    kept = []
    for i, (x, y) in enumerate(train):
        outside = all(
            ((x - tx) ** 2 + (y - ty) ** 2) ** 0.5 >= distance
            for tx, ty in test
        )
        if outside:
            kept.append(i)
    return kept
Training points retained
Fig 1. Seed-free toy geometry computed by the same core buffer function. Retention falls only when a radius crosses an actual train–test distance; this is not a smooth regularizer.

At radius 6, the toy keeps 5 of 8 training points. The production site runs the TypeScript implementation; Python and C++ above are faithful translations. No randomness is needed.

03

The absolute signal survives the harder exam

XGBoost feature setPR-AUC, 0 kmPR-AUC, 10 kmChange
Baseline0.5480.475−0.073
Baseline + Sentinel EO0.7650.684−0.081
Baseline + AlphaEarth0.7650.701−0.064
Baseline + TESSERA0.8400.727−0.113
Paper-reported pooled out-of-fold parcel PR-AUC. The 10 km arm removes nearby training parcels before refitting.
0 km buffer10 km buffer
Fig 2. Paper-reported PR-AUC for the four XGBoost feature stacks, before and after the 10 km buffer. The x-axis order is baseline, Sentinel EO, AlphaEarth, TESSERA.

The harder test does not erase the useful result. Relative to the topography-and-access baseline, every satellite-derived feature set still adds 0.21–0.25 PR-AUC at 10 km, with all reported 95% intervals above zero. The coordinate-only control falls to the 0.21 prevalence floor, while randomly removing half the training parcels costs less than 0.04. Proximity—not merely sample count—explains most of the drop.

04

The premium feature stack loses its clear edge

Without a buffer, TESSERA reaches 0.84 PR-AUC and beats both conventional Sentinel EO and AlphaEarth by 0.075. With 10 km separation, those point gaps shrink to 0.043 and 0.026. More importantly, their confidence intervals become −0.01 to 0.11 and −0.04 to 0.10: both include zero.

TESSERA − Sentinel EOTESSERA − AlphaEarth
Fig 3. Core-computed paired point differences from paper-reported PR-AUC. Buffering narrows TESSERA's apparent advantage; intervals, not shown as fabricated bars, determine the uncertainty claim.
05

PR-AUC asks the right rare-class question

Average precision sorts parcels by score and adds precision at each gain in recall. In the hand-verified ranking positive, negative, positive, negative, it equals 0.833. Unlike overall accuracy, its no-skill reference is prevalence: 0.21 in the labelled study sample. Across Europe, where primary and old-growth forest is under 3%, a useless always-negative classifier can report about 97% accuracy.

Fig 4. Paper-reported TESSERA PR-AUC before and after buffering, beside the study sample's no-skill prevalence. This is a performance comparison, not a claim that prevalence is identical outside the labelled sample.
06

What to probe next

The labels are a non-probability sample from one beech–spruce mountain landscape. The claimed area of applicability covers feature-space support, not verified accuracy outside that region; only 65% of wider Carpathian forest falls inside it. Next comes genuinely external field validation, more confirmed non-old-growth parcels, temporal transfer, and a pre-registered buffer chosen from spatial dependence rather than the result curve.

The primary pixel model is explained on the XGBoost model page. For a different failure caused by how evaluation examples are paired, see the forecast-protocol walkthrough.