A map can leak without sharing a single pixel
Old-growth forest is rare, spatially clustered, and expensive to label. Ratsakatika and colleagues map it across 211,893 hectares of Romania's Southern Carpathians. They compare ordinary Sentinel-1/2 satellite features with AlphaEarth and TESSERA geospatial foundation-model embeddings, using XGBoost to score individual 10 metre pixels before averaging predictions within forest parcels.
Their six validation folds are already spatial blocks. Adjacent parcels are kept together, tuning is nested inside the training folds, and the final metric is pooled out-of-fold precision-recall area under the curve (PR-AUC). Yet parcels on opposite sides of a fold boundary can still be neighbours. Smooth terrain, road access, labels, and embeddings can therefore make the test landscape look familiar.
The buffer is a geometric filter
For each held-out fold, remove every training parcel whose distance to any test parcel is less than the chosen radius. The comparison is strict about purpose: the same model is refitted after filtering, while a parcel-matched control removes the same number of training parcels at random. That separates proximity from the simpler penalty of having less data.
def buffered_indices(train, test, distance):
if distance < 0:
raise ValueError('distance must be non-negative')
kept = []
for i, (x, y) in enumerate(train):
outside = all(
((x - tx) ** 2 + (y - ty) ** 2) ** 0.5 >= distance
for tx, ty in test
)
if outside:
kept.append(i)
return keptAt radius 6, the toy keeps 5 of 8 training points. The production site runs the TypeScript implementation; Python and C++ above are faithful translations. No randomness is needed.
The absolute signal survives the harder exam
| XGBoost feature set | PR-AUC, 0 km | PR-AUC, 10 km | Change |
|---|---|---|---|
| Baseline | 0.548 | 0.475 | −0.073 |
| Baseline + Sentinel EO | 0.765 | 0.684 | −0.081 |
| Baseline + AlphaEarth | 0.765 | 0.701 | −0.064 |
| Baseline + TESSERA | 0.840 | 0.727 | −0.113 |
The harder test does not erase the useful result. Relative to the topography-and-access baseline, every satellite-derived feature set still adds 0.21–0.25 PR-AUC at 10 km, with all reported 95% intervals above zero. The coordinate-only control falls to the 0.21 prevalence floor, while randomly removing half the training parcels costs less than 0.04. Proximity—not merely sample count—explains most of the drop.
The premium feature stack loses its clear edge
Without a buffer, TESSERA reaches 0.84 PR-AUC and beats both conventional Sentinel EO and AlphaEarth by 0.075. With 10 km separation, those point gaps shrink to 0.043 and 0.026. More importantly, their confidence intervals become −0.01 to 0.11 and −0.04 to 0.10: both include zero.
PR-AUC asks the right rare-class question
Average precision sorts parcels by score and adds precision at each gain in recall. In the hand-verified ranking positive, negative, positive, negative, it equals 0.833. Unlike overall accuracy, its no-skill reference is prevalence: 0.21 in the labelled study sample. Across Europe, where primary and old-growth forest is under 3%, a useless always-negative classifier can report about 97% accuracy.
What to probe next
The labels are a non-probability sample from one beech–spruce mountain landscape. The claimed area of applicability covers feature-space support, not verified accuracy outside that region; only 65% of wider Carpathian forest falls inside it. Next comes genuinely external field validation, more confirmed non-old-growth parcels, temporal transfer, and a pre-registered buffer chosen from spatial dependence rather than the result curve.
The primary pixel model is explained on the XGBoost model page. For a different failure caused by how evaluation examples are paired, see the forecast-protocol walkthrough.
References
- T. Ratsakatika, M. Zotta, S. Keshav, and E. R. Lines (2026). Geospatial embeddings detect old-growth forests but buffered spatial validation narrows their advantage over Sentinel features. arXiv:2609.28194 [cs.LG]
- D. R. Roberts et al. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography 40(8), 913–929
- T. Saito and M. Rehmsmeier (2015). The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE 10(3), e0118432