Engineering Robust Geospatial Prediction Under Spatial Data Constraints
Date
2026
Authors
Advisors
Journal Title
Journal ISSN
Volume Title
Repository Usage Stats
views
downloads
Attention Stats
Abstract
Many problems in climate monitoring, economic development, and public health require understanding how key quantities vary across geographic regions where direct measurements are sparse or unavailable. Because collecting high-quality observations is often expensive or logistically difficult, measurements tend to concentrate in regions where data collection is most feasible, leaving other regions only partially observed. Estimating global building-sector greenhouse gas emissions exemplifies this challenge. In developing methods to estimate global building emissions, we encountered methodological challenges that arise throughout the geospatial prediction pipeline. Predictive error can emerge at multiple stages of this pipeline, including spatial dataset construction, geospatial data processing, and model generalization across geographic regions. This dissertation investigates these sources of error through empirical analysis of these stages.
First, we develop a method to generate a high-spatial-resolution global building emissions dataset at approximately 1km-by-1km spatial resolution by disaggregating EDGAR v8.0 emissions estimates using satellite-derived building floor area and energy consumption activity data. The resulting dataset improves sub-national emissions estimates while maintaining consistency with national inventories, and is currently published monthly through the ClimateTRACE initiative, a global coalition producing openly accessible, high-resolution greenhouse gas emissions estimates used by governments and organizations worldwide to support emissions monitoring and mitigation planning.
Second, we examine how methodological choices in spatial data processing affect predictive accuracy. These choices are critical because many geospatial analyses aggregate gridded data within administrative boundaries where small methodological differences can substantially affect the resulting estimates. For example, building emissions estimates are often stored in gridded raster form but aggregated across administrative polygon boundaries defining cities or municipalities. Using controlled experiments with synthetic spatial point process data, we evaluate several raster-to-polygon aggregation strategies and quantify how estimation error varies as a function of grid resolution and polygon size. These experiments demonstrate that commonly used centroid-based aggregation can introduce substantial error under many spatial configurations and motivate the introduction of a Confidence Factor that quantifies the expected error associated with different aggregation choices.
Finally, we investigate how ML models behave when they are trained on data from one geographic region and applied to data from another. Because residential energy use intensity (EUI) cannot be observed at global scale, we follow IPCC conventions and decompose the computation of emissions into the product of three distinct parameters -- one of which is EUI. Because EUI is the least observable component of building emissions estimation at global scale, it presents a natural opportunity to explore whether ML models can infer EUI and thereby reduce emissions estimation error. Across a range of ML model architectures, we first evaluate whether EUI can be estimated from available regional characteristics and whether explicitly modeling spatial relationships can improve predictive performance. To do this, we consider tree-based models, neural networks, transformers, and graph neural networks (GNNs), including both a standard GNN and our proposed Graph Neural Network with Multiple Aggregations and Residuals (GNNMAR). While some regional improvements are observed, predictive performance remains unstable under geographically partitioned evaluation. This motivates a broader analysis across four geospatial tabular regression datasets, where predictive variability emerges across ML model architectures when moving from random to spatial cross-validation. We subsequently introduce an operational taxonomy to better understand these failures, and explore whether context engineering for the prior-fitted network TabPFN can influence predictive performance.
In conclusion, this dissertation demonstrates that predictive error in geospatial regression often arises from interactions across spatial data construction, geospatial data processing, and model generalization under spatial data constraints, rather than from model architecture alone. By developing improved spatial emissions datasets, quantifying the uncertainty introduced by spatial aggregation methods, and diagnosing how ML models fail under geographic partitioning, this work provides a framework for understanding and mitigating error throughout the geospatial prediction pipeline.
Type
Department
Description
Provenance
Subjects
Citation
Permalink
Citation
Markakis, Paul James (2026). Engineering Robust Geospatial Prediction Under Spatial Data Constraints. Dissertation, Duke University. Retrieved from https://hdl.handle.net/10161/35304.
Collections
Except where otherwise noted, student scholarship that was shared on DukeSpace after 2009 is made available to the public under a Creative Commons Attribution / Non-commercial / No derivatives (CC-BY-NC-ND) license. All rights in student work shared on DukeSpace before 2009 remain with the author and/or their designee, whose permission may be required for reuse.
