UHI-Bench: Benchmarking Dual-Source Urban Heat Island Modeling Across Cities in Diverse Climate Regimes
UHI-Bench is introduced as the first benchmark for dual-source Urban Heat Island (UHI) modeling, integrating both land surface temperature (LST-UHI) and near-surface air temperature (AirT-UHI) observations The benchmark evaluates over 20 baselines from four model families across five tasks, 20 cities, and nine Köppen climate classes using a unified signal, mechanism, and transfer framework Foundation models demonstrate consistently competitive and stable performance, though no single model is un
Analysis
TL;DR
- UHI-Bench is introduced as the first benchmark for dual-source Urban Heat Island (UHI) modeling, integrating both land surface temperature (LST-UHI) and near-surface air temperature (AirT-UHI) observations
- The benchmark evaluates over 20 baselines from four model families across five tasks, 20 cities, and nine Köppen climate classes using a unified signal, mechanism, and transfer framework
- Foundation models demonstrate consistently competitive and stable performance, though no single model is uniformly best across all tasks and cities
- Environmental covariates generally improve model performance, but their utility varies significantly across data sources and prediction tasks
- Cross-city transferability is better explained by overlap in UHI regimes rather than by climate-zone similarity, challenging conventional assumptions about climate-based generalization
Why It Matters
This benchmark addresses a critical gap in climate and urban heat research by providing a standardized evaluation framework for dual-source UHI modeling, which is essential for accurate human thermal exposure assessment. For AI practitioners working in climate science and geospatial ML, UHI-Bench offers a rigorous testbed for evaluating model generalization across diverse climates and urban morphologies. The findings have direct implications for deploying ML models in real-world urban heat mitigation and public health applications.
Technical Details
- Dual-source modeling framework: The benchmark jointly addresses LST-UHI (satellite-derived land surface temperature) and AirT-UHI (ground station near-surface air temperature), recognizing that these capture physically distinct aspects of urban heat and substituting one for the other can substantially bias heat exposure estimates
- Unified evaluation structure: Follows a three-part framework—signal (data representation), mechanism (model architecture and environmental covariate integration), and transfer (cross-city generalization)—enabling systematic comparison across model families
- Scale and diversity: Evaluates 20+ baselines across 20 cities spanning nine Köppen climate classes, addressing spatiotemporal incompatibilities between dynamic meteorological drivers and static urban morphology features, as well as cloud gaps in LST and sparse AirT station networks
- Key finding on transferability: Cross-city transfer performance correlates more strongly with UHI regime overlap than with climate-zone similarity, suggesting that mechanistic similarity in heat dynamics matters more than broad climatic classification for model portability
Industry Insight
- Climate AI practitioners should prioritize UHI regime similarity over climate-zone matching when selecting source cities for transfer learning, as this yields more reliable cross-city generalization than traditional climate-based stratification
- The consistent competitiveness of foundation models suggests they are strong candidates for deployment in urban heat modeling, but task-specific fine-tuning with appropriate environmental covariates remains essential for optimal performance
- The benchmark highlights the importance of data equity in climate research—cities with sparse monitoring infrastructure can benefit from cross-city transfer, but only if benchmarking frameworks explicitly account for diverse climate regimes and UHI mechanisms
Disclaimer: The above content is generated by AI and is for reference only.