-
Granger causal inference for climate change attribution
Authors:
Mark D. Risser,
Mohammed Ombadi,
Michael F. Wehner
Abstract:
Climate change detection and attribution (D&A) is concerned with determining the extent to which anthropogenic activities have influenced specific aspects of the global climate system. D&A fits within the broader field of causal inference, the collection of statistical methods that identify cause and effect relationships. There are a wide variety of methods for making attribution statements, each…
▽ More
Climate change detection and attribution (D&A) is concerned with determining the extent to which anthropogenic activities have influenced specific aspects of the global climate system. D&A fits within the broader field of causal inference, the collection of statistical methods that identify cause and effect relationships. There are a wide variety of methods for making attribution statements, each of which require different types of input data and focus on different types of weather and climate events and each of which are conditional to varying extents. Some methods are based on Pearl causality while others leverage Granger causality, and the causal framing provides important context for how the resulting attribution conclusion should be interpreted. However, while Granger-causal attribution analyses have become more common, there is no clear statement of their strengths and weaknesses and no clear consensus on where and when Granger-causal perspectives are appropriate. In this prospective paper, we provide a formal definition for Granger-based approaches to trend and event attribution and a clear comparison with more traditional methods for assessing the human influence on extreme weather and climate events. Broadly speaking, Granger-causal attribution statements can be constructed quickly from observations and do not require computationally-intesive dynamical experiments. These analyses also enable rapid attribution, which is useful in the aftermath of a severe weather event, and provide multiple lines of evidence for anthropogenic climate change when paired with Pearl-causal attribution. Confidence in attribution statements is increased when different methodologies arrive at similar conclusions. Moving forward, we encourage the D&A community to embrace hybrid approaches to climate change attribution that leverage the strengths of both Granger and Pearl causality.
△ Less
Submitted 13 August, 2024;
originally announced August 2024.
-
Impossible temperatures are not as rare as you think
Authors:
Mark D. Risser,
Likun Zhang,
Michael F. Wehner
Abstract:
The last decade has seen numerous record-shattering heatwaves in all corners of the globe. In the aftermath of these devastating events, there is interest in identifying worst-case thresholds or upper bounds that quantify just how hot temperatures can become. Generalized Extreme Value theory provides a data-driven estimate of extreme thresholds; however, upper bounds may be exceeded by future even…
▽ More
The last decade has seen numerous record-shattering heatwaves in all corners of the globe. In the aftermath of these devastating events, there is interest in identifying worst-case thresholds or upper bounds that quantify just how hot temperatures can become. Generalized Extreme Value theory provides a data-driven estimate of extreme thresholds; however, upper bounds may be exceeded by future events, which undermines attribution and planning for heatwave impacts. Here, we show how the occurrence and relative probability of observed events that exceed a priori upper bound estimates, so-called "impossible" temperatures, has changed over time. We find that many unprecedented events are actually within data-driven upper bounds, but only when using modern spatial statistical methods. Furthermore, there are clear connections between anthropogenic forcing and the "impossibility" of the most extreme temperatures. Robust understanding of heatwave thresholds provides critical information about future record-breaking events and how their extremity relates to historical measurements.
△ Less
Submitted 9 August, 2024;
originally announced August 2024.
-
Huge Ensembles Part I: Design of Ensemble Weather Forecasts using Spherical Fourier Neural Operators
Authors:
Ankur Mahesh,
William Collins,
Boris Bonev,
Noah Brenowitz,
Yair Cohen,
Joshua Elms,
Peter Harrington,
Karthik Kashinath,
Thorsten Kurth,
Joshua North,
Travis OBrien,
Michael Pritchard,
David Pruitt,
Mark Risser,
Shashank Subramanian,
Jared Willard
Abstract:
Studying low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computatio…
▽ More
Studying low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1,000-10,000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part I, we construct an ensemble weather forecasting system based on Spherical Fourier Neural Operators (SFNO), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. Using large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states, and the ML ensemble thus passes a crucial spectral test in the literature. The IFS and ML ensembles have similar Extreme Forecast Indices, and we show that the ML extreme weather forecasts are reliable and discriminating.
△ Less
Submitted 6 August, 2024;
originally announced August 2024.
-
Huge Ensembles Part II: Properties of a Huge Ensemble of Hindcasts Generated with Spherical Fourier Neural Operators
Authors:
Ankur Mahesh,
William Collins,
Boris Bonev,
Noah Brenowitz,
Yair Cohen,
Peter Harrington,
Karthik Kashinath,
Thorsten Kurth,
Joshua North,
Travis OBrien,
Michael Pritchard,
David Pruitt,
Mark Risser,
Shashank Subramanian,
Jared Willard
Abstract:
In Part I, we created an ensemble based on Spherical Fourier Neural Operators. As initial condition perturbations, we used bred vectors, and as model perturbations, we used multiple checkpoints trained independently from scratch. Based on diagnostics that assess the ensemble's physical fidelity, our ensemble has comparable performance to operational weather forecasting systems. However, it require…
▽ More
In Part I, we created an ensemble based on Spherical Fourier Neural Operators. As initial condition perturbations, we used bred vectors, and as model perturbations, we used multiple checkpoints trained independently from scratch. Based on diagnostics that assess the ensemble's physical fidelity, our ensemble has comparable performance to operational weather forecasting systems. However, it requires several orders of magnitude fewer computational resources. Here in Part II, we generate a huge ensemble (HENS), with 7,424 members initialized each day of summer 2023. We enumerate the technical requirements for running huge ensembles at this scale. HENS precisely samples the tails of the forecast distribution and presents a detailed sampling of internal variability. For extreme climate statistics, HENS samples events 4$σ$ away from the ensemble mean. At each grid cell, HENS improves the skill of the most accurate ensemble member and enhances coverage of possible future trajectories. As a weather forecasting model, HENS issues extreme weather forecasts with better uncertainty quantification. It also reduces the probability of outlier events, in which the verification value lies outside the ensemble forecast distribution.
△ Less
Submitted 2 August, 2024;
originally announced August 2024.
-
A Unifying Perspective on Non-Stationary Kernels for Deeper Gaussian Processes
Authors:
Marcus M. Noack,
Hengrui Luo,
Mark D. Risser
Abstract:
The Gaussian process (GP) is a popular statistical technique for stochastic function approximation and uncertainty quantification from data. GPs have been adopted into the realm of machine learning in the last two decades because of their superior prediction abilities, especially in data-sparse scenarios, and their inherent ability to provide robust uncertainty estimates. Even so, their performanc…
▽ More
The Gaussian process (GP) is a popular statistical technique for stochastic function approximation and uncertainty quantification from data. GPs have been adopted into the realm of machine learning in the last two decades because of their superior prediction abilities, especially in data-sparse scenarios, and their inherent ability to provide robust uncertainty estimates. Even so, their performance highly depends on intricate customizations of the core methodology, which often leads to dissatisfaction among practitioners when standard setups and off-the-shelf software tools are being deployed. Arguably the most important building block of a GP is the kernel function which assumes the role of a covariance operator. Stationary kernels of the Matérn class are used in the vast majority of applied studies; poor prediction performance and unrealistic uncertainty quantification are often the consequences. Non-stationary kernels show improved performance but are rarely used due to their more complicated functional form and the associated effort and expertise needed to define and tune them optimally. In this perspective, we want to help ML practitioners make sense of some of the most common forms of non-stationarity for Gaussian processes. We show a variety of kernels in action using representative datasets, carefully study their properties, and compare their performances. Based on our findings, we propose a new kernel that combines some of the identified advantages of existing kernels.
△ Less
Submitted 18 September, 2023;
originally announced September 2023.
-
A flexible class of priors for orthonormal matrices with basis function-specific structure
Authors:
Joshua S. North,
Mark D. Risser,
F. Jay Breidt
Abstract:
Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. Howe…
▽ More
Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system's internal variability.
△ Less
Submitted 23 May, 2024; v1 submitted 25 July, 2023;
originally announced July 2023.
-
Explaining the unexplainable: leveraging extremal dependence to characterize the 2021 Pacific Northwest heatwave
Authors:
Likun Zhang,
Mark D. Risser,
Michael F. Wehner,
Travis A. O'Brien
Abstract:
In late June, 2021, a devastating heatwave affected the US Pacific Northwest and western Canada, breaking numerous all-time temperature records by large margins and directly causing hundreds of fatalities. The observed 2021 daily maximum temperature across much of the U.S. Pacific Northwest exceeded upper bound estimates obtained from single-station temperature records even after accounting for an…
▽ More
In late June, 2021, a devastating heatwave affected the US Pacific Northwest and western Canada, breaking numerous all-time temperature records by large margins and directly causing hundreds of fatalities. The observed 2021 daily maximum temperature across much of the U.S. Pacific Northwest exceeded upper bound estimates obtained from single-station temperature records even after accounting for anthropogenic climate change, meaning that the event could not have been predicted under standard univariate extreme value analysis assumptions. In this work, we utilize a flexible spatial extremes model that considers all stations across the Pacific Northwest domain and accounts for the fact that many stations simultaneously experience extreme temperatures. Our analysis incorporates the effects of anthropogenic forcing and natural climate variability in order to better characterize time-varying changes in the distribution of daily temperature extremes. We show that greenhouse gas forcing, drought conditions and large-scale atmospheric modes of variability all have significant impact on summertime maximum temperatures in this region. Our model represents a significant improvement over corresponding single-station analysis, and our posterior medians of the upper bounds are able to anticipate more than 96% of the observed 2021 high station temperatures after properly accounting for extremal dependence.
△ Less
Submitted 7 July, 2023;
originally announced July 2023.
-
Exact Gaussian Processes for Massive Datasets via Non-Stationary Sparsity-Discovering Kernels
Authors:
Marcus M. Noack,
Harinarayan Krishnan,
Mark D. Risser,
Kristofer G. Reyes
Abstract:
A Gaussian Process (GP) is a prominent mathematical framework for stochastic function approximation in science and engineering applications. This success is largely attributed to the GP's analytical tractability, robustness, non-parametric structure, and natural inclusion of uncertainty quantification. Unfortunately, the use of exact GPs is prohibitively expensive for large datasets due to their u…
▽ More
A Gaussian Process (GP) is a prominent mathematical framework for stochastic function approximation in science and engineering applications. This success is largely attributed to the GP's analytical tractability, robustness, non-parametric structure, and natural inclusion of uncertainty quantification. Unfortunately, the use of exact GPs is prohibitively expensive for large datasets due to their unfavorable numerical complexity of $O(N^3)$ in computation and $O(N^2)$ in storage. All existing methods addressing this issue utilize some form of approximation -- usually considering subsets of the full dataset or finding representative pseudo-points that render the covariance matrix well-structured and sparse. These approximate methods can lead to inaccuracies in function approximations and often limit the user's flexibility in designing expressive kernels. Instead of inducing sparsity via data-point geometry and structure, we propose to take advantage of naturally-occurring sparsity by allowing the kernel to discover -- instead of induce -- sparse structure. The premise of this paper is that GPs, in their most native form, are often naturally sparse, but commonly-used kernels do not allow us to exploit this sparsity. The core concept of exact, and at the same time sparse GPs relies on kernel definitions that provide enough flexibility to learn and encode not only non-zero but also zero covariances. This principle of ultra-flexible, compactly-supported, and non-stationary kernels, combined with HPC and constrained optimization, lets us scale exact GPs well beyond 5 million data points.
△ Less
Submitted 18 May, 2022;
originally announced May 2022.
-
Nonstationary Bayesian modeling for a large data set of derived surface temperature return values
Authors:
Mark Risser
Abstract:
Heat waves resulting from prolonged extreme temperatures pose a significant risk to human health globally. Given the limitations of observations of extreme temperature, climate models are often used to characterize extreme temperature globally, from which one can derive quantities like return values to summarize the magnitude of a low probability event for an arbitrary geographic location. However…
▽ More
Heat waves resulting from prolonged extreme temperatures pose a significant risk to human health globally. Given the limitations of observations of extreme temperature, climate models are often used to characterize extreme temperature globally, from which one can derive quantities like return values to summarize the magnitude of a low probability event for an arbitrary geographic location. However, while these derived quantities are useful on their own, it is also often important to apply a spatial statistical model to such data in order to, e.g., understand how the spatial dependence properties of the return values vary over space and emulate the climate model for generating additional spatial fields with corresponding statistical properties. For these objectives, when modeling global data it is critical to use a nonstationary covariance function. Furthermore, given that the output of modern global climate models can be on the order of $\mathcal{O}(10^4)$, it is important to utilize approximate Gaussian process methods to enable inference. In this paper, we demonstrate the application of methodology introduced in Risser and Turek (2020) to conduct a nonstationary and fully Bayesian analysis of a large data set of 20-year return values derived from an ensemble of global climate model runs with over 50,000 spatial locations. This analysis uses the freely available BayesNSGP software package for R.
△ Less
Submitted 27 April, 2020;
originally announced May 2020.
-
The effect of geographic sampling on evaluation of extreme precipitation in high resolution climate models
Authors:
Mark D. Risser,
Michael F. Wehner
Abstract:
Traditional approaches for comparing global climate models and observational data products typically fail to account for the geographic location of the underlying weather station data. For modern high-resolution models, this is an oversight since there are likely grid cells where the physical output of a climate model is compared with a statistically interpolated quantity instead of actual measure…
▽ More
Traditional approaches for comparing global climate models and observational data products typically fail to account for the geographic location of the underlying weather station data. For modern high-resolution models, this is an oversight since there are likely grid cells where the physical output of a climate model is compared with a statistically interpolated quantity instead of actual measurements of the climate system. In this paper, we quantify the impact of geographic sampling on the relative performance of high resolution climate models' representation of precipitation extremes in Boreal winter (DJF) over the contiguous United States (CONUS), comparing model output from five early submissions to the HighResMIP subproject of the CMIP6 experiment. We find that properly accounting for the geographic sampling of weather stations can significantly change the assessment of model performance. Across the models considered, failing to account for sampling impacts the different metrics (extreme bias, spatial pattern correlation, and spatial variability) in different ways (both increasing and decreasing). We argue that the geographic sampling of weather stations should be accounted for in order to yield a more straightforward and appropriate comparison between models and observational data sets, particularly for high resolution models. While we focus on the CONUS in this paper, our results have important implications for other global land regions where the sampling problem is more severe.
△ Less
Submitted 2 June, 2020; v1 submitted 12 November, 2019;
originally announced November 2019.
-
Bayesian inference for high-dimensional nonstationary Gaussian processes
Authors:
Mark D. Risser,
Daniel Turek
Abstract:
In spite of the diverse literature on nonstationary spatial modeling and approximate Gaussian process (GP) methods, there are no general approaches for conducting fully Bayesian inference for moderately sized nonstationary spatial data sets on a personal laptop. For statisticians and data scientists who wish to learn about spatially-referenced data and conduct posterior inference and prediction wi…
▽ More
In spite of the diverse literature on nonstationary spatial modeling and approximate Gaussian process (GP) methods, there are no general approaches for conducting fully Bayesian inference for moderately sized nonstationary spatial data sets on a personal laptop. For statisticians and data scientists who wish to learn about spatially-referenced data and conduct posterior inference and prediction with appropriate uncertainty quantification, the lack of such approaches and corresponding software is a significant limitation. In this paper, we develop methodology for implementing formal Bayesian inference for a general class of nonstationary GPs. Our novel approach uses pre-existing frameworks for characterizing nonstationarity in a new way that is applicable for small to moderately sized data sets via modern GP likelihood approximations. Posterior sampling is implemented using flexible MCMC methods, with nonstationary posterior prediction conducted as a post-processing step. We demonstrate our novel methods on two data sets, ranging from several hundred to several thousand locations, and compare our methodology with related statistical methods that provide off-the-shelf software. All of our methods are implemented in the freely available BayesNSGP software package for R.
△ Less
Submitted 30 June, 2020; v1 submitted 30 October, 2019;
originally announced October 2019.
-
Detected changes in precipitation extremes at their native scales derived from in situ measurements
Authors:
Mark D. Risser,
Christopher J. Paciorek,
Travis A. O'Brien,
Michael F. Wehner,
William D. Collins
Abstract:
The gridding of daily accumulated precipitation -- especially extremes -- from ground-based station observations is problematic due to the fractal nature of precipitation, and therefore estimates of long period return values and their changes based on such gridded daily data sets are generally underestimated. In this paper, we characterize high-resolution changes in observed extreme precipitation…
▽ More
The gridding of daily accumulated precipitation -- especially extremes -- from ground-based station observations is problematic due to the fractal nature of precipitation, and therefore estimates of long period return values and their changes based on such gridded daily data sets are generally underestimated. In this paper, we characterize high-resolution changes in observed extreme precipitation from 1950 to 2017 for the contiguous United States (CONUS) based on in situ measurements only. Our analysis utilizes spatial statistical methods that allow us to derive gridded estimates that do not smooth extreme daily measurements and are consistent with statistics from the original station data while increasing the resulting signal to noise ratio. Furthermore, we use a robust statistical technique to identify significant pointwise changes in the climatology of extreme precipitation while carefully controlling the rate of false positives. We present and discuss seasonal changes in the statistics of extreme precipitation: the largest and most spatially-coherent pointwise changes are in fall (SON), with approximately 33% of CONUS exhibiting significant changes (in an absolute sense). Other seasons display very few meaningful pointwise changes (in either a relative or absolute sense), illustrating the difficulty in detecting pointwise changes in extreme precipitation based on in situ measurements. While our main result involves seasonal changes, we also present and discuss annual changes in the statistics of extreme precipitation. In this paper we only seek to detect changes over time and leave attribution of the underlying causes of these changes for future work.
△ Less
Submitted 13 August, 2019; v1 submitted 15 February, 2019;
originally announced February 2019.
-
A probabilistic gridded product for daily precipitation extremes over the United States
Authors:
Mark D. Risser,
Christopher J. Paciorek,
Michael F. Wehner,
Travis A. O'Brien,
William D. Collins
Abstract:
Gridded data products, for example interpolated daily measurements of precipitation from weather stations, are commonly used as a convenient substitute for direct observations because these products provide a spatially and temporally continuous and complete source of data. However, when the goal is to characterize climatological features of extreme precipitation over a spatial domain (e.g., a map…
▽ More
Gridded data products, for example interpolated daily measurements of precipitation from weather stations, are commonly used as a convenient substitute for direct observations because these products provide a spatially and temporally continuous and complete source of data. However, when the goal is to characterize climatological features of extreme precipitation over a spatial domain (e.g., a map of return values) at the native spatial scales of these phenomena, then gridded products may lead to incorrect conclusions because daily precipitation is a fractal field and hence any smoothing technique will dampen local extremes. To address this issue, we create a new "probabilistic" gridded product specifically designed to characterize the climatological properties of extreme precipitation by applying spatial statistical analyses to daily measurements of precipitation from the GHCN over CONUS. The essence of our method is to first estimate the climatology of extreme precipitation based on station data and then use a data-driven statistical approach to interpolate these estimates to a fine grid. We argue that our method yields an improved characterization of the climatology within a grid cell because the probabilistic behavior of extreme precipitation is much better behaved (i.e., smoother) than daily weather. Furthermore, the spatial smoothing innate to our approach significantly increases the signal-to-noise ratio in the estimated extreme statistics relative to an analysis without smoothing. Finally, by deriving a data-driven approach for translating extreme statistics to a spatially complete grid, the methodology outlined in this paper resolves the issue of how to properly compare station data with output from earth system models. We conclude the paper by comparing our probabilistic gridded product with a standard extreme value analysis of the Livneh gridded daily precipitation product.
△ Less
Submitted 2 January, 2019; v1 submitted 11 July, 2018;
originally announced July 2018.
-
Spatially-Dependent Multiple Testing Under Model Misspecification, with Application to Detection of Anthropogenic Influence on Extreme Climate Events
Authors:
Mark D. Risser,
Christopher J. Paciorek,
Daithi Stone
Abstract:
The Weather Risk Attribution Forecast (WRAF) is a forecasting tool that uses output from global climate models to make simultaneous attribution statements about whether and how greenhouse gas emissions have contributed to extreme weather across the globe. However, in conducting a large number of simultaneous hypothesis tests, the WRAF is prone to identifying false "discoveries." A common technique…
▽ More
The Weather Risk Attribution Forecast (WRAF) is a forecasting tool that uses output from global climate models to make simultaneous attribution statements about whether and how greenhouse gas emissions have contributed to extreme weather across the globe. However, in conducting a large number of simultaneous hypothesis tests, the WRAF is prone to identifying false "discoveries." A common technique for addressing this multiple testing problem is to adjust the procedure in a way that controls the proportion of true null hypotheses that are incorrectly rejected, or the false discovery rate (FDR). Unfortunately, generic FDR procedures suffer from low power when the hypotheses are dependent, and techniques designed to account for dependence are sensitive to misspecification of the underlying statistical model. In this paper, we develop a Bayesian decision theoretic approach for dependent multiple testing and a nonparametric hierarchical statistical model that flexibly controls false discovery and is robust to model misspecification. We illustrate the robustness of our procedure to model error with a simulation study, using a framework that accounts for generic spatial dependence and allows the practitioner to flexibly specify the decision criteria. Finally, we apply our procedure to several seasonal forecasts and discuss implementation for the WRAF workflow.
△ Less
Submitted 14 November, 2017; v1 submitted 29 March, 2017;
originally announced March 2017.
-
Review: Nonstationary Spatial Modeling, with Emphasis on Process Convolution and Covariate-Driven Approaches
Authors:
Mark D. Risser
Abstract:
In many environmental applications involving spatially-referenced data, limitations on the number and locations of observations motivate the need for practical and efficient models for spatial interpolation, or kriging. A key component of models for continuously-indexed spatial data is the covariance function, which is traditionally assumed to belong to a parametric class of stationary models. Whi…
▽ More
In many environmental applications involving spatially-referenced data, limitations on the number and locations of observations motivate the need for practical and efficient models for spatial interpolation, or kriging. A key component of models for continuously-indexed spatial data is the covariance function, which is traditionally assumed to belong to a parametric class of stationary models. While convenient, the assumption of stationarity is rarely realistic; as a result, there is a rich literature on alternative methodologies which capture and model the nonstationarity present in most environmental processes. This review document provides a rigorous and concise description of the existing literature on nonstationary methods, paying particular attention to process convolution (also called kernel smoothing or moving average) approaches. A summary is also provided of more recent methods which leverage covariate information and yield both interpretational and computational benefits.
Note: the article is borrowed from Chapters 1 and 2 of the author's Ph.D. dissertation, joint with Catherine A. Calder.
△ Less
Submitted 7 October, 2016;
originally announced October 2016.
-
Nonstationary Spatial Prediction of Soil Organic Carbon: Implications for Stock Assessment Decision Making
Authors:
Mark D. Risser,
Catherine A. Calder,
Veronica J. Berrocal,
Candace Berrett
Abstract:
The Rapid Carbon Assessment (RaCA) project was conducted by the US Department of Agriculture's National Resources Conservation Service between 2010-2012 in order to provide contemporaneous measurements of soil organic carbon (SOC) across the US. Despite the broad extent of the RaCA data collection effort, direct observations of SOC are not available at the high spatial resolution needed for studyi…
▽ More
The Rapid Carbon Assessment (RaCA) project was conducted by the US Department of Agriculture's National Resources Conservation Service between 2010-2012 in order to provide contemporaneous measurements of soil organic carbon (SOC) across the US. Despite the broad extent of the RaCA data collection effort, direct observations of SOC are not available at the high spatial resolution needed for studying carbon storage in soil and its implications for important problems in climate science and agriculture. As a result, there is a need for predicting SOC at spatial locations not included as part of the RaCA project. In this paper, we compare spatial prediction of SOC using a subset of the RaCA data for a variety of statistical methods. We investigate the performance of methods with off-the-shelf software available (both stationary and nonstationary) as well as a novel nonstationary approach based on partitioning relevant spatially-varying covariate processes. Our new method addresses open questions regarding (1) how to partition the spatial domain for segmentation-based nonstationary methods, (2) incorporating partially observed covariates into a spatial model, and (3) accounting for uncertainty in the partitioning. In applying the various statistical methods we find that there are minimal differences in out-of-sample criteria for this particular data set, however, there are major differences in maps of uncertainty in SOC predictions. We argue that the spatially-varying measures of prediction uncertainty produced by our new approach are valuable to decision makers, as they can be used to better benchmark mechanistic models, identify target areas for soil restoration projects, and inform carbon sequestration projects.
△ Less
Submitted 10 June, 2018; v1 submitted 19 August, 2016;
originally announced August 2016.
-
Quantifying the effect of interannual ocean variability on the attribution of extreme climate events to human influence
Authors:
Mark D. Risser,
Daithi A. Stone,
Christopher J. Paciorek,
Michael F. Wehner,
Oliver Angelil
Abstract:
In recent years, the climate change research community has become highly interested in describing the anthropogenic influence on extreme weather events, commonly termed "event attribution." Limitations in the observational record and in computational resources motivate the use of uncoupled, atmosphere/land-only climate models with prescribed ocean conditions run over a short period, leading up to…
▽ More
In recent years, the climate change research community has become highly interested in describing the anthropogenic influence on extreme weather events, commonly termed "event attribution." Limitations in the observational record and in computational resources motivate the use of uncoupled, atmosphere/land-only climate models with prescribed ocean conditions run over a short period, leading up to and including an event of interest. In this approach, large ensembles of high-resolution simulations can be generated under factual observed conditions and counterfactual conditions that might have been observed in the absence of human interference; these can be used to estimate the change in probability of the given event due to anthropogenic influence. However, using a prescribed ocean state ignores the possibility that estimates of attributable risk might be a function of the ocean state. Thus, the uncertainty in attributable risk is likely underestimated, implying an over-confidence in anthropogenic influence.
In this work, we estimate the year-to-year variability in calculations of the anthropogenic contribution to extreme weather based on large ensembles of atmospheric model simulations. Our results both quantify the magnitude of year-to-year variability and categorize the degree to which conclusions of attributable risk are qualitatively affected. The methodology is illustrated by exploring extreme temperature and precipitation events for the northwest coast of South America and northern-central Siberia; we also provides results for regions around the globe. While it remains preferable to perform a full multi-year analysis, the results presented here can serve as an indication of where and when attribution researchers should be concerned about the use of atmosphere-only simulations.
△ Less
Submitted 28 September, 2016; v1 submitted 28 June, 2016;
originally announced June 2016.
-
Local likelihood estimation for covariance functions with spatially-varying parameters: the convoSPAT package for R
Authors:
Mark D. Risser,
Catherine A. Calder
Abstract:
In spite of the interest in and appeal of convolution-based approaches for nonstationary spatial modeling, off-the-shelf software for model fitting does not as of yet exist. Convolution-based models are highly flexible yet notoriously difficult to fit, even with relatively small data sets. The general lack of pre-packaged options for model fitting makes it difficult to compare new methodology in n…
▽ More
In spite of the interest in and appeal of convolution-based approaches for nonstationary spatial modeling, off-the-shelf software for model fitting does not as of yet exist. Convolution-based models are highly flexible yet notoriously difficult to fit, even with relatively small data sets. The general lack of pre-packaged options for model fitting makes it difficult to compare new methodology in nonstationary modeling with other existing methods, and as a result most new models are simply compared to stationary models. Using a convolution-based approach, we present a new nonstationary covariance function for spatial Gaussian process models that allows for efficient computing in two ways: first, by representing the spatially-varying parameters via a discrete mixture or "mixture component" model, and second, by estimating the mixture component parameters through a local likelihood approach. In order to make computation for a convolution-based nonstationary spatial model readily available, this paper also presents and describes the convoSPAT package for R. The nonstationary model is fit to both a synthetic data set and a real data application involving annual precipitation to demonstrate the capabilities of the package.
△ Less
Submitted 3 February, 2017; v1 submitted 30 July, 2015;
originally announced July 2015.
-
Regression-based covariance functions for nonstationary spatial modeling
Authors:
Mark D. Risser,
Catherine A. Calder
Abstract:
In many environmental applications involving spatially-referenced data, limitations on the number and locations of observations motivate the need for practical and efficient models for spatial interpolation, or kriging. A key component of models for continuously-indexed spatial data is the covariance function, which is traditionally assumed to belong to a parametric class of stationary models. How…
▽ More
In many environmental applications involving spatially-referenced data, limitations on the number and locations of observations motivate the need for practical and efficient models for spatial interpolation, or kriging. A key component of models for continuously-indexed spatial data is the covariance function, which is traditionally assumed to belong to a parametric class of stationary models. However, stationarity is rarely a realistic assumption. Alternative methods which more appropriately model the nonstationarity present in environmental processes often involve high-dimensional parameter spaces, which lead to difficulties in model fitting and interpretability. To overcome this issue, we build on the growing literature of covariate-driven nonstationary spatial modeling. Using process convolution techniques, we propose a Bayesian model for continuously-indexed spatial data based on a flexible parametric covariance regression structure for a convolution-kernel covariance matrix. The resulting model is a parsimonious representation of the kernel process, and we explore properties of the implied model, including a description of the resulting nonstationary covariance function and the interpretational benefits in the kernel parameters. Furthermore, we demonstrate that our model provides a practical compromise between stationary and highly parameterized nonstationary spatial covariance functions that do not perform well in practice. We illustrate our approach through an analysis of annual precipitation data.
△ Less
Submitted 4 February, 2015; v1 submitted 6 October, 2014;
originally announced October 2014.