Geostatistics
Branch of statistics for spatial and spatiotemporal datasets.
Geostatistics is a field of statistics that deals with data linked to locations in space or space-time. It was first created to estimate the probability of finding different ore grades in mining, but today it is used across many fields, such as petroleum and hydrogeology, meteorology, oceanography, forestry, soil science, and precision agriculture. It also plays a role in geography, especially for tracking disease spread, planning logistics for commerce or the military, and designing efficient spatial networks.
- field
- Statistics
- known_for
- Spatial and spatiotemporal data analysis, interpolation methods, kriging
- applications
- Mining, petroleum geology, hydrogeology, meteorology, geography, epidemiology, logistics
Lore & Background
Geostatistics is intimately related to interpolation methods but extends far beyond simple interpolation problems. Geostatistical techniques rely on statistical models based on random function (or random variable) theory to model the uncertainty associated with spatial estimation and simulation. A number of simpler interpolation methods/algorithms, such as inverse distance weighting, bilinear interpolation and nearest-neighbor interpolation, were already well known before geostatistics. Geostatistics goes beyond the interpolation problem by considering the studied phenomenon at unknown locations as a set of correlated random variables.
Let Z(x) be the value of the variable of interest at a certain location x. This value is unknown (e.g., temperature, rainfall, piezometric level, geological facies, etc.). Although there exists a value at location x that could be measured, geostatistics considers this value as random since it was not measured or has not been measured yet. However, the randomness of Z(x) is not complete. Still, it is defined by a cumulative distribution function (CDF) that depends on certain information that is known about the value Z(x). Typically, if the value of Z is known at locations close to x (or in the neighborhood of x) one can constrain the CDF of Z(x) by this neighborhood: if a high spatial continuity is assumed, Z(x) can only have values similar to the ones found in the neighborhood. Conversely, in the absence of spatial continuity Z(x) can take any value.
By applying a single spatial model on an entire domain, one makes the assumption that Z is a stationary process. It means that the same statistical properties are applicable on the entire domain. Several geostatistical methods provide ways of relaxing this stationarity assumption. In this framework, one can distinguish two modeling goals: estimating the value for Z(x), typically by the expectation, the median or the mode of the CDF f(z,x), and sampling from the entire probability density function f(z,x) by actually considering each possible outcome of it at each location. This is generally done by creating several alternative maps of Z, called realizations. Each realization is considered as a possible scenario of what the real variable could be. All associated workflows are then considering ensemble of realizations, and consequently ensemble of predictions that allow for probabilistic forecasting.
Reader's Guide
Geostatistics is significant as a specialized branch of statistics that addresses the unique challenges of spatial and spatiotemporal data. Its development for mining ore-grade prediction has expanded into a wide range of fields, including petroleum geology, hydrogeology, meteorology, oceanography, geography, forestry, environmental control, and precision agriculture. The discipline provides a rigorous framework for interpolation that goes beyond simpler methods by treating unknown values as random variables with spatial correlation, allowing for uncertainty quantification. Key techniques include kriging for estimation and various simulation methods for generating multiple realizations. Geostatistical algorithms are incorporated into geographic information systems (GIS), making them widely accessible. The field's legacy lies in its ability to model spatial continuity, handle uncertainty, and produce probabilistic forecasts, which are essential for decision-making in resource management, environmental science, and logistics. Its methods continue to evolve with Bayesian inference and machine learning approaches, maintaining relevance in modern data analysis.
Did You Know?
- Geostatistics was originally developed to predict probability distributions of ore grades for mining operations.
- Geostatistical algorithms are incorporated in many places, including geographic information systems (GIS).
- Geostatistics is applied in varied branches of geography, particularly those involving the spread of diseases (epidemiology), the practice of commerce and military planning (logistics), and the development of efficient s
- Geostatistics goes beyond the interpolation problem by considering the studied phenomenon at unknown locations as a set of correlated random variables.
More in Spatial Analysis & Gis 1-24
Spotted an error? Know more?
This is a living reference — every entry is fact-audited, and reader corrections feed straight into our audit queue. Suggest an edit · See this site's audit record
