Concerning the Measurement of Global Mean Surface Temperature

Agreement concerning basic definitions, models, and measurement procedures is an objective of good science. Below are some questions and a paper concerning GMST. It is left to the reader to ascertain if these answers by Grok are reasonable.

Question 1.

Consider, for example, the concept of length in a Galilean spacetime. If a sequence of independent, identically distributed measurements X_1, X_2, \ldots ,X_n is made of the length μ of a rigid rod, then, by the central limit theorem, the sampling distribution of the mean \bar{X} of the sequence converges to a normal distribution with mean equal to the mean of the sample distributions. In the absence of bias, the mean of the sample distribution will be the exact length μ. Moreover, the variance of the data can be compared to that expected from knowledge of experimental errors inherent in the details of the measurement process. Unexpected discrepancies can then be evaluated and possible model weaknesses analyzed or unexpected sources of error sought.

Application of this measurement procedure to, for example, the temperature measurement of a homogeneous fluid body in a constant state would quite likely raise few concerns. On the other hand, the surface of the earth is not a homogeneous fluid body in a constant state. If a number of different observers each using different sets of surface points make estimates of the global mean surface temperature, then the sequence of values so obtained does not represent a list of independent, identically distributed random variables and the hypotheses of the central limit theorem do not hold. Rather, this situation is like unto a set of observers, ignorant of the precepts of special relativity, travelling at different, but high, velocities relative to each other, and attempting to ascribe a length to a rigid rod which does not lie in their respective rest frames. Since the model hypothesis of a Galilean spacetime is false, the very concept of length is physically meaningless and differences in length obtained by different observers cannot be attributed merely to technical difficulties connected with the measuring process but rather to the incorrect model assumption of a Galilean spacetime. Due to the relativity of simultaneity, different observers are, in reality, measuring the separations between different pairs of spacetime events. Discrepancies obtained by such observers cannot in this case be ascribed to experimental error which can in principle be overcome by sufficient ingenuity; rather, their explanation requires a fundamental change of philosophy and the recognition that each is measuring a different parameter. In this case, knowledge of the experimental errors involved and of the discrepancies between observers would hopefully lead to a change in the model hypothesis and the discovery of special relativity. This, in turn, would enable the observers to transform their measurements to the rest frame of the rod, thereby yielding consistent estimates of the true rest length.

Unfortunately, the measurement of global mean surface temperature turns out to be far more complex and discrepancies between different observers measuring different distributions cannot be so easily transformed away. Some authors, attribute statistical bias in global mean temperature estimates to an uneven distribution of weather stations. This attribution implicitly assumes that such errors are amenable to correction by more uniform sampling or, perhaps, by better statistical processing. Is there evidence to the earth’s continually temperature distribution does not posses similar mathematical pathology?

Question 2

Let us imagine once more that two observers, with distinct but very large relative velocity, make repeated measurements on the length of the same rigid rod so as to obtain confidence intervals. Assume also that both observers take the utmost possible care to avoid bias in their measurement techniques. After a certain number of measurements their respective confidence intervals for the length will inevitably be disjoint even though both have avoided any procedural bias. This is a simple consequence of the central limit theorem and the fact that, due to relativistic length dilation, the two observers are in effect sampling from distributions with distinct means. Suppose now that one of the observers attempts to reduce this discrepancy by taking additional measurements. No matter how careful he is, or how many measurements are made, the most that can be achieved, will be a reduction in the size of his confidence interval. The problem of disjoint intervals will not disappear. Compare this to a situation in which two climatologists independently measure the average temperature in some region, using different set of sample points, and find a similar unreconcilable discrepancy between their respective results. Since all climatologists are honest, both will dutifully increase the number of sample points in order to obtain, say, a more uniform or, perhaps, less biased sample. Notwithstanding their undoubted honesty, their best efforts and good intentions are doomed to failure. As in the relativistic situation, the distributions which each climatologist samples are distinct, and so the discrepancy cannot be so eliminated by further measurement. In reality, the temperature distribution over the surface of the earth, at any moment, is quite irregular and it seems improbable that two distinct sets of points on the surface taken at different times (or even at the same time) yield samples from the same temperature distribution. In our relativistic analogy, the discrepancy is easily corrected by a transformation to the rest frame of the rod. In the case of temperature, it would seem that one is compelled to make ad hoc assumptions concerning the regularity of the global temperature distribution. What assumptions concerning the regularity of the global temperature distribution have been made?

Question 3

To obtain a confidence interval for the (time) average temperature at a single point, assumptions have to be made concerning the temporal regularity of the temperature at the said point. In fact, the very existence an average temperature underlying the observed temporal fluctuations has to be assumed —after all, not every random variable has a mean. Thus, for the sake of simplicity, let us eliminate this difficulty from the discussion by considering the best possible scenario of a sphere whose surface temperature distribution does not change with time but is purely a function of position. Unfortunately, even in this case there exist temperature distributions which are intractable to any conceivable sampling procedure. Consider, for example, the case of a distribution which is equal to a constant c over the surface of a sphere except for a single point P at which the temperature is a delta function of weight normalized so that the average temperature over the sphere is exactly equal to \beta + c where β is non zero. Then any uniform random sample of points on the sphere would, with probability 1, fail to contain the point P and so yield a sample mean temperature equal to c. Thus, the sample mean would differ from the true mean by the “bias” β which could be any pre-prescribed large number. Infinitely many variants of this construction are possible. Moreover, weaker (physical) versions could be constructed without recourse to the theory of distributions. On the surface of the earth an infinity of sharp temperature jumps occur in such places as the boundaries of icebergs floating in warmer water, at shorelines which are ever shifting and changing and, indeed, at the border of every shadow. Moreover, when temperature is measured, there exists the eminently practical problem of deciding how much weight to allocate to sample points which might be so unlucky as to fall in such hostile and inconvenient places as on an iceberg in the middle of the ocean or at the center of a volcano. The human decisions necessary in such circumstances inevitably lead to a degree of arbitrariness which might well explain the “observed” 0.6 degrees Celsius per century (or whatever it is) temperature increase. Incidently, two icebergs of the same surface area may have considerably different volumes and, hence, correspondingly different thermodynamic contributions but that is another issue entirely.

In summary, there exists the possibility that discrepancies in the measurement of global mean surface temperature may be measure-theoretic in origin, and, therefore not amenable to correction by, say, increases in sample size and quality. Such bias is intrinsic to the situation and independent of the measurements of even the most conscientious climatologist. A finite sample of points cannot, reliably estimate a distribution of unknown form over a set of positive measure. Similarly, surface measurements cannot reliably estimate an unknown volume distribution. How do climatologists handle this problem without relying on an ad hoc model assumption?

© Philip Pennance

For a related discussion, please see the following pdf: