The Nature of Statistics and Reality

Statistics is fundamentally about probabilities and the likelihood of different outcomes. Imagine spending a day at Belmont Park, where you decide to bet on the winner of a horse race. The betting board shows that Secretariat is the favorite with odds of 1.5. The lower the odds, the more likely the horse is considered to win. If you bet $100, a winning ticket would return $150.

But before placing your bet, ask yourself one simple question: can you be certain that Secretariat will win?

The answer is no.

Why? Because no one can control the countless random events that may influence the outcome. Secretariat may become unsettled before the race, suffer a minor injury, get boxed in by another horse, or simply have an unusually poor performance that day. These are stochastic events—random influences that cannot be predicted or controlled with certainty.

Statistics does not eliminate uncertainty. Instead, it quantifies uncertainty and helps us make the best possible decisions based on the information available.

Now consider another question. How does a bookmaker decide that Secretariat should have odds of 1.5? The answer is simple: by analyzing information.

By analyzing such information, bookmakers reduce uncertainty. They cannot predict the outcome with certainty, but they can make informed estimates that allow them to stay in business over the long term.

The objective of any survey is to describe reality as accurately as possible. Since it is rarely practical to observe every individual, we rely on carefully designed samples to represent the whole population. Reality exists independently of our measurements. Statistics is our best attempt to describe it as faithfully as possible.

Why Sampling Design Matters

This is perhaps the most important aspect to consider. Imagine that you want to estimate the average stem diameter of all trees in a forest stand. The stand contains 204 trees, each of which can be assigned a unique number from 1 to 204. Rather than measuring every tree, we select a random sample. The selection should be completely random, much like the numbers drawn in a lottery. Every tree should have an equal probability of being selected. The next obvious question is: How many trees should we sample? In other words, how many randomly selected trees are needed to obtain a reliable estimate? That brings us to the next section.

Why Sample Size Matters

How many samples you need is a crucial part in designing your survey.

The answer depends on the precision required to answer the question. Every sample is influenced by random variation, often referred to as stochastic effects. With only a few observations, estimates such as the mean, variance and the underlying distribution may vary considerably simply due to chance. As the sample size increases, these random effects become less influential and the precision of the estimates improves.

A simple way to illustrate this is by repeatedly resampling the same data set with replacement (bootstrapping). Estimates based on a sample size of five observations typically show much greater variation than estimates based on twenty or fifty observations. The larger the sample, the more stable and reliable the estimates become.

The important question is therefore not “How many samples can I collect?” but rather “What level of precision is required to support the decision?”

Figure 1 illustrates how increasing the sample size progressively reduces the influence of random variation on statistical estimates.

Figure 1  shows the standard deviation changes as the sample size increases. The important insight is that stochastic fluctuations diminish as the sample size increases. With small sample sizes (n < 30), your survey may not represent reality in a convincing way. Once the fluctuations remain within the dashed lines, the estimates provide an acceptable representation of reality. The standard deviation indicates the precision of the statistical estimates. Notice that even at sample sizes greater than 500, there are still clear signs of stochastic variation.


Precision Comes at a Cost

Increasing the sample size improves the precision of the estimates, but it also increases the time and cost required to collect the data. Fortunately, the gain in precision is not proportional to the increase in sample size. Figure 2 illustrates how increasing the sample size progressively reduces the influence of random variation on the estimates.

Figure 2  shows the effect of sample size on standard error (SE). This statistic provides information about the precision of the estimated mean. Mathematically, the standard error is used to construct confidence intervals around the estimated mean. The curve in the plot pretty much follows the inverse function (1/x) where the precision of your mean estimate is very poor when based on few samples. However the good thing is that the precision improves fast by increased sample sizes. At sample sizes between 50 and 100 the improvement in precision starts to fade. In practical terms this implies that chasing better precision not only elevates the cost of the your survey - the payback in improved precision is almost negligible. So clearly, the trade-off decision would be at the knee-point, marked by the circle, of the curve (n = 50 to 100 samples).


As shown in Figure 2, precision improves rapidly when the sample size increases from only a few observations. Beyond that, however, each additional sample contributes progressively less to the overall precision. This phenomenon is known as the point of diminishing returns.

The important question is therefore not how many samples can I afford. Instead you should ask yourself how much uncertainty you may accept. The answer is a trade-off between what you need and what you can afford. At sample sizes of 500 and above the rate of return in reduced uncertainty practically diminish, it will just drain your wallet.

The example shown in the figure 2, most of the improvement in precision is achieved with approximately 50–100 observations. Beyond that point, the confidence intervals become only marginally narrower. However, this should not be interpreted as a universal recommendation. The required sample size always depends on the natural variation of the population being studied and the level of precision needed to support the final decision. By calculating the the variance based on samples you can, by setting your demand of uncertainty, calculate the number of samples need to get precision you need,

Another aspect is what precision and accuracy do you need. To complicate it a bit for you–according to trust able dictionaries these word are synonyms. However, we need introduce a distinction between meaning of the words, Accuracy describes how well a measure is to it’s true size while precision tells you something about repeated sampling of items. An invented example might be the pulse frequency of pacemakers. A flaw in production process cause slightly different frequencies–but each pacemaker keeps ticking accurately at its frequency. Figure 3 shows an example of that, however, based on what statisticians use to describe as stochastic processes which are random deviations from expected probabilities. Let’s take an example from the casino. A roulette player has a strategy to bet on red numbers first after the ball has landed on black numbers five times repeatedly. Will the player become a winner using that strategy? Of course not! Why? Because every spin of the wheel is independent any prior result. If we skip the green slot on the wheel for simplicity, the probability is fifty-fifty between black and red slots. As stochastic processes doesn’t adjust probabilities, the likelihood stays the same for each spin of the wheel. however, regarding the law of large numbers, the actual probability will eventually close in at the expected probability.

Figure 3  shows the accuracy of the assessed mean value a sampled population and how it improves by increasing number of samples. The accuracy (gray line) shows how the estimate becomes more regressed to the true population mean (red line). The stochastic effects are still there but the influence hampers with increasing sample sizes. We can, with a bit good will, see different stages. At sample sizes less than 20 the accuracy of the sample mean may route you to wrong decision. To rely on small sample size is risky business. At sample sizes from 50 - 100 you will be at safer grounds, however, the level safety depends on the consequences of what a wrongful decision have. If you are to decide whether management actions are needed or not, then you would make a safe enough decision. If your survey is about lethal risks concerning the dosage of a drug, then you would need 1000 samples or more to reach a reliable decision. Lastly, notice that sampling errors are still observable at large sample sizes though considerably less influential. If you can’t accept this level of accuracy, the only solution is measure the complete population.


Designing Studies That Answer the Right Question

Collecting data is relatively easy. Collecting the right data is crucial.

Before a single observation is made, the objective of the survey must be clearly defined. What question needs to be answered? What level of precision is required? Which variables need to be measured? How should the samples be collected to provide an unbiased representation of reality?

Figure 4  shows a normal distribution based on standardized values of a variable of interest. The mean is therefore located at 0 (blue line). Values smaller than the mean become negative, and values larger than the mean become positive, while the original variation of the data remains unchanged. The original values can easily be recovered by adding the original mean to the standardized values. The y-axis represents the probability of observing different values. Observations close to the mean are the most likely, whereas increasingly extreme values become progressively less likely. The gray shaded area represents virtually all possible outcomes. Mathematically, the tails of the distribution never actually reach zero—they approach it asymptotically. This is not a weakness of statistics. In practice, the probabilities in the extreme tails become so small that they have no influence on real-world decisions. An equally important property is that the entire gray shaded area represents every possible outcome, therefore the total probability equals 1. Consequently, the probability of observing a value below the mean is 0.5, leaving a probability of 0.5 of observing a value above the mean. Having established how probabilities are distributed, we can now understand confidence limits. The red shaded tails together contain only 1% of the total probability (p = 0.01), leaving 99% within the central part of the distribution. Because the distribution is symmetric, each tail contains a probability of only 0.005. The black dotted lines indicate commonly used confidence limits. They are tools for reducing uncertainty in decision-making. A confidence limit of 68.3% corresponds to one standard deviation (SD). To calculate confidence limits you need the standard error (SE) of the mean estimate. That statistic gives the precision of your mean estimate and is primarily used to quantify the precision of an estimated mean (see Figure 2). You can calculate any desired confidence limit of desire but 90%, 95%, and 99% are common alternatives, where uncertainty is reduced to 10%, 5% and 1% respectively. Your job is to assess the risk that comes with a wrongful decision.


A well-designed survey often requires fewer samples than a poorly planned one while producing more reliable answers. Careful planning before fieldwork begins therefore saves both time and money while increasing the value of the collected information.

Good surveys do not happen by chance. They are designed to answer the right question.

Correlation, Causation and Common Sense

Finding a statistical relationship between variables does not necessarily mean that one causes the other. Ice cream sales and drowning accidents both increase during summer, but buying ice cream does not cause people to drown. They simply share a common cause: warm weather.

Statistics can help identify and quantify relationships, but common sense and subject knowledge are needed to interpret them correctly. This is why statistical analyses should always be interpreted together with knowledge of the subject at hand.

Numbers alone rarely tell the whole story. Reliable conclusions are reached by combining appropriate statistical methods with critical thinking and an understanding of the underlying causes.

Final Notes

Statistics is often perceived as a collection of mathematical techniques. In reality, it is a way of reducing uncertainty so that better decisions can be made. Good statistics begins long before the first analysis is performed. It begins by asking the right questions, designing an appropriate survey, collecting representative data, and recognizing the uncertainty that will always remain.

Statistical methods do not create certainty, but when applied correctly, they provide the most reliable foundation for sound decisions. Therefore, collecting data that genuinely helps explain the question at hand is paramount.