Analysis · Concept hubDeep story

Statistics and Inference

1662 CE19th-century France and Britain (Laplace and Fisher)

Concept

The mathematics of reasoning from observed cases to a wider population, process, or effect. Its central task is not merely gathering more numbers but exposing who was counted, what was measured, and which comparisons and assumptions support a conclusion.

Understand it in one breath

"How confidently can we draw conclusions about a population from a sample?" Sample means, confidence intervals, and p-values appear in medical trials, political polls, and A/B tests. A standard Central Limit Theorem says that for independent, identically distributed observations with finite variance, the standardized sample mean approaches a normal distribution as the sample grows. A large sample does not repair a bad frame, nonresponse, or measurement bias; it can be precisely wrong.

At a glance

-4-3-2-101234

Key formula

Xˉnμσ/ndN(0,1)(a standard Central Limit Theorem)\frac{\bar X_n-\mu}{\sigma/\sqrt n} \xrightarrow{d} N(0,1) \quad \text{(a standard Central Limit Theorem)}

Worked examples

  1. 1

    Q.A random sample of 100 has mean 50 and standard deviation 10. What is the approximate normal 95% confidence interval?

Key moments

1662 CE

Graunt — reading a city through mortality bills

John Graunt assembled London mortality bills to compare patterns by cause, sex, and place. It was an important inference from incomplete administrative records, not a modern sample survey or census.

1812 CE

Laplace — error, probability, and population in one calculus

The Théorie analytique des probabilités joined astronomical error, inverse probability, and population data. It was a powerful synthesis of several traditions, not one book inventing statistics alone.

1900 CE

Pearson — comparing observed and expected cells

The chi-square goodness-of-fit test measured disagreement between categorical counts and model expectations. It was not the first test of every kind, and a small p-value is not the probability that a hypothesis is false.

1908 CE

Gosset — accounting for uncertainty in small samples

Guinness brewer William Gosset published the t distribution as “Student” for comparing means when population variance is unknown and samples are small. Industrial secrecy and repeated production shaped the problem.

1926 CE

Fisher — designing the comparison through randomization

At Rothamsted, randomized treatment allocation was joined to replication and blocking. The emphasis shifted from choosing a formula after observation toward designing data that could sustain a comparison.

1950 CE

Mahalanobis and India’s National Sample Survey

Stratified and multistage designs sought to represent regional, rural, and urban diversity without enumerating everyone. Representativeness comes from frames, selection probabilities, and fieldwork—not sample size alone.

Modern applications

Clinical trials, opinion polling, machine-learning metrics, A/B tests, and diagnostic accuracy — evidence-based decision-making itself.

Beyond MathVoyage

Loading…