Spatial atlas

VOYAGE ELEVEN · HOW NUMBERS CAME TO SEE PEOPLE

When Data Began to Read People — Who Was Counted, Who Disappeared?

Walk 934 years from land books and knotted records to sampling, randomization, algorithm audits, and privacy. Track when more numbers improve evidence—and when they more firmly hide missing people and chosen categories.

QUESTION FOR THE ROUTE

To reason from some people's records to a population, what should be counted, who should be sampled, which comparison should be built, and how should responsibility for the resulting decision remain visible?

WHAT THE LINE DOES NOT CLAIM

This line is neither a genealogy in which statistics was invented in one European city and diffused outward nor a progress story in which numbers neutrally copy reality. It separates administrative records, mathematical methods, field designs, and political categories while keeping public benefit and discriminatory use of the same tools in view.

The camera rests on each city while you read, then eases through the runway between scenes. Select any marker or scene link to travel in either direction.

The same route, four questions

A lens never hides a scene or proves a cause. It changes which places you compare first, and the URL preserves your choice.

The whole-route view keeps problem, cognitive change, place, movement, and evidence boundary at equal weight. Choose a lens when you want to test a different explanation against the same scenes.

18 scroll-controlled map scenes

Live map · 지도를 불러오는 중…

07 / 18 · 1854 CE

London

  1. 01 · 1086 CE

    Winchester · Administrative compilation

    Counting Land and Value before Counting People — Domesday Book

    William I's officials asked across England about landholders, ploughland, livestock, and value, then assembled the answers into a royal record. Turning local testimony into comparable fields let the crown see taxable and defensive resources from afar. This was not a census enumerating every resident, and almost all named individuals were landholders. What gets counted already reveals the purpose of rule.

    PAUSE AND ASK

    When a king wants to know a realm's 'value,' what becomes a field in the table before people themselves do?

    How the idea changed

    Turn local testimony and land memory into repeated questions about tenure, cultivation, livestock, and value that a distant ruler can compare.

    What this place made possible

    Winchester's royal treasury and scribal administration helped assemble nationwide returns into volumes, though the local inquiries did not all occur there.

    How it moved

    Royal questions → itinerant officials and local juries → estate records → Winchester compilation and custody → taxation and defense

    Do not overclaim

    Detailed as it was, this was not a census of every resident; unnamed people must not be read as nonexistent people.

    Evidence sources
    Stable link to this scene
    WinchesterCusco
  2. 02 · c. 1500 CE

    Cusco · Main activity

    Binding Imperial Quantities without Paper — Inka Khipu

    Inka record keepers, or khipukamayuq, encoded households, stored goods, labor, and tribute through the positions, knots, and colors of cords, reporting through the road network. Evidence that several keepers maintained matching accounts also suggests layers of verification and control. Cusco is an editorial anchor for imperial administration, not the production site of one surviving khipu, and not every meaning beyond numerical records has been deciphered.

    PAUSE AND ASK

    Without paper or an alphabetic script, how could a large empire check and recount households, labor, and food?

    How the idea changed

    Place quantities in cord hierarchy, position, knots, and color rather than flat writing, creating portable and aggregable administrative memory.

    What this place made possible

    Cusco was the imperial administrative capital where roads, storehouses, and tribute reports converged, while surviving khipu have varied Andean provenances.

    How it moved

    Local production, households, and labor → khipukamayuq accounts → duplicate records and upward reports → road network → storage, tribute, and mobilization

    Do not overclaim

    Numerical and accounting uses are well supported; not every color and knot has been decoded as narrative, nor was one system instantly invented at Cusco.

    Evidence sources
    Stable link to this scene
    CuscoLondon
  3. 03 · 1662 CE

    London · Publication

    Reading Urban Regularity from Weekly Death Lists — John Graunt

    London parishes issued weekly mortality bills for epidemic surveillance; Graunt recombined many years of them. Comparing causes of death, seasons, sex counts, and city size revealed population-level regularities invisible in any one death. Cause labels relied on searchers rather than modern medical examiners and many people were missing, so regularity in a table was neither complete registration nor a causal law.

    PAUSE AND ASK

    One death is a tragedy; why do thousands layered across years begin to look like recurring urban patterns?

    How the idea changed

    Reassemble lists of individual events into frequencies by cause, season, sex, and year, then reason about population-level rates and regularities.

    What this place made possible

    London's parish bills, plague surveillance, print market, and Royal Society made a record network that could be compared over time and debated publicly.

    How it moved

    Household death → parish searchers and weekly bills → print and sale → Graunt's reclassification and comparison → life tables, insurance, and demography

    Do not overclaim

    Regularities found in incomplete, nonmedical cause labels are not enlarged into complete registration, individual prediction, or causal effects.

    Evidence sources
    Stable link to this scene
    LondonStockholm
  4. 04 · 1749 CE

    Stockholm · Administrative compilation

    Joining Parish Tables into a Continuous National Record — Tabellverket

    Sweden's Tabellverket regularly combined parish reports of population, births, and deaths, letting the state compare itself over time. Repeated forms, rather than one grand total, exposed rates of change and regional differences. The tables served health and administration but rested on church and state categories and omissions; they were not a neutral modern database.

    PAUSE AND ASK

    What becomes visible when a country repeats the same tables each year instead of conducting one grand count?

    How the idea changed

    Move from a still photograph of total population to a continuous record of births, deaths, regional differences, and change over time.

    What this place made possible

    Stockholm's central administration combined standardized parish returns into national tables and institutionalized continuity of comparison.

    How it moved

    Parish registers → standardized local tables → central Tabellverket compilation → annual comparison → health, military, and administrative debate

    Do not overclaim

    Even a long official time series is not the whole reality independent of church-state categories, reporting capacity, and omission.

    Evidence sources
    Stable link to this scene
    StockholmLondon
  5. 05 · 1801 CE

    London · Administrative compilation

    Counting a Country to Ask How Many It Could Feed — The 1801 Census

    Amid war, bad harvests, and debate after Malthus, Parliament passed the 1800 Act and conducted an official census in England, Wales, and Scotland in 1801. Counts of households, people, and broad occupations became numbers for national planning. Household schedules completed by residents took shape only in 1841, and groups such as some soldiers and sailors were excluded, so 1801 was not yet a fully modern census.

    PAUSE AND ASK

    If a state debates food shortage without knowing how many people exist, what machinery of inquiry does it build?

    How the idea changed

    Move from fragments of estimates, taxes, and church records toward direct national enumeration by one date and broad categories.

    What this place made possible

    Parliament and central administration in London turned war, grain, and population debate into law and a national count, while local officials and schoolmasters collected returns.

    How it moved

    Food and population debate → 1800 Census Act → local household enumeration → central clerical processing → national totals for Parliament and administration

    Do not overclaim

    The first official count was not identical to the individual household schedules of 1841; excluded groups and coarse occupation categories remain visible.

    Evidence sources
    Stable link to this scene
    LondonBrussels
  6. 06 · 1835 CE

    Brussels · Publication

    Turning an Average from Summary into a Human Figure — Quetelet's Average Man

    Quetelet carried error curves from astronomy into averages and regularities in height, crime, and marriage, making the 'average man' central to his 1835 social physics. A powerful way to summarize groups appeared, but treating averages across many traits as one real person or desirable norm erases individual variation and institutional causes. An average answers a question; it is not humanity's master copy.

    PAUSE AND ASK

    Is an average across people a useful summary, or a 'normal human' who somehow really exists?

    How the idea changed

    Carry error curves and probability from astronomical observations into social data, seeking stable group patterns behind individual variation.

    What this place made possible

    Brussels' Royal Observatory, Belgian statistical administration, and international congress work supplied data and authority linking measurement to social theory.

    How it moved

    Laplace and Fourier on probability and error → astronomy → Belgian social statistics → 1835 social physics → international statistical standards

    Do not overclaim

    Stability of an average proves neither an ideal individual, racial essence, nor natural law of crime, and it does not erase variation or subgroups.

    Evidence sources
    Stable link to this scene
    BrusselsLondon
  7. 07 · 1854 CE

    London · Field investigation

    Layering Death Addresses on a Map to Suspect the Water — John Snow

    Snow investigated cholera death addresses and water use in Soho, comparing the concentration around the Broad Street pump with distant drinkers and exceptions such as a workhouse with its own well. The map made the case visible but was not the whole evidence. The outbreak was already waning when the handle was removed, and the map alone did not finally prove germ theory or a single cause.

    PAUSE AND ASK

    When dots cluster around one pump, what beyond the dots must be asked before making a causal claim?

    How the idea changed

    Combine death counts with addresses, water exposure, and exceptions, turning spatial clustering into comparative epidemiological evidence.

    What this place made possible

    London's street addresses, detailed maps, death registration, and competing water systems made it possible to compare who drank which water where.

    How it moved

    Death certificates and addresses → household field interviews → pump and water-supply comparison → map and tables → local committee and public-health debate

    Do not overclaim

    Neither the famous map nor handle removal was a lone proof; timing of decline and Snow's broader water-supply comparisons remain part of the evidence.

    Evidence sources
    Stable link to this scene
    LondonLondon
  8. 08 · 1886 CE

    London · Publication

    Seeing a Return in the Cloud of Parents and Children — Galton's Regression

    Galton plotted family heights and noticed that children of exceptionally tall or short parents tended, on average, to lie closer to the overall mean, calling the pattern regression toward mediocrity. Reading how two variables move together became a new visual and mathematical project. Correlation does not establish causation, and Galton joined these tools to discriminatory political projects of hereditary ranking and eugenics. Useful mathematics and harmful purpose cannot be separated by celebration.

    PAUSE AND ASK

    When two values move together, is the pattern a cause, a prediction, or a trace of how the data were selected?

    How the idea changed

    Move from separate row and column averages to reading relationship strength and regression toward a center in a cloud of paired observations.

    What this place made possible

    London's exhibitions, Royal Society, anthropometric laboratory, and family-data network gathered measurements while restricting who was measured and which traits counted as valuable.

    How it moved

    Quetelet's distributions and averages → Galton's family, seed, and height data → regression and correlation → Pearson's formalization → biometrics and social science

    Do not overclaim

    Correlation is not causation, and regression supplies no warrant for hereditary ranking or a eugenic command that states should select people.

    Evidence sources
    Stable link to this scene
    LondonWashington DC
  9. 09 · 1890 CE

    Washington DC · Administrative compilation

    Turning Human Answers into Holes a Machine Could Read — Hollerith

    The U.S. Census transferred each person's age, sex, race, marital status, citizenship, and other answers into positions on punched cards, then used electrical readers and tabulators. Faster processing made it possible to recount combinations of categories and helped launch a data-processing industry. The machine was no more neutral than the questionnaire: political choices had already determined which questions and racial categories became card fields.

    PAUSE AND ASK

    When questionnaire answers become holes in cards, calculation speeds up—but what about people becomes fixed?

    How the idea changed

    Move from clerks summing written answers to electrical machines repeatedly reading and cross-tabulating standardized individual records.

    What this place made possible

    Washington's federal census organization combined national schedules, machine contracts, and category definitions into one operating network for large-scale tabulation.

    How it moved

    Household schedules → person-level punched cards → pins, mercury, and electric counters → cross-tabulations → office-machine and data-processing industries

    Do not overclaim

    Innovation in processing speed did not make race, sex, and citizenship categories natural facts, and most original 1890 schedules were later destroyed by fire.

    Evidence sources
    Stable link to this scene
    Washington DCLondon
  10. 10 · 1900 CE

    London · Publication

    Measuring the Mismatch between Observed and Expected Cells — Pearson's Chi-Square

    Karl Pearson published a general method that combined differences between observed category counts and model-expected counts into one goodness-of-fit statistic. It moved judgment from 'looks close' toward comparison with sampling variation. It did not single-handedly invent every hypothesis test or today's culture of p-values, and Pearson's biometric methods and institutions were deeply entangled with a eugenic program.

    PAUSE AND ASK

    When observed and model-expected counts differ, how large must the mismatch be before chance alone looks implausible?

    How the idea changed

    Scale each cell's discrepancy by its expected count, sum one statistic, and compare model fit with sampling variation.

    What this place made possible

    UCL's biometric laboratory, data collection, teaching, and specialist journals supplied an institution for repeated calculation and training in new tests.

    How it moved

    Galton on variation and correlation → Weldon's biological data → Pearson's distributions and fit calculations → 1900 chi-square paper → testing in biology, medicine, and social science

    Do not overclaim

    One chi-square statistic did not complete all hypothesis testing or modern p-value practice; conditions such as adequate expected counts and independence still matter.

    Evidence sources
    Stable link to this scene
    LondonDublin
  11. 11 · 1908 CE

    Dublin · Main activity

    Learning Small-Sample Uncertainty from a Pint of Beer — 'Student'

    At Guinness, William Gosset faced barley and malt experiments where only a few observations were practical and population variance was unknown. His 1908 paper under the name 'Student' described a distribution whose heavier tails change with sample size. The resulting methods still rely on assumptions such as independence and a chosen population model; they do not turn a small convenience sample into a representative one.

    PAUSE AND ASK

    With a small sample and unknown population spread, what goes missing if uncertainty is calculated as though the sample were large?

    How the idea changed

    Reflect uncertainty in the estimated standard deviation through degrees of freedom and heavier tails, making small-sample uncertainty more honest.

    What this place made possible

    The scale of Guinness and the cost of testing ingredients supplied a recurring industrial problem: making quality decisions from few experiments and long records.

    How it moved

    Brewing and barley quality → Gosset's study with Pearson → 1908 'Student' paper → Fisher's degrees-of-freedom interpretation and extension → experiments and estimation

    Do not overclaim

    The t distribution is no magic license for tiny samples; independence, model assumptions, and sampling bias must be checked separately.

    Evidence sources
    Stable link to this scene
    DublinHarpenden
  12. 12 · 1926 CE

    Harpenden · Experiment

    Letting Chance Choose the Plot instead of the Researcher — Rothamsted

    Fisher and Rothamsted colleagues combined replication, blocks, and deliberate random allocation to separate fertilizer or crop effects from irregular soil. Chance supplied a reference comparison instead of letting the investigator pick plots that looked favorable. The 1926 paper is a landmark formulation, not the absolute first randomized experiment, and Fisher's active support for eugenics remains part of the history.

    PAUSE AND ASK

    If plots differ before treatment, how can fertilizer effects be separated from the effects of already-good soil?

    How the idea changed

    Use replication, blocking, and random allocation to handle pretreatment variation in design, making error estimation a condition of comparison rather than an afterthought.

    What this place made possible

    Rothamsted's long-running fields, varied soils, agricultural team, and accumulated yields formed a workshop for testing design against real heterogeneity.

    How it moved

    Long-term field data → debate among Gosset, Fisher, and field researchers → 1926 allocation principles → randomized blocks and ANOVA → scientific and industrial experiments

    Do not overclaim

    Fisher is not made the first person ever to imagine randomization; powerful design work and his eugenic activism are both recorded.

    Evidence sources
    Stable link to this scene
    HarpendenPrinceton
  13. 13 · 1936 CE

    Princeton · Field investigation

    When Millions of Replies Lost to a Smaller Poll — Literary Digest and Gallup

    The Literary Digest mailed roughly ten million ballots and received more than 2.3 million replies, yet wrongly forecast a Landon victory. Its frame and response process were both biased. Gallup's much smaller quota sample, run from Princeton, correctly called Roosevelt's victory. This was not a simple triumph of modern probability sampling: the Gallup design used quotas and its method suffered a major failure in 1948.

    PAUSE AND ASK

    If tens of thousands of answers can beat 2.3 million, what must be asked before sample size?

    How the idea changed

    Shift the center of accuracy from 'How many replied?' to the frame, selection mechanism, and who did not respond.

    What this place made possible

    The Princeton-based American Institute of Public Opinion combined newspaper syndication, demographic quotas, and commercial survey methods into repeatable election forecasts.

    How it moved

    Telephone and automobile lists plus mail response → Literary Digest mass sample → Gallup quota sample and newspaper network → election outcome comparison → modern survey-method debate

    Do not overclaim

    The failure is not reduced to an affluent frame alone; nonresponse bias also matters, and Gallup's quota sample is not rewritten as a modern probability sample.

    Evidence sources
    Stable link to this scene
    PrincetonLondon
  14. 14 · 1948 CE

    London · Experiment

    Separating Treatment Hope from Comparison — The MRC Streptomycin Trial

    The British MRC tuberculosis trial used central allocation based on random numbers to assign scarce streptomycin and compared a treatment group with a bed-rest control group. Radiograph readers assessed images without knowing allocation. It was a major transition in clinical trials, but not the first controlled trial and not double-blind in every respect. A statistical difference also does not promise the same benefit to every patient.

    PAUSE AND ASK

    To keep hope for a new drug and clinician choice from entering the comparison, who should know or control allocation?

    How the idea changed

    Build comparison groups through central random allocation and blinded image reading, reducing selection and assessment bias by design.

    What this place made possible

    The MRC's London committee and statistical center coordinated allocation of a scarce drug, a common protocol, and independent reading across hospitals.

    How it moved

    American streptomycin supply → British MRC multicenter trial → central random numbers and concealed allocation → radiographic and clinical comparison → 1948 BMJ report

    Do not overclaim

    A landmark in central randomization, it is not labeled the first controlled trial or fully double-blind, and its access and ethical context is kept distinct from today's.

    Evidence sources
    Stable link to this scene
    LondonKolkata
  15. 15 · 1950 CE

    Kolkata · Field investigation

    Trying to See Regional Difference without Counting an Entire Country — India's National Sample Survey

    Mahalanobis and the Indian Statistical Institute network launched the National Sample Survey in 1950 to repeatedly measure consumption, work, land, and production across rural and urban India. Stratified multistage designs and field organization produced timely planning evidence for a large and diverse new nation. A sample survey is not merely a cheaper miniature census but a separate design carrying uncertainty, and a national average cannot stand in for every region or group.

    PAUSE AND ASK

    When hundreds of millions across many regions cannot all be surveyed, how can a smaller sample retain national diversity?

    How the idea changed

    Design a sample survey as its own inferential instrument with strata, multistage selection, weights, and field checks—not a shrunken census.

    What this place made possible

    The Indian Statistical Institute in Kolkata linked theory, training, field investigators, and planning into infrastructure for repeated continental-scale surveys.

    How it moved

    Bengal crop surveys and sampling research → ISI methods and training → central-government support → 1950 NSS field rounds → consumption, employment, land, and planning data

    Do not overclaim

    The NSS is neither one person's lone invention nor perfect representation of every region; sampling error, nonsampling error, and state categories remain visible.

    Evidence sources
    Stable link to this scene
    KolkataBerkeley
  16. 16 · 1975 CE

    Berkeley · Publication

    When the Overall Rate and Department Rates Pointed Opposite Ways — Berkeley Admissions

    In UC Berkeley graduate admissions data, women's overall admission rate appeared lower, while department-level comparisons greatly reduced or sometimes reversed the gap. Applicants had entered departments with very different selectivity at unequal rates, allowing aggregation to reverse a pattern. The analysis was not an automatic verdict that discrimination was absent. Which variables to condition on depends on a causal question, including constraints shaping department choice.

    PAUSE AND ASK

    When the overall admission rate and department rates point in opposite directions, which number should be trusted?

    How the idea changed

    Stop treating one aggregate table as the conclusion; decompose group composition and selection paths, then ask which comparison answers the causal question.

    What this place made possible

    Berkeley's actual admissions records and statistical community turned a textbook paradox into a real question about organization and department choice.

    How it moved

    Graduate applications and admissions → aggregate comparison by sex → department disaggregation → 1975 analysis → teaching confounding, selection, and causal inference

    Do not overclaim

    Aggregate reversal does not automatically prove or disprove discrimination; whether to condition on department depends on a causal model including what shaped application choices.

    Evidence sources
    Stable link to this scene
    BerkeleyCambridge, MA
  17. 17 · 2018 CE

    Cambridge, MA · Performance audit

    Recounting the Faces Hidden under Average Accuracy — Gender Shades

    Joy Buolamwini and Timnit Gebru evaluated commercial gender classifiers separately across intersections of perceived skin type and gender. Large performance gaps and the lighter-skinned, male-heavy composition of existing benchmarks became visible beneath one average accuracy. The study audited four classifiers and constructed datasets at a particular time; it was neither a theory of every face-recognition system nor a justification for reducing gender identity to a binary classification task.

    PAUSE AND ASK

    Even when overall accuracy looks high, how can we reveal which intersectional group bears most errors?

    How the idea changed

    Move from one average score to error tables disaggregated jointly by perceived skin type and gender, making dataset composition and performance gaps auditable.

    What this place made possible

    MIT Media Lab's research and public-paper setting, plus a new benchmark built from parliamentarian images, enabled external comparison of commercial APIs.

    How it moved

    Audit existing face datasets → Pilot Parliaments Benchmark → intersectional evaluation of commercial classifiers → public paper and company response → algorithm-auditing movement

    Do not overclaim

    Results for particular 2018 gender classifiers are not permanent rankings of all face recognition, nor do they naturalize binary gender classification itself.

    Evidence sources
    Stable link to this scene
    Cambridge, MAWashington DC
  18. 18 · 2020 CE

    Washington DC · Data release

    Using Mathematics to Obscure the People Revealed by Publication — Differentially Private Census Tables

    The U.S. Census Bureau applied a differential-privacy-based disclosure avoidance system to parts of the 2020 Census release to limit the risk of combining tables to reidentify people. Adding controlled random noise makes a privacy budget explicit but creates tension with accuracy, especially for small places and groups. Differential privacy is a framework for designing risk and utility together, not one magic algorithm that guarantees anonymity.

    PAUSE AND ASK

    If many exact small-area tables reveal individuals, through which number should privacy and usefulness be negotiated?

    How the idea changed

    Move from deleting names to a mathematical release contract limiting how much one person's inclusion can change the output distribution.

    What this place made possible

    The Census Bureau in Washington had to reconcile legal confidentiality, massive geographic tables, reidentification research, and public data demand in one release system.

    How it moved

    Reconstruction research on 2010 data → differential-privacy framework → 2020 Disclosure Avoidance System → noise, invariants, and quality metrics → researcher and community feedback

    Do not overclaim

    Differential privacy is neither unconditional anonymity nor error-free release; implementation, privacy budget, and small-group accuracy require public scrutiny.

    Evidence sources
    Stable link to this scene

PAUSE THE FILM · AUDIT THE DATA

A large sample can still miss the population

A town of 1,000 people has a true support rate of 52% in this constructed example. Change how people enter the sample and how many are counted. Then split the same data into groups and watch the conclusion reverse.

Who is the population?

Check whether the people receiving the conclusion match the actual list.

Who could enter?

Inspect the frame, selection, and nonresponse before sample size.

What will this number decide?

Who bears the cost of error is part of the result.

Same population, different ways into the sample

The values are deterministic teaching examples, not a promise that every random sample behaves exactly this way.

Absolute error: 8%

In this constructed sequence, a larger probability sample reduces random fluctuation. It does not repair bad measurement or nonresponse by itself.

TOUCH THE MATHEMATICS

When Data Began to Read People — Counting, Sampling, and the Politics of Inference

Begin with a land survey at Winchester and Inka khipu, then move through mortality bills, national tables, censuses, the average person, a cholera map, correlation, punched cards, sampling distributions, and randomized experiments. Across a failed mass poll, India's National Sample Survey, an aggregate reversal, intersectional algorithm audits, and differential privacy, eighteen scenes ask how data can be both a mirror of people and a lens built from categories, samples, and decisions.

Continue as an eighteen-scene cinematic journey

OPEN THE FULL MAP

Reconnect the conditions of statistics and inference

Revisit samples, uncertainty, intervals, and tests through formulas and examples that state what each can and cannot say.

Explore the full map