John W. Tukey




Image source: wikimedia.org

  • Born in 1915, in New Bedford, Massachusetts.
  • Mum was a private tutor who home-schooled John. Dad was a Latin teacher.
  • BA and MSc in Chemistry, and PhD in Mathematics
  • Awarded the National Medal of Science in 1973, by President Nixon
  • By some reports, his home-schooling was unorthodox and contributed to his thinking and working differently.

Taking a glimpse back in time

is possible with the American Statistical Association video lending library.


We’re going to watch John Tukey talking about exploring high-dimensional data with an amazing new computer in 1973, four years before the EDA book.

Look out for these things:

Tukey’s expertise is described as for trial and error learning and the computing equipment.

First 4.25 minutes

Today we’re taking this seriously: grab a pen and paper, because for the rest of the lecture, you’ll be doing the arithmetic and drawing the pictures yourself.

Setting the frame of mind

Excerpt from the introduction

This book is based on an important principle.


It is important to understand what you CAN DO before you learn to measure how WELL you seem to have DONE it.


Learning first what you can do will help you to work more easily and effectively.


This book is about exploratory data analysis, about looking at data to see what it seems to say. It concentrates on simple arithmetic and easy-to-draw pictures. It regards whatever appearances we have recognized as partial descriptions, and tries to look beneath them for new insights. Its concern is with appearance, not with confirmation.


Examples, NOT case histories


The book does not exist to make the case that exploratory data analysis is useful. Rather it exists to expose its readers and users to a considerable variety of techniques for looking more effectively at one’s data. The examples are not intended to be complete case histories. Rather they should isolated techniques in action on real data. The emphasis is on general techniques, rather than specific problems.

A basic problem about any body of data is to make it more easily and effectively handleable by minds – our minds, her mind, his mind. To this general end:

  • anything that make a simpler description possible makes the description more easily handleable.
  • anything that looks below the previously described surface makes the description more effective.


So we shall always be glad (a) to simplify description and (b) to describe one layer deeper. In particular,

  • to be able to say that we looked one layer deeper, and found nothing, is a definite step forward – though not as far as to be able to say that we looked deeper and found thus-and-such.
  • to be able to say that “if we change our point of view in the following way … things are simpler” is always a gain–though not quite so much as to be able to say “if we don’t bother to change our point of view (some other) things are equally simple.”



Consistent with this view, we believe, is a clear demand that pictures based on exploration of data should force their messages upon us. Pictures that emphasize what we already know–“security blankets” to reassure us–are frequently not worth the space they take. Pictures that have to be gone over with a reading glass to see the main point are wasteful of time and inadequate of effect. The greatest value of a picture is when it forces us to notice what we never expected to see.


Confirmation


The principles and procedures of what we call confirmatory data analysis are both widely used and one of the great intellectual products of our century. In their simplest form, these principles and procedures look at a sample–and at what that sample has told us about the population from which it came–and assess the precision with which our inference from sample to population is made. We can no longer get along without confirmatory data analysis. But we need not start with it.


The best way to understand what CAN be done is not longer–if it ever was–to ask what things could, in the current state of our skill techniques, be confirmed (positively or negatively). Even more understanding is lost if we consider each thing we can do to data only in terms of some set of very restrictive assumptions under which that thing is best possible–assumptions we know we CANNOT check in practice.

Exploration AND confirmation

Once upon a time, statisticians only explored. Then they learned to confirm exactly–to confirm a few things exactly, each under very specific circumstances. As they emphasized exact confirmation, their techniques inevitably became less flexible. The connection of the most used techniques with past insights was weakened. Anything to which confirmatory procedure was not explicitly attached was decried as “mere descriptive statistics”, no matter how much we learned from it.


Today, the flexibility of (approximate) confirmation by the jacknife makes it relatively easy to ask, for almost any clearly specified exploration, “How far is it confirmed?”


Today, exploratory and confirmatory can–and should–proceed side by side. This book, of course, considers only exploratory techniques, leaving confirmatory techniques to other accounts.


About the problems


The teacher needs to be careful about assigning problems. Not too many, please. They are likely to take longer than you think. The number supplied is to accommodate diversity of interest, not to keep everybody busy.


Besides the length of our problems, both teacher and student need to realise that many problems do not have a single “right answer”. There can be many ways to approach a body of data. Not all are equally good. For some bodies of data this may be clear, but for others we may not be able to tell from a single body of data which approach is preferred. Even several bodies of data about very similar situations may not be enough to show which approach should be preferred. Accordingly, it will often be quite reasonable for different analysts to reach somewhat different analyses.


Yet more–to unlock the analysis of a body of day, to find the good way to approach it, may require a key, whose finding is a creative act. Not everyone can be expected to create the key to any one situation. And to continue to paraphrase Barnum, no one can be expected to create a key to each situation he or she meets.


To learn about data analysis, it is right that each of us try many things that do not work–that we tackle more problems than we make expert analyses of. We often learn less from an expertly done analysis than from one where, by not trying something, we missed–at least until we were told about it–an opportunity to learn more. Each teacher needs to recognize this in grading and commenting on problems.


Precision

The teacher who heeds these words and admits that there need be no one correct approach may, I regret to contemplate, still want whatever is done to be digit perfect. (Under such a requirement, the write should still be able to pass the course, but it is not clear whether she would get an “A”.) One does, from time to time, have to produce digit-perfect, carefully checked results, but forgiving techniques that are not too distributed by unusual data are also, usually, little disturbed by SMALL arithmetic errors. The techniques we discuss here have been chosen to be forgiving. It is hoped, then, that small arithmetic errors will take little off the problem’s grades, leaving severe penalties for larger errors, either of arithmetic or concept.

Outline

  1. Scratching down numbers
  2. Schematic summary
  3. Easy re-expression
  4. Effective comparison
  5. Plots of relationship
  6. Straightening out plots (using three points)
  7. Smoothing sequences
  8. Parallel and wandering schematic plots
  9. Delineations of batches of points
  10. Using two-way analyses
  1. Making two-way analyses
  2. Advanced fits
  3. Three way fits
  4. Looking in two or more ways at batched of points
  5. Counted fractions
  6. Better smoothing
  7. Counts in bin after bin
  8. Product-ratio plots
  9. Shapes of distributions
  10. Mathematical distributions

Today we’re covering the first three (in bold), by hand.

Looking at numbers with Tukey

What’s a stem-and-leaf plot?

A quick way to sketch the shape of a batch of numbers using only the digits themselves:

  • Split each number into a stem (leading digit(s)) and a leaf (the next digit).
  • Write the stems down a column, smallest to largest.
  • Add each leaf to its stem’s row, in the order the data arrive.
  • Then re-sort the leaves within each row, smallest to largest.

Stem-and-leaf plot: still seen in introductory statistics texts, and still one of the fastest ways to see a small batch of numbers with just pen and paper.

The plot doubles as a histogram turned on its side, but keeps every original digit, so you can still read off the numbers, the median, quartiles, and so on.

Your turn: petrol prices

Unleaded 91 price (cents/litre) at 39 Sydney service stations, most recent report as of end of July 2026.

Real data from the NSW Government’s FuelCheck scheme, via the FuelCheck open data API (data.nsw.gov.au).

Show code
options(width = 60)
print(fuel_main$price, digits = 4)
 [1] 190.9 180.7 182.9 179.9 190.9 199.9 182.7 182.7 189.9
[10] 181.9 194.9 186.9 185.9 184.9 185.7 183.7 186.9 181.9
[19] 189.9 183.9 187.7 189.7 184.5 192.7 181.7 190.9 185.9
[28] 189.9 187.9 184.5 185.9 195.9 187.9 189.9 206.9 192.9
[37] 188.9 191.9 194.7

Pen & paper exercise

By hand:

  1. Decide on a sensible stem (tens? hundreds?).
  2. Write the stems in a column.
  3. Add a leaf for each price, then sort the leaves within each stem.
  4. What shape do you see? Is it symmetric, or does it lean one way?

And, in R …

Show code
stem(fuel_main$price)

  The decimal point is 1 digit(s) to the right of the |

  17 | 
  18 | 0122233344
  18 | 5556666778889
  19 | 00000111233
  19 | 556
  20 | 0
  20 | 7

Compare this to your hand-drawn version. Does the shape match what you expected?

Refining the scale

The stem doesn’t have to be a single digit. Splitting or stretching stems can reveal more, or less, structure.

Show code
stem(fuel_main$price, scale = 2)

  The decimal point is at the |

  178 | 9
  180 | 7799
  182 | 77979
  184 | 5597999
  186 | 99799
  188 | 979999
  190 | 9999
  192 | 79
  194 | 799
  196 | 
  198 | 9
  200 | 
  202 | 
  204 | 
  206 | 9

What changed compared to the default?

Summary: stem-and-leaf

  • Stem-and-leaf plots carry similar information to a histogram, but keep the raw digits.
  • Because the numbers are still there, it’s easy to then read off the median, Q1, or Q3 by hand.
  • It’s great for small data sets, when you only have pencil and paper.
  • Alternatives: histogram, (jittered) dot plot, density plot, box plot, violin plot, letter-value plot.

A different style of number-scratching

for categorical variables

We know about

but it’s too easy to

make a mistake

Is this easier?

or harder?

Your turn: getting to campus

24 postgrad students were asked, in a 2026 survey, how they usually get to campus:

Illustrative values, constructed for this exercise – not from an actual survey.

Show code
options(width = 60)
transport$mode
 [1] "Train" "Car"   "Bike"  "Train" "Walk"  "Car"   "Train"
 [8] "Bus"   "Bike"  "Train" "Car"   "Walk"  "Train" "Bike" 
[15] "Car"   "Train" "Bus"   "Walk"  "Car"   "Train" "Bike" 
[22] "Train" "Car"   "Walk" 

Pen & paper exercise

By hand, tally these into counts for each mode of transport – try the five-bar-gate tally first, then try the squares approach. Which method did you find less error-prone?

Actually, wait. This will be a competition, one group using five-bar-gate tally, and the other group using the squares.

And, in R …

Show code
transport |> count(mode)
# A tibble: 5 × 2
  mode      n
  <chr> <int>
1 Bike      4
2 Bus       2
3 Car       6
4 Train     8
5 Walk      4

What does it mean to “feel what the data are like?”

This is a stem and leaf of the height of the highest peak in each of the 50 US states.

Data: US Geological Survey elevation records, as tabulated in Tukey (1977).


The states roughly fall into three groups.


It’s not really surprising, but we can imagine this grouping. Alaska is in a group of its own, with a much higher high peak. Then the Rocky Mountain states, California, Washington and Hawaii also have high peaks, and the rest of the states lump together.

Exploratory data analysis is detective work – in the purest sense – finding and revealing the clues.

More summaries of numerical values

Hinges and 5-number summaries

You know the median is the middle number. What’s a hinge?

  • Find the median: the middle value.
  • Split the data at the median into a lower half and an upper half (include the median itself in both halves if \(n\) is odd).
  • The hinge is the median of each half.

Hinges are almost always close to Q1 and Q3, but are simpler to find by hand.

Our 39 petrol prices, sorted:

Show code
options(width = 25)
print(sort(fuel_main$price), digits = 4)
 [1] 179.9 180.7 181.7
 [4] 181.9 181.9 182.7
 [7] 182.7 182.9 183.7
[10] 183.9 184.5 184.5
[13] 184.9 185.7 185.9
[16] 185.9 185.9 186.9
[19] 186.9 187.7 187.9
[22] 187.9 188.9 189.7
[25] 189.9 189.9 189.9
[28] 189.9 190.9 190.9
[31] 190.9 191.9 192.7
[34] 192.9 194.7 194.9
[37] 195.9 199.9 206.9

With 39 values: the median is the ??, and each hinge is the middle of the remaining ?? (the ?? from each end).

Your turn: hinges and the 5-number summary

Pen & paper exercise

Using the stem-and-leaf:


  The decimal point is at the |

  178 | 9
  180 | 7799
  182 | 77979
  184 | 5597999
  186 | 99799
  188 | 979999
  190 | 9999
  192 | 79
  194 | 799
  196 | 
  198 | 9
  200 | 
  202 | 
  204 | 
  206 | 9
  1. Find the median.
  2. Find the lower and upper hinge.
  3. Find the minimum and maximum.
  4. Write out the full 5-number summary: minimum, lower hinge, median, upper hinge, maximum.

Check in R

Show code
print(fivenum(fuel_main$price), digits = 4)
[1] 179.9 184.2 187.7
[4] 190.9 206.9

fivenum() returns min, lower hinge, median, upper hinge, max – using the same depth rule you just used by hand.

Box-and-whisker display

Starting from your 5-number summary:

  • Draw a box from the lower hinge to the upper hinge.
  • Mark the median with a line inside the box.
  • Draw whiskers out to the min and max.

Pen & paper exercise

Using your 5-number summary for the petrol prices, draw the box-and-whisker plot by hand on a simple number line.

Show code
ggplot(fuel_main, aes(x = "", y = price)) +
  geom_boxplot() +
  xlab("") 

Your turn: two more stations turn up

Two more NSW stations show up in the data: Independent Goulburn (200.9c/L) and Thredbo Service Station (287.0c/L).

Goulburn is a regional town on the Hume Highway; Thredbo is a remote alpine resort village – both plausible reasons for pricier fuel.

Pen & paper exercise

  1. Using the H-spread from your hinges, compute the inner and outer fences.
  2. Where do 200.9 and 287.0 fall: inside, outside, or far out?
  3. Would either of the original values have been flagged, if you hadn’t already seen them as “normal”?
Show code
fuel |>
  ggplot(aes(x = "", y = price)) +
  geom_boxplot() +
  xlab("")

Points plotted beyond the whiskers are exactly the “outside” and “far out” values from your fence rule.

Isn’t this imposing a belief?

There is no excuse for failing to plot and look

Another Tukey wisdom drop

New statistics: trimeans

The number that comes closest to

\[\frac{\text{lower hinge} + 2\times \text{median} + \text{upper hinge}}{4}\] is the trimean – a measure of centre that leans on the median, but nudges toward the hinges too.


Think about trimmed means, where we might drop the highest and lowest 5% of observations, as a related idea.

Your turn: compute the trimean

Pen & paper exercise

Using your median and hinges from the petrol prices, compute the trimean by hand. How does it compare to the plain mean of the values?

Check in R

Show code
fn <- fivenum(fuel_main$price)
print((fn[2] + 2 * fn[3] + fn[4]) / 4, digits = 4)
[1] 187.6
Show code
print(mean(fuel_main$price), digits = 4)
[1] 188.1

Letter-value plots: today’s solution

Why break the data into quarters? Why not eighths, sixteenths? k-number summaries?

What does a 7-number summary look like?

How would you make an 11-number summary? (This one is easier to explore with more data than by hand – we’ll let R do the arithmetic.)

Petrol price (cents/litre), by fuel type, for 7919 NSW service station reports in July 2026 – still more than you’d want to sort by hand.

Real data: latest price per station and fuel type, via the NSW FuelCheck open data API.

Show code
fuel_by_type <- fuel_by_type |>
  mutate(fuel_type = factor(fuel_type, 
    levels = c("DL", "E10", "U91", "P95", "P98")))
ggplot(fuel_by_type, aes(fuel_type, price)) +
  geom_lv(aes(fill = after_stat(LV))) +
  scale_fill_brewer() +
  xlab("fuel type")

Your turn: choosing k

geom_lv() picks the number of letter values (k) automatically from the sample size, but you can also set it yourself.

Using R

  1. Re-run the plot with geom_lv(aes(fill = after_stat(LV)), k = 3), then again with k = 6 and k = 10.
  2. What changes as k gets small? As it gets larger?
  3. Is there a k that feels like “too few” letter values? Too many? Why?
Show code
ggplot(fuel_by_type, aes(fuel_type, price)) +
  geom_lv(aes(fill = after_stat(LV)), k = 3) +
  scale_fill_brewer() +
  xlab("fuel type")

Show code
ggplot(fuel_by_type, aes(fuel_type, price)) +
  geom_lv(aes(fill = after_stat(LV)), k = 6) +
  scale_fill_brewer() +
  xlab("fuel type")

Show code
ggplot(fuel_by_type, aes(fuel_type, price)) +
  geom_lv(aes(fill = after_stat(LV)), k = 10) +
  scale_fill_brewer() +
  xlab("fuel type")

Box plots are ubiquitous in use today.



- 🐈🐩 Mostly used to compare distributions, multiple subsets of the data.

  • Puts the emphasis on the middle 50% of observations, although variations can put emphasis on other aspects.

Easy re-expression

Logs, square roots, reciprocals

What you need to know about logs?

  • how to find good enough logs fast and easily
  • that equal differences in logs correspond to equal ratios of raw values.

(This means that wherever you find people using products or ratios– even in such things as price indexes–using logs–thus converting producers to sums and ratios to differences–is likely to help.)

The most common transformations are logs, sqrt root, reciprocals, reciprocals of square roots

-1, -1/2, +1/2, +1

What happened to ZERO?

It turns out that the role of a zero power, is for the purposes of re-expression, neatly solved by the logarithm.

Re-express to symmetrize the distribution

Your turn: fixing the skew

Your stem-and-leaf of the petrol prices likely showed a mild tail on the high side (the Edgecliff station stands apart from the rest).

Usinf R

  1. Take \(\log_{10}\) of each of the prices (a calculator is fine here).
  2. Sketch a rough stem-and-leaf of the logged values.
  3. Compare the shape to your original stem-and-leaf. Is it more symmetric?
Show code
stem(fuel_main$price, scale=2)

  The decimal point is at the |

  178 | 9
  180 | 7799
  182 | 77979
  184 | 5597999
  186 | 99799
  188 | 979999
  190 | 9999
  192 | 79
  194 | 799
  196 | 
  198 | 9
  200 | 
  202 | 
  204 | 
  206 | 9
Show code
stem(log10(fuel_main$price), scale=2)

  The decimal point is 2 digit(s) to the left of the |

  225 | 579
  226 | 002224
  226 | 56679999
  227 | 22344
  227 | 689999
  228 | 1113
  228 | 559
  229 | 02
  229 | 
  230 | 1
  230 | 
  231 | 
  231 | 6

Power ladder



⬅️ fix RIGHT-skewed values

-2, -1, -1/2, 0 (log), 1/3, 1/2, 1, 2, 3, 4


fix LEFT-skewed values ➡️

We now regard re-expression as a tool, something to let us do a better job of grasping. The grasping is done with the eye and the better job is through a more symmetric appearance.

Another Tukey wisdom drop

Linearising bivariate relationships


Surprising observation: The small fluctuations in later years.

What might be possible reasons?

Linearising bivariate relationships


See some fluctuations in the early years, too. Note that the log transformation couldn’t linearise.

Whatever the data, we can try to gain by straightening or by flattening.

When we succeed in doing one or both, we almost always see more clearly what is going on.

Rules and advice

  1. Graphics are friendly.
  2. Arithmetic often exists to make graphs possible.
  3. Graphs force us to notice the unexpected; nothing could be more important.
  4. Different graphs show us quite different aspects of the same data.
  5. There is no more reason to expect one graph to “tell all” than to expect one number to do the same.
  6. “Plotting \(y\) against \(x\)” involves significant choices–how we express one or both variables can be crucial.
  1. The first step in penetrating plotting is to straighten out the dependence or point scatter as much as reasonable.
  2. Plotting \(y^2\), \(\sqrt{y}\), \(log(y)\), \(-1/y\) or the like instead of \(y\) is one plausible step to take in search of straightness.
  3. Plotting \(x^2\), \(\sqrt{x}\), \(log(x)\), \(-1/x\) or the like instead of \(x\) is another.
  4. Once the plot is straightened, we can usually gain much by flattening it, usually by plotting residuals.
  5. When plotting scatters, we may need to be careful about how we express \(x\) and \(y\) in order to avoid concealment by crowding.

Summary: what we learned today

Scratching down numbers

  • Stem-and-leaf plots: keep every digit, still readable as a histogram shape, by hand.
  • Tallying: counting categorical data without losing track (five-bar gate, squares).

Schematic summary

  • Hinges and the 5-number summary: median + hinges, found by hand via depth.
  • Box-and-whisker plots: built directly from the 5-number summary.
  • Fences: H-spread-based rules for flagging “outside” and “far out” values.
  • Trimean: a robust measure of centre, weighting the median more than the hinges.
  • Letter-value plots: extend the idea to more letter values (k) for larger data sets.

Easy re-expression

  • Logs, square roots, and reciprocals can symmetrize a skewed distribution.
  • The power ladder gives a systematic way to choose a transformation, in either direction.
  • Straightening and flattening help us see relationships more clearly.

Underlying theme

  • Tukey’s tools are simple arithmetic and easy-to-draw pictures, designed to be done with pencil and paper.
  • The goal throughout is exploration: looking at data to see what it seems to say, before any formal confirmation.

The meandering of time

Each step below is a reaction to the one before it:

Hypothesis testing formalised

Tukey’s EDA Reaction to too much emphasis on hypothesis testing.

Testing methodology matures Reaction to the ad-hoc nature of data analysis.

Machine learning Reaction to statistical models missing non-linear relationships.

Renewed data exploration Reaction to the need to explain complex models.

Throughout, it has become much easier to accomplish thanks to computers. “Exploratory data analysis” as commonly used today is unfortunately synonymous with “descriptive statistics”, but it is truly much more – it is the tooling to discover what you don’t know.

Resources