A crowd at a glance
If you glance at the marbles in a jar and close your eyes, you can't say where each marble was. But you can usually say roughly how big the marbles were on average, and whether they were all about the same size or mixed. The visual system summarises crowds with a kind of statistics.
In this experiment you look at twelve circles or twelve short lines for only half a second. Then you set the average size or the average tilt. Twice there is a surprise question: which of two circles was in the set? One really was in the set; the other never was, but its size was exactly the set's average.
Ariely's spots
In a study published in Psychological Science in 2001, Dan Ariely showed two observers sets of 4, 8, 12 or 16 spots for half a second. In one task he asked whether a single spot that followed was larger or smaller than the set's average; in another, whether a spot had been in the set at all.
The result was striking. The observers could discriminate the average very finely: the threshold was about 4–6 per cent of the spot size for similar spots and 6–12 per cent for spots of clearly different sizes. But they could tell whether a spot had been in the set only at near-chance level, even though members and non-members differed in size by at least 18 per cent. They knew the set's average, but not its members.
Short display, solid average
In 2003 Chong and Treisman tested this finding with sets of twelve circles of different sizes. The threshold for judging the set's average was close to the thresholds measured for sets of identical circles and for a single circle. Cutting the display time to 50 milliseconds, or delaying the answer by 2 seconds, changed the result very little.
The twelve circles and the half-second display in this experiment are inspired by these studies. The circle sizes come in four steps, spread symmetrically around the average, so the average is a size that is not in the set at all. The closer your circle is to that average, the higher your score.
Average orientation and crowding
Averages are not just about size. In 2001 Parkes, Lund, Angelucci, Solomon and Morgan showed oriented patterns packed closely together in peripheral vision in Nature Neuroscience. The patterns were so crowded that observers could not report the orientation of the central one; yet they could reliably report the average orientation of all of them. Information that could not be read individually was not lost; it went into the average.
In two rounds of this experiment we test the same idea with tilt: the orientations of twelve short lines are spread evenly on both sides of an average tilt, and you set that average. Ensemble perception also works for faces: in 2007 Haberman and Whitney found that people rapidly extract the average emotion and gender of a set of faces.
Recognising a circle you never saw
If the visual system keeps only the average, an interesting error is expected: a circle that was never in the set but is exactly the average size should feel more familiar than a real member. In 2009 de Fockert and Wolfenstein showed this with faces: participants who saw the faces of four different people said 'it was in the set' more often to an average face morphed from those four faces than to the real members.
The surprise question in this experiment follows the same logic. One of the two circles is a real member; the other was not in the set but is exactly the average size. If you guess, you will pick the real member half the time; a memory that drifts toward the average pulls you toward the average-sized circle. One person's two answers say little; we will look at which circle the crowd picks more often.
A summary, or a few samples?
How to interpret these findings is debated. One view is that the visual system processes all the items in a set in parallel and extracts a special summary. Myczek and Simons challenged this in 2008: with experiments and simulations they showed that focusing on a few items and averaging only those could explain most findings on average size. Ariely replied the same year and the debate continued; it is still not settled.
In their 2018 review, Whitney and Yamanashi Leib argued that ensemble perception works at many levels, from colour and size to facial expressions, and that it may be a way for a limited memory to summarise the world. Whichever route it takes, the practical result is the same: when you glance at a crowd, what stays with you is usually not the individual faces but the crowd itself.