Method card · Estimation

How many candies are in the jar?.

Without counting the candies in the jar, guess how many there are and type your answer.

Go to experiment2 min long12 real first answers

Original study

  1. Galton, F. (1907). Vox populi. Nature, 75(1949), 450–451. Source
  2. Galton, F. (1907). The ballot-box. Nature, 75(1952), 509–510. Source
  3. Lorenz, J., Rauhut, H., Schweitzer, F., & Helbing, D. (2011). How social influence can undermine the wisdom of crowd effect. Proceedings of the National Academy of Sciences, 108(22), 9020–9025. Source
  4. Kao, A. B., Berdahl, A. M., Hartnett, A. T., Lutz, M. J., Bak-Coleman, J. B., Ioannou, C. C., Giam, X., & Couzin, I. D. (2018). Counteracting estimation bias and social influence to improve the wisdom of crowds. Journal of the Royal Society Interface, 15(141), 20180130. Source
  5. Izard, V., & Dehaene, S. (2008). Calibrating the mental number line. Cognition, 106(3), 1221–1247. Source
  6. Treynor, J. L. (1987). Market efficiency and the bean jar experiment. Financial Analysts Journal, 43(3), 50–53. Source
  7. Becker, J., Brackbill, D., & Centola, D. (2017). Network dynamics of social influence in the wisdom of crowds. Proceedings of the National Academy of Sciences, 114(26), E5070–E5076. Source
  8. Yaniv, I., & Kleinberger, E. (2000). Advice taking in decision making: Egocentric discounting and reputation formation. Organizational Behavior and Human Decision Processes, 83(2), 260–281. Source
  9. Soll, J. B., & Larrick, R. P. (2009). Strategies for revising judgment: How (and how well) people use others' opinions. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(3), 780–805. Source

Our adaptation

An adaptation of Galton's (1907) crowd estimate and the social-influence design of Lorenz et al. (2011) with jars drawn on screen. Round 1 asks for an estimate of a jar with a fixed layout of 487 candies, round 2 a differently shaped jar with 1,240 candies; the input field shows no example number. Round 3 shows the median round-1 estimate of previous visitors (excluding the participant) and asks for a new estimate of jar 1. The true counts are revealed after round 3.

Differences from the original study

  • The jar is drawn on screen; there is no depth and screen size varies between people.
  • Social information is a single value (the median of previous estimates); Lorenz et al. used repeated rounds with different information conditions.
  • There is no reward.

Measured variables

The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.

ColumnTypeDescription
vintegerRound 1: estimate for jar 1 (true count 487; main measure).
v2numberRound 2: estimate for jar 2 (true count 1,240).
v1bnumberRound 3: new estimate for jar 1 after seeing the crowd median.
woanumberWeight of advice: (new estimate − first estimate) / (shown median − first estimate); 0 = no shift, 1 = moved fully to the median; 0 if the median equals the first estimate.
Columns present in every file (14)
row_idstringRow code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments.
experimentstringShort name of the experiment (URL slug).
experiment_versionintegerVersion of the answer format. For experiments whose format changed, only the current format is published.
dateYYYY-MM-DD | YYYY-Www | YYYY-MMDay the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few.
date_precisionday | week | monthPrecision of the date column.
langtr | en | (boş)Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.
devicedesktop | mobile | tablet | (boş)Device class reported by the browser (class only; browser details are neither stored here nor published).
sourcesite | embedWhere the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published.
modefree | daily | challenge | race | session | (boş)Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it.
color_visionnormal | rg | by | contrast | (boş)In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments.
first_play1Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data.
seen_beforeyes | no | unsure | (boş)Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask).
duration_snumberTime from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s).
play_scorenumberLeaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments.

Exclusion criteria

The server accepts the round-1 estimate as a whole number from 1 to 100,000. Values from the extra rounds are collected in the browser. Because estimates are skewed, we recommend a log scale and the median in analysis.

  • Only each participant's first answer to this experiment; replays are not included.
  • Simulated rows, bots, banned and sample accounts are not included.
  • Answers before 3 October 2026 are not included (the first day the open data notice was live).

Scoring

Leaderboard score: percentage error from the mean logarithmic error over both jars, exp(mean(|ln(estimate / truth)|)) − 1, × 100; only jar 1 if there is no jar-2 answer. Lower is better.

Leaderboard measure: estimation error (lower is better). The play_score column in the open data is this score for the first play.

Version notes

  • October 2026

    This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.

  • October 2026

    The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.