How many candies are in the jar?.
Without counting the candies in the jar, guess how many there are and type your answer.
Original study
- Galton, F. (1907). Vox populi. Nature, 75(1949), 450–451. Source
- Galton, F. (1907). The ballot-box. Nature, 75(1952), 509–510. Source
- Lorenz, J., Rauhut, H., Schweitzer, F., & Helbing, D. (2011). How social influence can undermine the wisdom of crowd effect. Proceedings of the National Academy of Sciences, 108(22), 9020–9025. Source
- Kao, A. B., Berdahl, A. M., Hartnett, A. T., Lutz, M. J., Bak-Coleman, J. B., Ioannou, C. C., Giam, X., & Couzin, I. D. (2018). Counteracting estimation bias and social influence to improve the wisdom of crowds. Journal of the Royal Society Interface, 15(141), 20180130. Source
- Izard, V., & Dehaene, S. (2008). Calibrating the mental number line. Cognition, 106(3), 1221–1247. Source
- Treynor, J. L. (1987). Market efficiency and the bean jar experiment. Financial Analysts Journal, 43(3), 50–53. Source
- Becker, J., Brackbill, D., & Centola, D. (2017). Network dynamics of social influence in the wisdom of crowds. Proceedings of the National Academy of Sciences, 114(26), E5070–E5076. Source
- Yaniv, I., & Kleinberger, E. (2000). Advice taking in decision making: Egocentric discounting and reputation formation. Organizational Behavior and Human Decision Processes, 83(2), 260–281. Source
- Soll, J. B., & Larrick, R. P. (2009). Strategies for revising judgment: How (and how well) people use others' opinions. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(3), 780–805. Source
Our adaptation
An adaptation of Galton's (1907) crowd estimate and the social-influence design of Lorenz et al. (2011) with jars drawn on screen. Round 1 asks for an estimate of a jar with a fixed layout of 487 candies, round 2 a differently shaped jar with 1,240 candies; the input field shows no example number. Round 3 shows the median round-1 estimate of previous visitors (excluding the participant) and asks for a new estimate of jar 1. The true counts are revealed after round 3.
Differences from the original study
- The jar is drawn on screen; there is no depth and screen size varies between people.
- Social information is a single value (the median of previous estimates); Lorenz et al. used repeated rounds with different information conditions.
- There is no reward.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
v | integer | Round 1: estimate for jar 1 (true count 487; main measure). |
v2 | number | Round 2: estimate for jar 2 (true count 1,240). |
v1b | number | Round 3: new estimate for jar 1 after seeing the crowd median. |
woa | number | Weight of advice: (new estimate − first estimate) / (shown median − first estimate); 0 = no shift, 1 = moved fully to the median; 0 if the median equals the first estimate. |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server accepts the round-1 estimate as a whole number from 1 to 100,000. Values from the extra rounds are collected in the browser. Because estimates are skewed, we recommend a log scale and the median in analysis.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: percentage error from the mean logarithmic error over both jars, exp(mean(|ln(estimate / truth)|)) − 1, × 100; only jar 1 if there is no jar-2 answer. Lower is better.
Leaderboard measure: estimation error (lower is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.