Heads or tails: can you be random?.
Can you behave randomly, predict chance, and tell a sequence made by a human apart from a real coin?
Original study
- Wagenaar, W. A. (1972). Generation of random sequences by human subjects: A critical survey of literature. Psychological Bulletin, 77(1), 65–72. Source
- Falk, R., & Konold, C. (1997). Making sense of randomness: Implicit encoding as a basis for judgment. Psychological Review, 104(2), 301–318. Source
- Nickerson, R. S. (2002). The production and perception of randomness. Psychological Review, 109(2), 330–357. Source
- Kahneman, D., & Tversky, A. (1972). Subjective probability: A judgment of representativeness. Cognitive Psychology, 3(3), 430–454. Source
- Gauvrit, N., Zenil, H., Soler-Toscano, F., Delahaye, J.-P., & Brugger, P. (2017). Human behavioral complexity peaks at age 25. PLOS Computational Biology, 13(4), e1005408. Source
- Neuringer, A. (1986). Can people behave “randomly?”: The role of feedback. Journal of Experimental Psychology: General, 115(1), 62–75. Source
- Croson, R., & Sundali, J. (2005). The gambler's fallacy and the hot hand: Empirical data from casinos. Journal of Risk and Uncertainty, 30(3), 195–209. Source
- Ayton, P., & Fischer, I. (2004). The hot hand fallacy and the gambler's fallacy: Two faces of subjective randomness? Memory & Cognition, 32(8), 1369–1378. Source
- Miller, J. B., & Sanjurjo, A. (2018). Surprised by the hot hand fallacy? A truth in the law of small numbers. Econometrica, 86(6), 2019–2047. Source
Our adaptation
A three-round adaptation of studies on how people produce (Wagenaar, 1972) and perceive (Falk and Konold, 1997) random sequences. In round 1 a sequence is produced by tapping Heads or Tails 20 times; the sequence does not accumulate on screen, only a counter is shown. In round 2 the participant predicts each of 20 real tosses made with the browser's cryptographic random number generator. In round 3 two sequences are shown three times (one a truly random computer sequence, the other a previous visitor's round-1 sequence) and the participant picks the human-made one.
Differences from the original study
- 20 tosses are short for judging one person; even for a fair coin the alternation rate usually ranges from 0.32 to 0.68.
- The instruction is to 'produce a random sequence'; as Nickerson (2002) stresses, results are sensitive to instructions.
- The sequence is not shown on screen, so that a visible sequence does not invite balancing.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
seq | string | Round 1: produced sequence (Y = heads, T = tails, from the Turkish yazı / tura; 20 characters). |
alt | number | Round 1 alternation rate: share of the 19 transitions where consecutive tosses differ (0.5 expected for a fair coin). |
longest | integer | Round 1: longest run of identical outcomes (≈ 4.66 expected for a fair coin). |
pred_acc | number | Round 2: share of correct predictions over 20 real tosses. |
gambler | number | Round 2: share of predictions of the opposite outcome right after three identical outcomes in a row (gambler's fallacy); empty if no such moment occurred. |
streaks | integer | Round 2: number of moments after three identical outcomes in a row (the denominator of gambler). |
spot_1 | number | Round 3: whether the human-made sequence was identified in each pair (1 = correct, 0 = wrong). (item 1) |
spot_2 | number | Round 3: whether the human-made sequence was identified in each pair (1 = correct, 0 = wrong). (item 2) |
spot_3 | number | Round 3: whether the human-made sequence was identified in each pair (1 = correct, 0 = wrong). (item 3) |
pred_outcomes | string | The 20 truly random tosses of round 2 (Y/T). |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server checks that the round-1 sequence consists of exactly 20 Y/T characters and computes the alternation rate and the longest run itself. Round 2 and 3 values are computed in the browser.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: the round-1 randomness score plus 10 points for each correct identification in round 3. Randomness score = max(0, 100 − |alternation rate − 0.5| × 250 − |longest run − 4.66| × 8). Higher is better.
Leaderboard measure: randomness score (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.