Method card · Randomness

Heads or tails: can you be random?.

Can you behave randomly, predict chance, and tell a sequence made by a human apart from a real coin?

Go to experiment3 min long6 real first answers

Original study

  1. Wagenaar, W. A. (1972). Generation of random sequences by human subjects: A critical survey of literature. Psychological Bulletin, 77(1), 65–72. Source
  2. Falk, R., & Konold, C. (1997). Making sense of randomness: Implicit encoding as a basis for judgment. Psychological Review, 104(2), 301–318. Source
  3. Nickerson, R. S. (2002). The production and perception of randomness. Psychological Review, 109(2), 330–357. Source
  4. Kahneman, D., & Tversky, A. (1972). Subjective probability: A judgment of representativeness. Cognitive Psychology, 3(3), 430–454. Source
  5. Gauvrit, N., Zenil, H., Soler-Toscano, F., Delahaye, J.-P., & Brugger, P. (2017). Human behavioral complexity peaks at age 25. PLOS Computational Biology, 13(4), e1005408. Source
  6. Neuringer, A. (1986). Can people behave “randomly?”: The role of feedback. Journal of Experimental Psychology: General, 115(1), 62–75. Source
  7. Croson, R., & Sundali, J. (2005). The gambler's fallacy and the hot hand: Empirical data from casinos. Journal of Risk and Uncertainty, 30(3), 195–209. Source
  8. Ayton, P., & Fischer, I. (2004). The hot hand fallacy and the gambler's fallacy: Two faces of subjective randomness? Memory & Cognition, 32(8), 1369–1378. Source
  9. Miller, J. B., & Sanjurjo, A. (2018). Surprised by the hot hand fallacy? A truth in the law of small numbers. Econometrica, 86(6), 2019–2047. Source

Our adaptation

A three-round adaptation of studies on how people produce (Wagenaar, 1972) and perceive (Falk and Konold, 1997) random sequences. In round 1 a sequence is produced by tapping Heads or Tails 20 times; the sequence does not accumulate on screen, only a counter is shown. In round 2 the participant predicts each of 20 real tosses made with the browser's cryptographic random number generator. In round 3 two sequences are shown three times (one a truly random computer sequence, the other a previous visitor's round-1 sequence) and the participant picks the human-made one.

Differences from the original study

  • 20 tosses are short for judging one person; even for a fair coin the alternation rate usually ranges from 0.32 to 0.68.
  • The instruction is to 'produce a random sequence'; as Nickerson (2002) stresses, results are sensitive to instructions.
  • The sequence is not shown on screen, so that a visible sequence does not invite balancing.

Measured variables

The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.

ColumnTypeDescription
seqstringRound 1: produced sequence (Y = heads, T = tails, from the Turkish yazı / tura; 20 characters).
altnumberRound 1 alternation rate: share of the 19 transitions where consecutive tosses differ (0.5 expected for a fair coin).
longestintegerRound 1: longest run of identical outcomes (≈ 4.66 expected for a fair coin).
pred_accnumberRound 2: share of correct predictions over 20 real tosses.
gamblernumberRound 2: share of predictions of the opposite outcome right after three identical outcomes in a row (gambler's fallacy); empty if no such moment occurred.
streaksintegerRound 2: number of moments after three identical outcomes in a row (the denominator of gambler).
spot_1numberRound 3: whether the human-made sequence was identified in each pair (1 = correct, 0 = wrong). (item 1)
spot_2numberRound 3: whether the human-made sequence was identified in each pair (1 = correct, 0 = wrong). (item 2)
spot_3numberRound 3: whether the human-made sequence was identified in each pair (1 = correct, 0 = wrong). (item 3)
pred_outcomesstringThe 20 truly random tosses of round 2 (Y/T).
Columns present in every file (14)
row_idstringRow code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments.
experimentstringShort name of the experiment (URL slug).
experiment_versionintegerVersion of the answer format. For experiments whose format changed, only the current format is published.
dateYYYY-MM-DD | YYYY-Www | YYYY-MMDay the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few.
date_precisionday | week | monthPrecision of the date column.
langtr | en | (boş)Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.
devicedesktop | mobile | tablet | (boş)Device class reported by the browser (class only; browser details are neither stored here nor published).
sourcesite | embedWhere the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published.
modefree | daily | challenge | race | session | (boş)Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it.
color_visionnormal | rg | by | contrast | (boş)In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments.
first_play1Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data.
seen_beforeyes | no | unsure | (boş)Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask).
duration_snumberTime from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s).
play_scorenumberLeaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments.

Exclusion criteria

The server checks that the round-1 sequence consists of exactly 20 Y/T characters and computes the alternation rate and the longest run itself. Round 2 and 3 values are computed in the browser.

  • Only each participant's first answer to this experiment; replays are not included.
  • Simulated rows, bots, banned and sample accounts are not included.
  • Answers before 3 October 2026 are not included (the first day the open data notice was live).

Scoring

Leaderboard score: the round-1 randomness score plus 10 points for each correct identification in round 3. Randomness score = max(0, 100 − |alternation rate − 0.5| × 250 − |longest run − 4.66| × 8). Higher is better.

Leaderboard measure: randomness score (higher is better). The play_score column in the open data is this score for the first play.

Version notes

  • October 2026

    This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.

  • October 2026

    The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.