How sure are you?.
Ten numerical questions: for each one, can you give a range that has a 90 percent chance of containing the right answer?
Original study
- Alpert, M., & Raiffa, H. (1982). A progress report on the training of probability assessors. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under uncertainty: Heuristics and biases (pp. 294–305). Cambridge University Press. Source
- Lichtenstein, S., Fischhoff, B., & Phillips, L. D. (1982). Calibration of probabilities: The state of the art to 1980. In D. Kahneman, P. Slovic, & A. Tversky (Eds.), Judgment under uncertainty: Heuristics and biases (pp. 306–334). Cambridge University Press. Source
- Soll, J. B., & Klayman, J. (2004). Overconfidence in interval estimates. Journal of Experimental Psychology: Learning, Memory, and Cognition, 30(2), 299–314. Source
- Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517. Source
- Haran, U., Moore, D. A., & Morewedge, C. K. (2010). A simple remedy for overprecision in judgment. Judgment and Decision Making, 5(7), 467–476. Source
- Moore, D. A., Tenney, E. R., & Haran, U. (2015). Overprecision in judgment. In G. Keren & G. Wu (Eds.), The Wiley Blackwell handbook of judgment and decision making (pp. 182–209). Wiley. Source
Our adaptation
A ten-question adaptation of Alpert and Raiffa's (1982) 90% confidence interval task. For each question a lower and an upper bound are entered such that the participant is 90% sure the true answer lies between them. The questions are split into two groups; at the end of the first group only the number of intervals containing the answer is shown, and the true answers with their sources are revealed at the end. The true answers are the values in the sources cited next to each question.
Differences from the original study
- Ten questions are few: chance can easily move the hit rate by 10–20 points.
- Answers can be looked up online; there is no time limit.
- Some questions concern Turkey and may be more familiar to some participants.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
intervals | JSON | Intervals for the ten questions: [[lower, upper] × 10] (swapped on the server if lower > upper). |
hits | integer | Number of intervals containing the true answer (0–10; about 9 expected for a well-calibrated person). |
hit_1 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 1) |
hit_2 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 2) |
hit_3 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 3) |
hit_4 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 4) |
hit_5 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 5) |
hit_6 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 6) |
hit_7 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 7) |
hit_8 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 8) |
hit_9 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 9) |
hit_10 | number | Whether the interval contains the true answer for each question (1 = yes); the number at the end of the column is the question order. (item 10) |
interval_score | number | Mean interval score (Gneiting and Raftery, 2007): (width + 20 × miss distance) / |true answer|; lower is better. |
points | integer | Calibration score: round(100 / (1 + interval_score)), 0–100. |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server checks that both bounds of each of the ten questions lie within that question's allowed range; incomplete submissions are rejected. Hits and scores are computed on the server.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: the calibration score (0–100); intervals that are both narrow and correct are rewarded. Higher is better.
Leaderboard measure: calibration (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.