The bat and the ball.
Answer seven short questions: will you write down the first thing that comes to mind, or think it over once more?
Original study
- Frederick, S. (2005). Cognitive reflection and decision making. Journal of Economic Perspectives, 19(4), 25–42. Source
- Toplak, M. E., West, R. F., & Stanovich, K. E. (2011). The Cognitive Reflection Test as a predictor of performance on heuristics-and-biases tasks. Memory & Cognition, 39(7), 1275–1289. Source
- Brañas-Garza, P., Kujal, P., & Lenkei, B. (2019). Cognitive reflection test: Whom, how, when. Journal of Behavioral and Experimental Economics, 82, 101455. Source
- Haigh, M. (2016). Has the standard Cognitive Reflection Test become a victim of its own success? Advances in Cognitive Psychology, 12(3), 145–149. Source
- Bialek, M., & Pennycook, G. (2018). The cognitive reflection test is robust to multiple exposures. Behavior Research Methods, 50(5), 1953–1959. Source
- Stagnaro, M. N., Pennycook, G., & Rand, D. G. (2018). Performance on the Cognitive Reflection Test is stable across time. Judgment and Decision Making, 13(3), 260–267. Source
- Thomson, K. S., & Oppenheimer, D. M. (2016). Investigating an alternate form of the cognitive reflection test. Judgment and Decision Making, 11(1), 99–113. Source
Our adaptation
An adaptation of Frederick's (2005) Cognitive Reflection Test (CRT) using Turkish lira and metric units: round 1 asks the three standard items in the standard order (bat and ball, machines, lily pads) with numeric input. Rounds 2 and 3 ask four items of Thomson and Oppenheimer's (2016) CRT-2, adapted to Turkish and the metric system. Before answering, participants are asked whether they had seen the questions before.
Differences from the original study
- The currency is TL (the 110 = 100 + 10 split is kept).
- The CRT-2 adaptations have not been tested in a separate validity study.
- Time and setting are not controlled; the items are very well known.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
a_1 | number | Round 1: the answer to each of the three CRT items read as a number (empty if it could not be read). (item 1) |
a_2 | number | Round 1: the answer to each of the three CRT items read as a number (empty if it could not be read). (item 2) |
a_3 | number | Round 1: the answer to each of the three CRT items read as a number (empty if it could not be read). (item 3) |
class_1 | correct | intuitive | other | Class of the bat and ball answer: correct (5) | intuitive (10) | other. |
class_2 | correct | intuitive | other | Class of the machines answer: correct (5) | intuitive (100) | other. |
class_3 | correct | intuitive | other | Class of the lily pads answer: correct (47) | intuitive (24) | other. |
score | integer | Round 1: number of correct answers (0–3). |
crt2_ok_1 | number | Rounds 2 and 3: whether each of the four CRT-2 items was answered correctly (1 = correct, 0 = wrong; empty if unanswered). (item 1) |
crt2_ok_2 | number | Rounds 2 and 3: whether each of the four CRT-2 items was answered correctly (1 = correct, 0 = wrong; empty if unanswered). (item 2) |
crt2_ok_3 | number | Rounds 2 and 3: whether each of the four CRT-2 items was answered correctly (1 = correct, 0 = wrong; empty if unanswered). (item 3) |
crt2_ok_4 | number | Rounds 2 and 3: whether each of the four CRT-2 items was answered correctly (1 = correct, 0 = wrong; empty if unanswered). (item 4) |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The answers themselves (free text) are not published; only the number read on the server and its class are. The server reads a single number from an answer, ignoring units and words ('5 TL', '5.00', 'five'). We recommend reporting participants who say they had seen the questions before (seen_before) separately.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: total correct over seven items (CRT 0–3 + CRT-2 0–4); higher is better. CRT-2 answers are re-scored on the server.
Leaderboard measure: correct (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
The answer parser was extended: a single number inside an answer with units or words (“8 sheep”, “five”) is now read. In earlier answers such entries may not have been read as numbers and fell into the “other” class.
- October 2026
This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.