Name the color, not the word.
Ignore what the word on the screen says; look at the color it is written in and press that color's button as fast as you can.
Original study
- Stroop, J. R. (1935). Studies of interference in serial verbal reactions. Journal of Experimental Psychology, 18(6), 643–662. Source
- MacLeod, C. M. (1991). Half a century of research on the Stroop effect: An integrative review. Psychological Bulletin, 109(2), 163–203. Source
- Augustinova, M., Parris, B. A., & Ferrand, L. (2019). The loci of Stroop interference and facilitation effects with manual and vocal responses. Frontiers in Psychology, 10, 1786. Source
- Crump, M. J. C., McDonnell, J. V., & Gureckis, T. M. (2013). Evaluating Amazon's Mechanical Turk as a tool for experimental behavioral research. PLoS ONE, 8(3), e57410. Source
- Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186. Source
- Durgin, F. H. (2000). The reverse Stroop effect. Psychonomic Bulletin & Review, 7(1), 121–125. Source
- Heitz, R. P. (2014). The speed-accuracy tradeoff: History, physiology, methodology, and behavior. Frontiers in Neuroscience, 8, 150. Source
Our adaptation
A touch and mouse adaptation of Stroop's (1935) color-word task. In round 1, over 16 trials a color word (RED, BLUE, GREEN, YELLOW) is shown in an ink color; half the trials are congruent and half incongruent, and the answer is given by tapping one of four color swatches in fixed positions. In round 2 (12 trials) the task is reversed: the word is read and word buttons printed in neutral ink are pressed. Round 3 (12 trials) returns to the classic task with a 1,200 ms limit per trial.
Differences from the original study
- Answers are given by tapping color swatches rather than aloud; this setup can make the Stroop difference smaller than in vocal-response studies.
- There are 8 trials per condition; this is noisy for individual differences and should be interpreted at the crowd level.
- Words are in Turkish or English depending on the interface language.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
cong | integer | Round 1: median response time of correct congruent trials (ms). |
incong | integer | Round 1: median response time of correct incongruent trials (ms). |
errors | integer | Round 1: number of wrong answers (out of 16 trials). |
interference | number | Round 1 Stroop effect: incong − cong (ms). |
rev_diff | number | Round 2 (read the word): incongruent minus congruent median time (ms). |
rev_err | integer | Round 2: number of wrong answers. |
press_rt | number | Round 3 (1,200 ms limit): median time of correct incongruent trials (ms). |
press_to | integer | Round 3: number of trials missed because time ran out. |
press_err | integer | Round 3: number of wrong answers given before time ran out. |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
Only correct answers at or above 200 ms enter the medians (faster responses count as anticipations). The server accepts round-1 medians between 200 and 5,000 ms and an error count between 0 and 64; anything else is rejected.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: the round-1 incongruent median time (ms); lower is better. Participants with more than 3 errors in 16 trials (accuracy below 80%) are not ranked; their answers remain in the data.
Leaderboard measure: incongruent median (lower is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
Plausibility bounds were tightened: anticipatory responses (below 200 ms for Stroop and rotation, below 100 ms for reaction time) are left out of the median on the client, and the server rejects shorter medians (the previous floor was 150 ms, 80 ms for reaction time).
- October 2026
This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.