Find the odd shade.
All the squares look the same color, but one is ever so slightly different: how small a difference can your eye catch?
Original study
- MacAdam, D. L. (1942). Visual sensitivities to color differences in daylight. Journal of the Optical Society of America, 32(5), 247–274. Source
- Mahy, M., Van Eycken, L., & Oosterlinck, A. (1994). Evaluation of uniform color spaces developed after the adoption of CIELAB and CIELUV. Color Research & Application, 19(2), 105–121. See also: Sharma, G., & Trussell, H. J. (1997). Digital color imaging. IEEE Transactions on Image Processing, 6(7), 901–932. https://doi.org/10.1109/83.597268 Source
- Levitt, H. (1971). Transformed up-down methods in psychoacoustics. The Journal of the Acoustical Society of America, 49(2B), 467–477. Source
- Farnsworth, D. (1943). The Farnsworth-Munsell 100-hue and dichotomous tests for color vision. Journal of the Optical Society of America, 33(10), 568–578. Source
- Birch, J. (2012). Worldwide prevalence of red-green color deficiency. Journal of the Optical Society of America A, 29(3), 313–320. Source
Our adaptation
Color discrimination thresholds measured with an adaptive staircase (Levitt, 1971), followed by a level game. Colors are generated in CIELAB (L* = 62, C* = 38, random hue) and differences are expressed as ΔE*ab (CIE76). Round 1 varies hue only and round 2 lightness only: the difference starts at ΔE 14, is multiplied by 0.75 after two correct answers in a row and by 1.33 after each error; a round ends after 8 reversals or at most 26 screens, and the threshold is the geometric mean of the last 6 reversals (about 70.7% correct). Round 3 is a level game: the grid grows from 2×2 to 7×7, the difference shrinks with every level, and there are three lives and a time limit.
Differences from the original study
- Screens are not calibrated; the same color code looks different across devices.
- The measure is a threshold estimate without the many trials and fixed viewing conditions of the lab.
- Plays in the selected color vision mode are evaluated in a separate bucket.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
th_1 | number | Discrimination threshold (ΔE*ab): th_1 = round 1 (hue), th_2 = round 2 (lightness). (item 1) |
th_2 | number | Discrimination threshold (ΔE*ab): th_1 = round 1 (hue), th_2 = round 2 (lightness). (item 2) |
acc_1 | number | Accuracy in the staircase rounds: acc_1 = hue, acc_2 = lightness. (item 1) |
acc_2 | number | Accuracy in the staircase rounds: acc_1 = hue, acc_2 = lightness. (item 2) |
level | integer | Number of levels passed in round 3. |
level_de | number | Color difference (ΔE*ab) at the level reached in round 3. |
tool_shuffle | integer | Times the 'Shuffle' helper was used in the level game. |
tool_rest | integer | Times the 'Rest' helper was used in the level game. |
end | hearts | time | max | Why the level game ended (hearts = out of lives, time = time ran out, max = last level). |
trials | JSON | Screens: [[round, ΔE, correct 1/0, time in ms], …]. |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
Only answers in the current format (format 2) are published. The server requires both thresholds between 0.05 and 80 and the level between 0 and the maximum; plays at a humanly impossible speed are not ranked.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
- Answers given in the old task format (format 1) are not included; only the current format is published.
Scoring
Leaderboard score: number of levels passed in round 3; higher is better.
Leaderboard measure: level reached (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
Switched to the arcade format (format 2): the task and the score scale changed, and older plays were removed from the leaderboard. The open data contains format 2 answers only.
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.