Method card · Color perception

Find the odd shade.

All the squares look the same color, but one is ever so slightly different: how small a difference can your eye catch?

Go to experiment3 min long0 real first answers

Original study

  1. MacAdam, D. L. (1942). Visual sensitivities to color differences in daylight. Journal of the Optical Society of America, 32(5), 247–274. Source
  2. Mahy, M., Van Eycken, L., & Oosterlinck, A. (1994). Evaluation of uniform color spaces developed after the adoption of CIELAB and CIELUV. Color Research & Application, 19(2), 105–121. See also: Sharma, G., & Trussell, H. J. (1997). Digital color imaging. IEEE Transactions on Image Processing, 6(7), 901–932. https://doi.org/10.1109/83.597268 Source
  3. Levitt, H. (1971). Transformed up-down methods in psychoacoustics. The Journal of the Acoustical Society of America, 49(2B), 467–477. Source
  4. Farnsworth, D. (1943). The Farnsworth-Munsell 100-hue and dichotomous tests for color vision. Journal of the Optical Society of America, 33(10), 568–578. Source
  5. Birch, J. (2012). Worldwide prevalence of red-green color deficiency. Journal of the Optical Society of America A, 29(3), 313–320. Source

Our adaptation

Color discrimination thresholds measured with an adaptive staircase (Levitt, 1971), followed by a level game. Colors are generated in CIELAB (L* = 62, C* = 38, random hue) and differences are expressed as ΔE*ab (CIE76). Round 1 varies hue only and round 2 lightness only: the difference starts at ΔE 14, is multiplied by 0.75 after two correct answers in a row and by 1.33 after each error; a round ends after 8 reversals or at most 26 screens, and the threshold is the geometric mean of the last 6 reversals (about 70.7% correct). Round 3 is a level game: the grid grows from 2×2 to 7×7, the difference shrinks with every level, and there are three lives and a time limit.

Differences from the original study

  • Screens are not calibrated; the same color code looks different across devices.
  • The measure is a threshold estimate without the many trials and fixed viewing conditions of the lab.
  • Plays in the selected color vision mode are evaluated in a separate bucket.

Measured variables

The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.

ColumnTypeDescription
th_1numberDiscrimination threshold (ΔE*ab): th_1 = round 1 (hue), th_2 = round 2 (lightness). (item 1)
th_2numberDiscrimination threshold (ΔE*ab): th_1 = round 1 (hue), th_2 = round 2 (lightness). (item 2)
acc_1numberAccuracy in the staircase rounds: acc_1 = hue, acc_2 = lightness. (item 1)
acc_2numberAccuracy in the staircase rounds: acc_1 = hue, acc_2 = lightness. (item 2)
levelintegerNumber of levels passed in round 3.
level_denumberColor difference (ΔE*ab) at the level reached in round 3.
tool_shuffleintegerTimes the 'Shuffle' helper was used in the level game.
tool_restintegerTimes the 'Rest' helper was used in the level game.
endhearts | time | maxWhy the level game ended (hearts = out of lives, time = time ran out, max = last level).
trialsJSONScreens: [[round, ΔE, correct 1/0, time in ms], …].
Columns present in every file (14)
row_idstringRow code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments.
experimentstringShort name of the experiment (URL slug).
experiment_versionintegerVersion of the answer format. For experiments whose format changed, only the current format is published.
dateYYYY-MM-DD | YYYY-Www | YYYY-MMDay the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few.
date_precisionday | week | monthPrecision of the date column.
langtr | en | (boş)Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.
devicedesktop | mobile | tablet | (boş)Device class reported by the browser (class only; browser details are neither stored here nor published).
sourcesite | embedWhere the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published.
modefree | daily | challenge | race | session | (boş)Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it.
color_visionnormal | rg | by | contrast | (boş)In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments.
first_play1Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data.
seen_beforeyes | no | unsure | (boş)Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask).
duration_snumberTime from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s).
play_scorenumberLeaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments.

Exclusion criteria

Only answers in the current format (format 2) are published. The server requires both thresholds between 0.05 and 80 and the level between 0 and the maximum; plays at a humanly impossible speed are not ranked.

  • Only each participant's first answer to this experiment; replays are not included.
  • Simulated rows, bots, banned and sample accounts are not included.
  • Answers before 3 October 2026 are not included (the first day the open data notice was live).
  • Answers given in the old task format (format 1) are not included; only the current format is published.

Scoring

Leaderboard score: number of levels passed in round 3; higher is better.

Leaderboard measure: level reached (higher is better). The play_score column in the open data is this score for the first play.

Version notes

  • October 2026

    Switched to the arcade format (format 2): the task and the score scale changed, and older plays were removed from the leaderboard. The open data contains format 2 answers only.

  • October 2026

    The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.