Catch the average.
After seeing twelve circles for half a second, can you set their average size? And can you recognize one of them?
Original study
- Ariely, D. (2001). Seeing sets: Representation by statistical properties. Psychological Science, 12(2), 157–162. Source
- Chong, S. C., & Treisman, A. (2003). Representation of statistical properties. Vision Research, 43(4), 393–404. Source
- Parkes, L., Lund, J., Angelucci, A., Solomon, J. A., & Morgan, M. (2001). Compulsory averaging of crowded orientation signals in human vision. Nature Neuroscience, 4(7), 739–744. Source
- Myczek, K., & Simons, D. J. (2008). Better than average: Alternatives to statistical summary representations for rapid judgments of average size. Perception & Psychophysics, 70(5), 772–788. Source
- de Fockert, J., & Wolfenstein, C. (2009). Rapid extraction of mean identity from sets of faces. Quarterly Journal of Experimental Psychology, 62(9), 1716–1722. Source
- Haberman, J., & Whitney, D. (2007). Rapid extraction of mean emotion and gender from sets of faces. Current Biology, 17(17), R751–R753. Source
- Whitney, D., & Yamanashi Leib, A. (2018). Ensemble perception. Annual Review of Psychology, 69, 105–129. Source
Our adaptation
An adaptation of mean size (Ariely, 2001; Chong and Treisman, 2003) and mean orientation (Parkes et al., 2001) tasks. In size rounds (1, 3, 5) twelve circles of four sizes appear for 500 ms and, after a 300 ms blank, a single circle is adjusted to the mean diameter; rounds 3 and 5 then add a surprise membership question: which of two circles was in the set (one a real member, the other an exactly mean-sized lure that was never in the set). In orientation rounds (2, 4) the mean direction of twelve lines spread ±24° around a mean tilt is reproduced.
Differences from the original study
- Mean size is the arithmetic mean of diameters; averaging areas would give about 3.5% larger.
- Screen size changes the visual angle of the circles.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
r | JSON | Rounds: [[type 0 = size / 1 = orientation, target, spread, answer, membership answer (−1 not asked, 0 chose the mean-sized lure, 1 chose the real member), member-to-mean ratio, time in ms] × 5] (size as diameter relative to field width, orientation in degrees). |
pts_1 | number | Round score (0–10): size 10 / (1 + (|ln(answer / target)| / 0.08)^1.6), orientation 10 / (1 + (error° / 6)^1.6). (item 1) |
pts_2 | number | Round score (0–10): size 10 / (1 + (|ln(answer / target)| / 0.08)^1.6), orientation 10 / (1 + (error° / 6)^1.6). (item 2) |
pts_3 | number | Round score (0–10): size 10 / (1 + (|ln(answer / target)| / 0.08)^1.6), orientation 10 / (1 + (error° / 6)^1.6). (item 3) |
pts_4 | number | Round score (0–10): size 10 / (1 + (|ln(answer / target)| / 0.08)^1.6), orientation 10 / (1 + (error° / 6)^1.6). (item 4) |
pts_5 | number | Round score (0–10): size 10 / (1 + (|ln(answer / target)| / 0.08)^1.6), orientation 10 / (1 + (error° / 6)^1.6). (item 5) |
points | number | Total score (0–50). |
size_err | number | Mean absolute error in size rounds (%). |
size_bias | number | Mean signed deviation in size rounds (%; positive = larger). |
ori_err | number | Mean absolute error in orientation rounds (°). |
member_pick | number | Share of membership questions in which the real member was chosen (chance level 0.5). |
probes | integer | Number of membership probes asked. |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server checks the type of each round and the placement of the probes against the plan, and values against their ranges; inconsistent submissions are rejected.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: total score (0–50); higher is better.
Leaderboard measure: eye for the average (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.