Method card · Spatial thinking

Mirror or not?.

Look at the rotated letters and shapes: is each one normal, or its mirror image? Answer as quickly and accurately as you can.

Go to experiment2 min long3 real first answers

Original study

  1. Shepard, R. N., & Metzler, J. (1971). Mental rotation of three-dimensional objects. Science, 171(3972), 701–703. Source
  2. Cooper, L. A., & Shepard, R. N. (1973). Chronometric studies of the rotation of mental images. In W. G. Chase (Ed.), Visual information processing (pp. 75–176). Academic Press. Source
  3. Quan, C., Li, C., Xue, J., Yue, J., & Zhang, C. (2017). Mirror-normal difference in the late phase of mental rotation: An ERP study. PLOS ONE, 12(9), e0184963. Source
  4. Voyer, D., Voyer, S., & Bryden, M. P. (1995). Magnitude of sex differences in spatial abilities: A meta-analysis and consideration of critical variables. Psychological Bulletin, 117(2), 250–270. Source
  5. Voyer, D. (2011). Time limits and gender differences on paper-and-pencil tests of mental rotation: A meta-analysis. Psychonomic Bulletin & Review, 18(2), 267–277. Source
  6. Uttal, D. H., Meadow, N. G., Tipton, E., Hand, L. L., Alden, A. R., Warren, C., & Newcombe, N. S. (2013). The malleability of spatial skills: A meta-analysis of training studies. Psychological Bulletin, 139(2), 352–402. Source
  7. Cooper, L. A. (1975). Mental rotation of random two-dimensional shapes. Cognitive Psychology, 7(1), 20–43. Source

Our adaptation

Cooper and Shepard's (1973) single-stimulus normal-versus-mirror task. In round 1, over 8 trials an asymmetric letter (F, G, J, R) is shown rotated by 0°, 60°, 120° or 180°, either normal or mirror-reversed, and the participant presses 'Normal' or 'Mirror'. Round 2 repeats the task with two-dimensional shapes made of squares, and round 3 uses angles between 0° and 300° with a 3-second limit per trial.

Differences from the original study

  • Letters and flat shapes are used instead of Shepard and Metzler's (1971) three-dimensional objects.
  • Each round has 8 trials; there are few trials per angle.
  • Touchscreen latency and device differences affect response times.

Measured variables

The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.

ColumnTypeDescription
rt_0numberRound 1: median response time of correct trials by angle (ms); the number at the end of the column is the angle (0, 60, 120, 180). (0)
rt_60numberRound 1: median response time of correct trials by angle (ms); the number at the end of the column is the angle (0, 60, 120, 180). (60)
rt_120numberRound 1: median response time of correct trials by angle (ms); the number at the end of the column is the angle (0, 60, 120, 180). (120)
rt_180numberRound 1: median response time of correct trials by angle (ms); the number at the end of the column is the angle (0, 60, 120, 180). (180)
accnumberRound 1: accuracy (0–1).
rt_slopenumberRound 1: difference between the 180° and 0° medians (ms); a rough measure of mental rotation.
shapes_accnumberRound 2 (shapes): accuracy.
shapes_rtnumberRound 2: median time of correct trials (ms).
hard_accnumberRound 3 (3-second limit, 0°–300°): accuracy.
hard_rtnumberRound 3: median time of correct trials (ms).
hard_timeoutsintegerRound 3: number of trials missed because time ran out.
Columns present in every file (14)
row_idstringRow code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments.
experimentstringShort name of the experiment (URL slug).
experiment_versionintegerVersion of the answer format. For experiments whose format changed, only the current format is published.
dateYYYY-MM-DD | YYYY-Www | YYYY-MMDay the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few.
date_precisionday | week | monthPrecision of the date column.
langtr | en | (boş)Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.
devicedesktop | mobile | tablet | (boş)Device class reported by the browser (class only; browser details are neither stored here nor published).
sourcesite | embedWhere the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published.
modefree | daily | challenge | race | session | (boş)Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it.
color_visionnormal | rg | by | contrast | (boş)In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments.
first_play1Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data.
seen_beforeyes | no | unsure | (boş)Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask).
duration_snumberTime from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s).
play_scorenumberLeaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments.

Exclusion criteria

Only correct answers at or above 200 ms enter the medians. The server requires a median for every angle between 200 and 8,000 ms and an accuracy between 0 and 1; incomplete or out-of-range submissions are rejected.

  • Only each participant's first answer to this experiment; replays are not included.
  • Simulated rows, bots, banned and sample accounts are not included.
  • Answers before 3 October 2026 are not included (the first day the open data notice was live).

Scoring

Leaderboard score: chance-corrected accuracy combined with speed: max(0, 2 × mean accuracy − 1) × 100 × 1500 / max(400, median of the angle medians). Mean accuracy is the mean over the three rounds. Higher is better.

Leaderboard measure: score (higher is better). The play_score column in the open data is this score for the first play.

Version notes

  • October 2026

    The leaderboard score was corrected for chance (2 × accuracy − 1 instead of accuracy, combined with speed); first-play scores were recomputed from the stored answers with the new formula. Answers did not change.

  • October 2026

    Plausibility bounds were tightened: anticipatory responses (below 200 ms for Stroop and rotation, below 100 ms for reaction time) are left out of the median on the client, and the server rejects shorter medians (the previous floor was 150 ms, 80 ms for reaction time).

  • October 2026

    This experiment started storing the play mode (free, daily, challenge, race, session) with the answer; the mode column is empty for earlier answers.

  • October 2026

    The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.