Method card · Hearing

Match the pitch.

Which of two tones played one after the other is higher? And how accurately can you match a tone you just heard by sliding a strip on the screen?

Go to experiment3 min long0 real first answers

Original study

  1. Wier, C. C., Jesteadt, W., & Green, D. M. (1977). Frequency discrimination as a function of frequency and sensation level. The Journal of the Acoustical Society of America, 61(1), 178–184. Source
  2. Micheyl, C., Xiao, L., & Oxenham, A. J. (2012). Characterizing the dependence of pure-tone frequency difference limens on frequency, duration, and level. Hearing Research, 292(1–2), 1–13. Source
  3. Micheyl, C., Delhommeau, K., Perrot, X., & Oxenham, A. J. (2006). Influence of musical and psychoacoustical training on pitch discrimination. Hearing Research, 219(1–2), 36–47. Source
  4. Hutchins, S., & Peretz, I. (2012). A frog in your throat or in your ear? Searching for the causes of poor singing. Journal of Experimental Psychology: General, 141(1), 76–97. Source
  5. Peretz, I., & Vuvan, D. T. (2017). Prevalence of congenital amusia. European Journal of Human Genetics, 25(5), 625–630. Source
  6. Levitt, H. (1971). Transformed up-down methods in psychoacoustics. The Journal of the Acoustical Society of America, 49(2B), 467–477. Source
  7. Hollingworth, H. L. (1910). The central tendency of judgment. The Journal of Philosophy, Psychology and Scientific Methods, 7(17), 461–469. Source

Our adaptation

A pitch discrimination threshold (Wier et al., 1977) and a pitch matching task. Round 1 asks which of two successive pure tones is higher; the base tone is drawn at random from 300–900 Hz, the difference starts at 100 cents, is multiplied by 0.7 after two correct answers and by 1.4 after each error; the round ends after 8 reversals or 30 trials and the threshold is the geometric mean of the last 6 reversals. In round 2 (5 trials) a target tone drawn from 220–880 Hz plays twice, and after a 1.5-second silent wait the tone is searched for by dragging a strip covering 150–1,300 Hz on a log scale.

Differences from the original study

  • Speakers, headphones and background noise are not controlled; headphones in a quiet room are recommended.
  • Tones are generated with Web Audio; level is not calibrated.
  • Matching is done by moving a strip, not by singing.

Measured variables

The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.

ColumnTypeDescription
thnumberRound 1: pitch discrimination threshold (cents; 100 cents = a semitone).
accnumberRound 1: accuracy.
tuneJSONRound 2 trials: [[target Hz, matched Hz, time in ms] × 5].
cents_1numberRound 2: deviation of the matched tone from the target (cents; positive = sharp); the number at the end of the column is the trial order. (item 1)
cents_2numberRound 2: deviation of the matched tone from the target (cents; positive = sharp); the number at the end of the column is the trial order. (item 2)
cents_3numberRound 2: deviation of the matched tone from the target (cents; positive = sharp); the number at the end of the column is the trial order. (item 3)
cents_4numberRound 2: deviation of the matched tone from the target (cents; positive = sharp); the number at the end of the column is the trial order. (item 4)
cents_5numberRound 2: deviation of the matched tone from the target (cents; positive = sharp); the number at the end of the column is the trial order. (item 5)
pts_1numberRound 2 trial score (0–10): 10 / (1 + (|deviation| / 45)^1.6). (item 1)
pts_2numberRound 2 trial score (0–10): 10 / (1 + (|deviation| / 45)^1.6). (item 2)
pts_3numberRound 2 trial score (0–10): 10 / (1 + (|deviation| / 45)^1.6). (item 3)
pts_4numberRound 2 trial score (0–10): 10 / (1 + (|deviation| / 45)^1.6). (item 4)
pts_5numberRound 2 trial score (0–10): 10 / (1 + (|deviation| / 45)^1.6). (item 5)
pointsnumberTotal score of round 2 (0–50).
biasnumberMean signed deviation (cents).
pullnumberPull toward the center: slope of the signed deviation against the target's position in octaves relative to 440 Hz (cents / octave).
Columns present in every file (14)
row_idstringRow code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments.
experimentstringShort name of the experiment (URL slug).
experiment_versionintegerVersion of the answer format. For experiments whose format changed, only the current format is published.
dateYYYY-MM-DD | YYYY-Www | YYYY-MMDay the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few.
date_precisionday | week | monthPrecision of the date column.
langtr | en | (boş)Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.
devicedesktop | mobile | tablet | (boş)Device class reported by the browser (class only; browser details are neither stored here nor published).
sourcesite | embedWhere the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published.
modefree | daily | challenge | race | session | (boş)Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it.
color_visionnormal | rg | by | contrast | (boş)In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments.
first_play1Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data.
seen_beforeyes | no | unsure | (boş)Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask).
duration_snumberTime from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s).
play_scorenumberLeaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments.

Exclusion criteria

Only answers in the current format (format 2) are published. The server requires a threshold between 0.5 and 1,200 cents and all five matching trials (target and answer 100–2,000 Hz); it computes deviations and scores itself.

  • Only each participant's first answer to this experiment; replays are not included.
  • Simulated rows, bots, banned and sample accounts are not included.
  • Answers before 3 October 2026 are not included (the first day the open data notice was live).
  • Answers given in the old task format (format 1) are not included; only the current format is published.

Scoring

Leaderboard score: total score of round 2 (0–50); higher is better.

Leaderboard measure: pitch matching (higher is better). The play_score column in the open data is this score for the first play.

Version notes

  • October 2026

    Switched to the arcade format (format 2): the task and the score scale changed, and older plays were removed from the leaderboard. The open data contains format 2 answers only.

  • October 2026

    The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.