Method card · Memory

The chimp test.

If a few numbers flash on screen for the blink of an eye and vanish, can you point to where they were in order, smallest to largest?

Go to experiment2 min long0 real first answers

Original study

  1. Inoue, S., & Matsuzawa, T. (2007). Working memory of numerals in chimpanzees. Current Biology, 17(23), R1004–R1005. Source
  2. Silberberg, A., & Kearns, D. (2009). Memory for the order of briefly presented numerals in humans as a function of practice. Animal Cognition, 12(2), 405–407. Source
  3. Cook, P., & Wilson, M. (2010). Do young chimpanzees have extraordinary working memory? Psychonomic Bulletin & Review, 17(4), 599–600. Source

Our adaptation

An adaptation of Inoue and Matsuzawa's (2007) numerical working memory task. On an 8 × 5 grid, tapping the central circle makes a few numbers from 1–9 appear at random positions; they turn into white squares shortly after and the squares are tapped in ascending order, with the trial ending at the first error. Plan: five numbers at 650, 430 and 210 ms, then six and seven numbers at 210 ms. Presentation time is measured frame by frame and stored.

Differences from the original study

  • The comparison with Ayumu is rough: his rate is based on hundreds of trials, this one on five.
  • Because of the screen's frame rate, numbers may stay visible a few milliseconds longer than intended.

Measured variables

The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.

ColumnTypeDescription
tJSONTrials: [[round, count of numbers, target display ms, measured display ms, correct before first error, answer ms], …].
pts_1numberRound score: 10 × mean of (correct before first error / count). (item 1)
pts_2numberRound score: 10 × mean of (correct before first error / count). (item 2)
pts_3numberRound score: 10 × mean of (correct before first error / count). (item 3)
pts_4numberRound score: 10 × mean of (correct before first error / count). (item 4)
pts_5numberRound score: 10 × mean of (correct before first error / count). (item 5)
pointsnumberTotal score (0–50).
acc_5_650numberShare of perfect trials per condition; column suffix 'count_time ms'. (5-650)
acc_5_430numberShare of perfect trials per condition; column suffix 'count_time ms'. (5-430)
acc_5_210numberShare of perfect trials per condition; column suffix 'count_time ms'. (5-210)
acc_6_210numberShare of perfect trials per condition; column suffix 'count_time ms'. (6-210)
acc_7_210numberShare of perfect trials per condition; column suffix 'count_time ms'. (7-210)
ayumunumberShare of perfect trials in the five numbers, 210 ms condition.
maxnintegerLargest count completed perfectly.
overnumberMean excess of measured over target display time (ms).
Columns present in every file (14)
row_idstringRow code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments.
experimentstringShort name of the experiment (URL slug).
experiment_versionintegerVersion of the answer format. For experiments whose format changed, only the current format is published.
dateYYYY-MM-DD | YYYY-Www | YYYY-MMDay the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few.
date_precisionday | week | monthPrecision of the date column.
langtr | en | (boş)Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.
devicedesktop | mobile | tablet | (boş)Device class reported by the browser (class only; browser details are neither stored here nor published).
sourcesite | embedWhere the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published.
modefree | daily | challenge | race | session | (boş)Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it.
color_visionnormal | rg | by | contrast | (boş)In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments.
first_play1Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data.
seen_beforeyes | no | unsure | (boş)Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask).
duration_snumberTime from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s).
play_scorenumberLeaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments.

Exclusion criteria

The server checks that the number of trials, the counts and the target times follow the plan and that the measured display time is within a plausible range.

  • Only each participant's first answer to this experiment; replays are not included.
  • Simulated rows, bots, banned and sample accounts are not included.
  • Answers before 3 October 2026 are not included (the first day the open data notice was live).

Scoring

Leaderboard score: total score (0–50); higher is better.

Leaderboard measure: number memory (higher is better). The play_score column in the open data is this score for the first play.

Version notes

  • October 2026

    The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.