The chimp test.
If a few numbers flash on screen for the blink of an eye and vanish, can you point to where they were in order, smallest to largest?
Original study
- Inoue, S., & Matsuzawa, T. (2007). Working memory of numerals in chimpanzees. Current Biology, 17(23), R1004–R1005. Source
- Silberberg, A., & Kearns, D. (2009). Memory for the order of briefly presented numerals in humans as a function of practice. Animal Cognition, 12(2), 405–407. Source
- Cook, P., & Wilson, M. (2010). Do young chimpanzees have extraordinary working memory? Psychonomic Bulletin & Review, 17(4), 599–600. Source
Our adaptation
An adaptation of Inoue and Matsuzawa's (2007) numerical working memory task. On an 8 × 5 grid, tapping the central circle makes a few numbers from 1–9 appear at random positions; they turn into white squares shortly after and the squares are tapped in ascending order, with the trial ending at the first error. Plan: five numbers at 650, 430 and 210 ms, then six and seven numbers at 210 ms. Presentation time is measured frame by frame and stored.
Differences from the original study
- The comparison with Ayumu is rough: his rate is based on hundreds of trials, this one on five.
- Because of the screen's frame rate, numbers may stay visible a few milliseconds longer than intended.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
t | JSON | Trials: [[round, count of numbers, target display ms, measured display ms, correct before first error, answer ms], …]. |
pts_1 | number | Round score: 10 × mean of (correct before first error / count). (item 1) |
pts_2 | number | Round score: 10 × mean of (correct before first error / count). (item 2) |
pts_3 | number | Round score: 10 × mean of (correct before first error / count). (item 3) |
pts_4 | number | Round score: 10 × mean of (correct before first error / count). (item 4) |
pts_5 | number | Round score: 10 × mean of (correct before first error / count). (item 5) |
points | number | Total score (0–50). |
acc_5_650 | number | Share of perfect trials per condition; column suffix 'count_time ms'. (5-650) |
acc_5_430 | number | Share of perfect trials per condition; column suffix 'count_time ms'. (5-430) |
acc_5_210 | number | Share of perfect trials per condition; column suffix 'count_time ms'. (5-210) |
acc_6_210 | number | Share of perfect trials per condition; column suffix 'count_time ms'. (6-210) |
acc_7_210 | number | Share of perfect trials per condition; column suffix 'count_time ms'. (7-210) |
ayumu | number | Share of perfect trials in the five numbers, 210 ms condition. |
maxn | integer | Largest count completed perfectly. |
over | number | Mean excess of measured over target display time (ms). |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server checks that the number of trials, the counts and the target times follow the plan and that the measured display time is within a plausible range.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: total score (0–50); higher is better.
Leaderboard measure: number memory (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.