Number line.
Where does 150 sit on a line from 0 to 1000? How accurately do you place numbers on the ruler in your head?
Original study
- Siegler, R. S., & Opfer, J. E. (2003). The development of numerical estimation: Evidence for multiple representations of numerical quantity. Psychological Science, 14(3), 237–243. Source
- Booth, J. L., & Siegler, R. S. (2006). Developmental and individual differences in pure numerical estimation. Developmental Psychology, 42(1), 189–201. Source
- Opfer, J. E., & Siegler, R. S. (2007). Representational change and children's numerical estimation. Cognitive Psychology, 55(3), 169–195. Source
- Dehaene, S., Izard, V., Spelke, E., & Pica, P. (2008). Log or linear? Distinct intuitions of the number scale in Western and Amazonian indigene cultures. Science, 320(5880), 1217–1220. Source
- Barth, H. C., & Paladino, A. M. (2011). The development of numerical estimation: Evidence against a representational shift. Developmental Science, 14(1), 125–135. Source
- Anobile, G., Cicchini, G. M., & Burr, D. C. (2012). Linear mapping of numbers onto space requires attention. Cognition, 122(3), 454–459. Source
Our adaptation
An adaptation of Siegler and Opfer's (2003) number-line estimation task. In each of five rounds four numbers are shown one at a time and placed on a line labeled only 0 and 100 at its ends (rounds 1 and 4) or 0 and 1,000 (rounds 2, 3, 5). Numbers are drawn from bins covering every region of the line; multiples of 50 on 0–1,000 and of 25 on 0–100 are not used.
Differences from the original study
- Participants are adults; the original studies examined children's development.
- Four or eight numbers are too few to reliably separate linear and logarithmic models for one person.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
d | JSON | Numbers: [[round, scale 100/1000, number, estimate, time in ms] × 20]. |
pts_1 | number | Round score: mean of four numbers, number score 10 / (1 + (|estimate − number| / scale / 0.05)^1.6). (item 1) |
pts_2 | number | Round score: mean of four numbers, number score 10 / (1 + (|estimate − number| / scale / 0.05)^1.6). (item 2) |
pts_3 | number | Round score: mean of four numbers, number score 10 / (1 + (|estimate − number| / scale / 0.05)^1.6). (item 3) |
pts_4 | number | Round score: mean of four numbers, number score 10 / (1 + (|estimate − number| / scale / 0.05)^1.6). (item 4) |
pts_5 | number | Round score: mean of four numbers, number score 10 / (1 + (|estimate − number| / scale / 0.05)^1.6). (item 5) |
points | number | Total score (0–50). |
pae | number | Mean percent absolute error (all numbers). |
pae1000 | number | Mean percent absolute error on the 0–1,000 line. |
pae100 | number | Mean percent absolute error on the 0–100 line. |
fit | JSON | Model fits: {"1000": {lin, log: R², a, b: linear coefficients, la, lb: logarithmic coefficients}, "100": {…}}. |
best | lin | log | Better-fitting model on the 0–1,000 line (lin | log). |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server checks that the twenty numbers follow the plan by round and scale, that each number is below the scale and each estimate within it.
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
Leaderboard score: total score (0–50); higher is better.
Leaderboard measure: number line (higher is better). The play_score column in the open data is this score for the first play.
Version notes
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.