Pen in mouth.
Can you rate four cartoons with a pen in your mouth?
Original study
- Strack, F., Martin, L. L., & Stepper, S. (1988). Inhibiting and facilitating conditions of the human smile: A nonobtrusive test of the facial feedback hypothesis. Journal of Personality and Social Psychology, 54(5), 768–777. Source
- Wagenmakers, E.-J., Beek, T., Dijkhoff, L., Gronau, Q. F., Acosta, A., Adams, R. B., Jr., … Zwaan, R. A. (2016). Registered Replication Report: Strack, Martin, & Stepper (1988). Perspectives on Psychological Science, 11(6), 917–928. Source
- Strack, F. (2016). Reflection on the smiling Registered Replication Report. Perspectives on Psychological Science, 11(6), 929–930. Source
- Noah, T., Schul, Y., & Mayo, R. (2018). When both the original study and its failed replication are correct: Feeling observed eliminates the facial-feedback effect. Journal of Personality and Social Psychology, 114(5), 657–664. Source
- Coles, N. A., Larsen, J. T., & Lench, H. C. (2019). A meta-analysis of the facial feedback literature: Effects of facial feedback on emotional experience are small and variable. Psychological Bulletin, 145(6), 610–651. Source
- Marsh, A. A., Rhoads, S. A., & Ryan, R. M. (2019). A multi-semester classroom demonstration yields evidence in support of the facial feedback effect. Emotion, 19(8), 1500–1504. Source
- Coles, N. A., March, D. S., Marmolejo-Ramos, F., Larsen, J. T., Arinze, N. C., Ndukaihe, I. L. G., … Liuzza, M. T. (2022). A multi-lab test of the facial feedback hypothesis by the Many Smiles Collaboration. Nature Human Behaviour, 6(12), 1731–1742. Source
Our adaptation
An unsupervised at-home adaptation of the facial feedback experiment of Strack, Martin and Stepper (1988). Participants are split into two groups with a personal seed: they hold a pen horizontally between their teeth (smile-like) or with their lips only (smile-inhibiting); nobody is told to smile or frown. With the pen in the mouth, four original cartoons drawn for balabs are rated from 0 to 9 in shuffled order. Afterwards participants report whether they could hold the pen, whether they understood the cartoons and what they think the experiment measured.
Differences from the original study
- Nobody sees whether the pen is held correctly; we rely on self-report.
- The cartoons are not the original Far Side cartoons but were drawn for this experiment.
- The page name and questions may reveal the purpose to some participants.
Measured variables
The columns of the open data file. List fields are expanded into numbered columns (for example rt_60, est_3); JSON columns contain only numbers and fixed stimulus names.
| Column | Type | Description |
|---|---|---|
cond | teeth | lips | Group: teeth = between the teeth, lips = with the lips. |
pen | integer | Used a pen or similar object (1 = yes). |
ratings | JSON | Cartoon ratings: [[cartoon, 0–9] × 4] in presentation order (penguin, pizza, cloud, maze). |
hold | all | most | few | Could hold the pen as described (all | most | few). |
understood | yes | some | no | Understood the cartoons (yes | some | no). |
guess | pen | read | taste | dk | Guess of what was measured (pen = effect of the pen, read = reading, taste = sense of humor, dk = don't know). |
mean | number | Mean of the four ratings. |
valid | 0 | 1 | Enters the main comparison: used a pen, could hold it, understood the cartoons and did not guess 'pen'. |
Columns present in every file (14)
row_id | string | Row code. Regenerated at random in every release; it does not identify a person and cannot be matched across releases or experiments. |
experiment | string | Short name of the experiment (URL slug). |
experiment_version | integer | Version of the answer format. For experiments whose format changed, only the current format is published. |
date | YYYY-MM-DD | YYYY-Www | YYYY-MM | Day the answer was given (Istanbul time). If fewer than 5 answers share the same day, language and device, it is coarsened to the ISO week, and to the month if that is still too few. |
date_precision | day | week | month | Precision of the date column. |
lang | tr | en | (boş) | Interface language. Language recording started in the first week of October 2026; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise. |
device | desktop | mobile | tablet | (boş) | Device class reported by the browser (class only; browser details are neither stored here nor published). |
source | site | embed | Where the answer came from: the balabs site or the experiment embedded on another site. The embedding site is not published. |
mode | free | daily | challenge | race | session | (boş) | Play mode in scored experiments: free play, daily round, challenge, race room, experiment session. Empty for unscored experiments and for older answers that did not store it. |
color_vision | normal | rg | by | contrast | (boş) | In color-based experiments, the color vision mode the participant chose: normal, red-green, blue-yellow or high contrast. Stimuli are generated along different axes per mode, so compare within a mode. Empty for other experiments. |
first_play | 1 | Every row is a participant’s first play of this experiment (always 1). Replays are stored only as leaderboard scores and are not part of the open data. |
seen_before | yes | no | unsure | (boş) | Participant’s own report of whether they had seen this experiment or its known answer before (in experiments that ask). |
duration_s | number | Time from the start of the experiment to submitting the answer, in seconds (measured in the browser, 0.1 s). |
play_score | number | Leaderboard score of the first play (scored experiments; defined on the method card). Empty for unscored experiments. |
Exclusion criteria
The server requires the group, a single rating for each of the four cartoons and the answers to the questions, and computes 'valid' itself. Only plays with valid = 1 are used in the main comparison (as in the Wagenmakers et al., 2016 replication).
- Only each participant's first answer to this experiment; replays are not included.
- Simulated rows, bots, banned and sample accounts are not included.
- Answers before 3 October 2026 are not included (the first day the open data notice was live).
Scoring
This experiment has no leaderboard score. The main measure is the difference in mean funniness ratings between the groups.
Version notes
- October 2026
The interface language (lang) started being stored with each answer; for earlier answers it is known only in experiments whose stimuli depend on the language, and empty otherwise.