Does a good word make grey look brighter?
An online brightness judgement experiment, with the demo, the timing and the power analysis on one page
People tend to link good with light and bad with dark. Meier, Robinson and Clore (2004) found that people judge words faster when positive words appear in white and negative words in black, and Meier et al. (2007) found the link runs the other way too: what a word means can push a later brightness judgement. This study tests that second claim online.
On each trial a participant reads a word, then sees a grey square and says whether it is brighter or darker than a reference square they saw at the start. In the main block every square is the reference grey, rgb(168, 168, 168). Nothing about the stimulus changes, so if positive words still get more “brighter” answers than negative ones, the effect comes from the viewer.
It runs on jsPsych (de Leeuw, 2015) and recruits through Prolific. You can take a shortened version, watch a simulated participant take it, or go straight to the results screen. The rest of this page lets you try the task here and look at two problems I had to solve along the way.
1. The task
Below is the main block cut down to twelve trials, four from each word category. Timing matches the real study: a cross for 800 ms, the word for 400 ms, then the square until you answer. Press F if the square looks brighter than the one shown before you start, J if darker. On a phone, tap the buttons.
Reference square. Keep it in mind.
| # | word | category | answer | rt | word ms |
|---|---|---|---|---|---|
| No trials yet. | |||||
All twelve squares were rgb(168, 168, 168), identical to the reference. Figure 2 shows how often you said brighter after each kind of word. With twelve trials this is mostly noise; the full study has 80 trials per person and the test is across people.
2. Frames
A browser can only change what is on screen once per frame. On a 60 Hz monitor a frame lasts about 16.7 ms, so a 400 ms word is really 24 frames, and one late frame makes it 417 ms. Bridges et al. (2020) measured this kind of timing error across experiment software and browsers.
The default jsPsych plugins use timers for this. I wrote a plugin that runs the cross, word and square as one trial and makes every change inside requestAnimationFrame. It changes the display on the frame where less than half a frame of the target time is left, and writes the measured durations into the data, so a dropped frame shows up in the row for that trial.
Measuring…
| display | frames | shown | error |
|---|
3. Sample size
Before running it I wanted to know how many people the study needs. Effects like this are small and people differ, so I simulated it. Each simulated person gets their own bias, drawn around the true effect, and answers each trial by coin flip weighted by that bias. The group is then tested with a one sample t test on each person's positive minus negative difference. Repeat that a thousand times and the share of significant results is the power.
4. Implementation
The site is plain HTML and JavaScript with no build step. The pieces worth reading:
- src/plugin-word-probe.js
- The trial plugin described in section 2. It records the repeat key and the brightness answer separately, which fixed a bug in the first version where pressing Space on a repeated word threw away the brightness answer. It also implements jsPsych's simulation mode, which is how the simulated participant works.
- src/analysis.js
- Per category summaries and reaction time exclusions. The repeat task is scored as d′ with the log linear correction from Hautus (1995), so a perfect score doesn't produce an infinite value. The same file runs at the end of the study, on Figure 2 above, and in the Node tests.
- src/design.js
- Trial order, repeated words, where the break goes (never between a word and its repeat) and data file names. Every session is seeded and the seed is saved, so adding ?seed= to the URL replays a session's exact order.
- tests/
- Unit tests in Node, and Playwright tests that run whole simulated sessions, press real keys, decline consent and check the CSV download. CI also checks every CDN script against a pinned version and its integrity hash.
Data is never sent anywhere. At the end of a session the browser saves it as a CSV file, and in Prolific mode that happens automatically before the participant is sent back.
References
- Bridges, D., Pitiot, A., MacAskill, M. R., & Peirce, J. W. (2020). The timing mega-study: comparing a range of experiment generators, both lab-based and online. PeerJ, 8, e9414.
- de Leeuw, J. R. (2015). jsPsych: A JavaScript library for creating behavioral experiments in a Web browser. Behavior Research Methods, 47(1).
- Hautus, M. J. (1995). Corrections for extreme proportions and their biasing effects on estimated values of d′. Behavior Research Methods, Instruments, & Computers, 27(1).
- Meier, B. P., Robinson, M. D., & Clore, G. L. (2004). Why good guys wear white: Automatic inferences about stimulus valence based on brightness. Psychological Science, 15(2).
- Meier, B. P., Robinson, M. D., Crawford, L. E., & Ahlvers, W. J. (2007). When “light” and “dark” thoughts become light and dark responses: Affect biases brightness judgments. Emotion, 7(2).