Replication of Gantman et al. (2020)

Author

Edwin Cortazo

Introduction

This is a replication of the study by Gantman et al. (2020). Replication studies play a crucial role in ensuring the reliability and validity of scientific findings. By conducting replications, researchers can assess the generality of previous results and confirm replicability. Replications also serve to identify potential sources of variation and confounding factors. In this study, we aim to replicate the findings of Gantman et al. (2020), who investigated the “moral pop-out effect” using event-related potentials (ERPs).

The moral pop-out effect suggests that the “morality” of a visual stimulus is prioritized in the early stages of perception, leading to faster recognition of moral words compared to non-moral words. Gantman et al. (2020) found marginal support for this effect, with participants responding slightly faster to moral words than non-moral words, although the difference was not statistically significant. They also observed that moral and non-moral words were distinguishable from one another as early as 300 ms after word presentation.

However, it remains unclear whether this effect is specific to the moral domain or if the observed effect could instead be explained by semantic priming resulting from seeing moral words repeatedly throughout the experiment. To account for this potential confound, we aim to replicate the study by incorporating an additional category of fashion words, as suggested by Firestone and Scholl (2015). If the pop-out effect holds true, we expect to replicate the behavioral and ERP effects for moral but not fashion words. By including this control condition and directly comparing moral and fashion words, this study aims to provide a more comprehensive understanding of the nature of the moral pop-out effect.

Methods

Participants

The data set contains 42 sessions from native English-speaking Vassar College students. Participants were aged 18 years or older and had normal or corrected vision and provided informed consent prior to the experiment. Participants were compensated with $20 and in some cases, course credit upon completion of experiment. This study was approved by the Vassar College Institutional Review Board. After applying the exclusion criteria, 34 subjects remain. 7 subjects were excluded for receiving fashion words twice due to a technical error, 1 subject was excluded for having poor EEG recordings, 0 were excluded for failing to reply on more than 50% of trials and 0 were excluded for having a non-word response rate greater than 90%.

Materials

The experiment was designed using jsPsych (www.jspsych.org) and utilized subsets of stimuli (word lists) from Gantman et al. (2020) and Firestone and Scholl (2015). The stimuli consisted of eight distinct categories: non-moral words, moral words, non-moral non-words, moral non-words, fashion words, fashion non-words, non-fashion words, and non-fashion non-words. Non-words were created by scrambling the letters of corresponding words.

The experiment was conducted using an ASUS VG248QE monitor with full HD 1080p resolution and a 144hz refresh rate. Participants used a Lenovo SK-8825 (L) wired black USB keyboard for response inputs. EEG data was recorded using a CGX Quick-20r wireless, battery-operated, full standard 10-20 montage EEG headset with dry sensor technology, sampling at 500Hz with 24-bit resolution. A CGX Wireless Stim Trigger was used for 16-bit simultaneous event marking with millisecond precision.

Procedure

Participants completed the experiment individually in 90 minute sessions (approx. 20 minutes in EEG) in a dimly lit room. The experiment began with 20 practice trials of 10 non-moral words and 10 non-moral non-words with decreasing intervals of 300, 100, 60, 30, 16ms. Participants were instructed to sit 60cm away from the screen and rest their arms in a comfortable position.

The main experiment consisted of two blocks of trials, each containing 300 trials (75 words and 75 non-words for each category), for a total of 600 trials. The order of the blocks (Block A: moral or fashion first) is randomly determined for each participant. Participants had a short break after every 100 trials.

Each trial followed the same structure:

  1. Fixation screen presented for 400-700ms
  2. Stimulus (letter string) presented for 16.6ms
  3. Fixation screen presented for 33.33ms
  4. Backward mask of ampersands (&) corresponding to the number of letters in the word, presented for 25ms
  5. Blank screen presented for 1500ms for participant response

Participants pressed the ‘1’ key if the string of letters appeared as a word and the ‘5’ key if it appeared as a non-word.

EEG Data was recorded from the Cz and Pz electrode, with additional sensors placed at C3, C4, P3, P4, Fz, F3, F4.

Analysis

Accuracy and mean ERP amplitude in each time window were modelled with generalized estimating equations (GEE, R package gee) with an exchangeable working correlation and participants as clusters. All SE, z and p values are robust (sandwich) estimates.

OSF Project and Preregistration

A preregistration for this study, stimuli and experiment scripts are available on the Open Science Framework at https://osf.io/9ygfj/. The preprocessed data are at https://osf.io/ngtha/.

Results

Behavioral

Figure 1: Accuracy for in-category words against out-of-category words, one point per participant. Points above the diagonal represent higher accuracy for in-category words. Overall accuracy was much higher than expected. We later discovered this was due to a difference in procedure between our study and the original study.

A pop-out effect was found for fashion words, but not moral words.

In the moral condition, participants were 92.7% accurate for moral words and 91.6% accurate for non-moral words. This difference was not statistically significant in the GEE model, β = -0.15, SE = 0.09, z = -1.65, p = .099.

In the fashion condition, participants were 92.9% accurate for fashion words and 88.4% accurate for non-fashion words. This difference was statistically significant in the GEE model, β = -0.55, SE = 0.12, z = -4.67, p < .001.

EEG

Figure 2: Grand average waveforms. Dotted lines mark the edges of the P2 (200 to 250 ms), N2 (250 to 350 ms), P3 (350 to 600 ms) and LPP (600 to 800 ms) windows.

Words vs. Non-Words

Following Gantman et al. (2020), we looked for word vs. non-word ERP effects at each time window at electrode Pz.

In the moral condition, words elicited a more positive ERP than non-words in the P2 window (β = 1.19, SE = 0.36, z = 3.31, p < .001), N2 window (β = 2.18, SE = 0.53, z = 4.12, p < .001) and P3 window (β = 2.82, SE = 0.71, z = 3.96, p < .001), but not in the LPP window (β = -0.30, SE = 0.62, z = -0.49, p = .625).

In the fashion condition, words elicited a more positive ERP than non-words in the P3 window (β = 2.02, SE = 0.71, z = 2.85, p = .004). There was no significant difference in the P2 window (β = 0.48, SE = 0.50, z = 0.95, p = .340), N2 window (β = 1.07, SE = 0.56, z = 1.90, p = .057) or LPP window (β = -0.57, SE = 0.68, z = -0.85, p = .398).

Pop-out effects

Also following Gantman et al. (2020), we looked for ERP differences related to the category vs. non-category distinction in all four time windows at electrode Cz. The coefficient is the non-category minus category difference.

In the moral condition, there were no significant differences between moral and non-moral words in the P2 window (β = 0.01, SE = 0.22, z = 0.06, p = .948), N2 window (β = 0.13, SE = 0.29, z = 0.44, p = .660), P3 window (β = 0.10, SE = 0.45, z = 0.22, p = .828) or LPP window (β = 0.32, SE = 0.40, z = 0.80, p = .421).

In the fashion condition, there were also no significant differences between fashion and non-fashion words in the P2 window (β = -0.05, SE = 0.42, z = -0.12, p = .905), N2 window (β = 0.12, SE = 0.37, z = 0.33, p = .744), P3 window (β = -0.78, SE = 0.45, z = -1.71, p = .087) or LPP window (β = -0.91, SE = 0.65, z = -1.39, p = .164).

Discussion

This replication study aimed to assess the generality and robustness of the moral pop-out effect as reported by Gantman et al. (2020), while taking into account potential confounding factors such as semantic priming. Our results do not support a moral pop-out effect.

In the behavioral data we found a pop-out effect for fashion words but not for moral words, with participants responding more accurately to fashion words than to non-fashion words. Gantman et al. reported a marginal effect for moral words, which we did not observe.

In the ERP data, words and non-words separated at Pz in the P2, N2 and P3 windows for moral words and in the P3 window for fashion words. This is a word recognition effect, not a pop-out effect. The pop-out comparison at Cz, category words against non-category words, showed no significant difference in any window for either category.

One notable difference between our study and Gantman et al. is the overall accuracy. Participants in our study responded much more accurately (~90%) than in Gantman et al. (~70%). We later discovered this was due to a procedural change where stimuli were presented for a longer duration than in the original study. This may have reduced the sensitivity of the task to detect differences between conditions. Taken together, our partial replication of Gantman et al. (2020) suggests that while moral words may be processed differently than non-moral words under certain conditions, this ‘pop-out effect’ is likely due to procedural variations and merits further research. Furthermore, this replication highlights the importance of carefully understanding procedures when attempting to replicate an effect.

Limitations

It is important to acknowledge the limitations of our study and consider how it might inform future research. First, our sample consisted entirely of college students from a single institution, which may limit the generalizability of the findings, despite being similar to the participant demographic used by Gantman et al. (2020). Second, while we aimed to control potential confounds, such as semantic priming, there may be other factors that could influence the pop-out effect. Additionally, this study focused on a specific set of moral and fashion words. To further delve into the generalizability of the pop-out effect, future studies could utilize a wider range of stimuli (perhaps testing the pop-out effect in a study with more than 2 categories), with varying semantic categories and stimulus types (e.g, other forms of stimulus such as sound or images and diverse categories). Finally, and as mentioned previously, our analysis is limited due to the procedural misunderstanding that led to the experiment not being accurately replicated in design. Future iterations of this research should take this into account.

Ultimately, our study contributes to the ongoing debate about the nature of moral perception. These findings highlight the importance of replication and call attention to the importance for careful control of procedural variables.

References

  1. Firestone, C. & Scholl, B. (2015). Enhanced visual awareness for morality and pajamas? Perception vs. memory in ‘top-down’ effects. Cognition, 136, 409-416.

  2. Gantman, A., Devraj-Kizuk, S., Mende-Siedlecki, P., Van Bavel, J., & Mathewson, K. (2020). The time course of moral perception: an ERP investigation of the moral pop-out effect. Social Cognitive and Affective Neuroscience, 15(2), 235-246.