-
PDF
- Split View
-
Views
-
Cite
Cite
Melanie A Mascarenhas, Konstantine K Zakzanis, Does performance invalidity undermine replication? Evidence from an undergraduate student participant sample, Archives of Clinical Neuropsychology, Volume 41, Issue 5, August 2026, acag044, https://doi.org/10.1093/arclin/acag044
Close - Share Icon Share
Abstract
Performance invalidity is well-documented in clinical and forensic neuropsychology, yet its potential influence on data quality in psychological research has received little attention. This raises concerns for reproducibility, particularly given widespread use of undergraduate students as research participants. The present study examined whether incorporating Performance Validity Tests (PVTs) into a research protocol provides a practical method for improving data quality and interpretation.
As part of a larger replication initiative, PVTs were embedded into a direct replication of a 2008 prescribed optimism study. 87 undergraduate students were classified as credible or non-credible based on PVT performance, and analyses were repeated within subgroups.
The effect replicated successfully across the sample (d = 0.98), consistent with the original study (d = 0.93). 17.2% of participants failed one or more PVTs. The credible subgroup yielded a larger effect (d = 1.02) than the full sample, whereas the non-credible subgroup produced a smaller effect (d = 0.87), and a directional reversal in the secondary analysis – a pattern inconsistent with prior literature and suggestive of systematic noise that disengaged responding may introduce into data. Further, credible and non-credible subgroups differed significantly on the primary outcome despite being demographically and experimentally equivalent.
Performance invalidity occurs at a non-trivial rate in undergraduate research samples and can influence study outcomes. Incorporating PVTs into research protocols offers a practical, scalable safeguard to improving data quality and transparency, highlighting how neuropsychology’s established tools could contribute solutions to reproducibility challenges in psychological science.
Introduction
Replication, often viewed as the “gold standard” (Jasny et al., 2011, p. 1225), is a crucial aspect of empirical research and generally involves recreating the methodology of an original study on another random sample from the same population (OSC, 2015). To this end, replication aims to confirm the validity of previous findings and determine the extent to which an effect is generalizable.
However, some maintain that disconfirmation is not necessarily a problem and view failed replications as an inevitable consequence of the scientific method. Indeed, disconfirmation can be viewed as a way to propel innovation in a given field (OSC, 2015). Others have pointed out that the way “success” is measured in a replication is often overly reductionistic, with statistically significant findings celebrated and nonsignificant findings labeled as failures (Anderson & Maxwell, 2016). Nonetheless, when research does not replicate, several considerations should be taken into account. Aspects of a theory may be incorrect (Brandt, 2013), there may be methodological issues or merely differences with a specific study or sample (Brandt et al., 2014), or an entire sub-discipline may be characterized by systemic issues (Chabris et al., 2012). From an empirical standpoint, this may generate mistrust in both the research process and the discipline as a whole.
In recent years, the replication crisis has been at the forefront of discussions on epistemological issues in psychological science. The crisis refers to systematic and methodological practices that may prevent successful replication of results (OSC, 2015). This discourse emerged in the early 2000s through a series of catalytic commentaries regarding falsely published findings (Ioannidis, 2005; Malich & Munafò, 2022) and with a resurgence of the discussion on prevalent publication bias and “file-drawer” problems in the field (Ferguson & Heene, 2012; Simmons et al., 2011). However, despite their importance, replications comprise only ~1% of studies (Makel et al., 2012).
A pivotal point in this discourse was the execution of various large replication projects, particularly the Many Labs Replication Project (Klein et al., 2014), the Pipeline Project (Schweinsberg et al., 2016), and the OSC Reproducibility Project. (OSC, 2015). In the Many Labs Replication Project (Klein et al., 2014), 36 groups of researchers worldwide independently tested 13 reputable psychological effects. They found that ten of the effects (77% of the studies) were consistently replicated across all 36 groups. The researchers explained the success of these results twofold: (1) they were studying effects that were proven to be reliable and stable, and (2) careful sample and methodological consistency across studies had contributed to successful replications. In the Pipeline Project (Schweinsberg et al., 2016), 10 studies on moral judgment (selected by researchers for lack of affect manipulation) were replicated before the original research was published. A total of 60% of studies replicated successfully across all laboratories, which was attributed to their straightforwardness and ease of administration. Finally, in the largest systematic replication effort to date, the OSC Reproducibility Project (OSC, 2015) replicated 100 studies from three major psychology journals, following a sampling process specified by the project coordinators to mitigate selection bias and match researchers’ expertise. All materials were obtained directly from the original researchers. Surprisingly, compared to 97% of original studies, only 36% of replications reported significant results.
Although the execution of replication projects is encouraging and suggests that the field is addressing this issue, there is no consensus on why the problem exists. There is a pressing need to understand which factors systematically contribute to the replication crisis. The present study proposes one such a priori factor: performance invalidity among undergraduate research participants. Performance invalidity, and its potential contribution to low replication rates, must first be understood in the context of who primarily comprises psychological research samples. The majority of psychological research is conducted with undergraduate students. Research indicates that an estimated 67%–86% of articles published in psychology journals involve undergraduate research participants (Arnett, 2008; Sears, 1986; Smart, 1966; Wintre et al., 2001). Henrich and coworkers (2010) demonstrated that it was 4,000 times more likely for an American undergraduate student to serve as a research participant than for a randomly selected individual outside this group.
Moreover, undergraduate students are not particularly incentivized to perform well or to participate in research experiments to the best of their ability. Firstly, undergraduate students most often participate in research studies to obtain non-contingent credit toward a university course (An et al., 2012). Their participation alone, not their performance, ensures that they receive course credit. In compliance with standard ethics protocols, participants are informed that even withdrawal from the study will not compromise their course credit, thereby communicating that credit is not contingent on performance. Secondly, undergraduate research participants are not likely to view the research as personally significant. Although many of these participants are first-year psychology students (An et al., 2012), there is no guarantee that they have a genuine interest in the research being conducted, which may lead to performance invalidity due to suboptimal effort (DeRight & Jorgensen, 2014).
In clinical and forensic settings, “performance invalidity” refers to behavior on cognitive tests that does not reflect an individual’s true ability. In this regard, Performance Validity Tests (PVTs) are objective cognitive measures designed to detect invalid performance by gauging the credibility of an individual’s score on a given test (Di Vico et al., 2023; McWhirter et al., 2020). When there is an external incentive to appear impaired on cognitive testing by failing PVTs, the term “malingering” is used to describe deliberate attempts to feign or exaggerate impairment for personal gain (Sherman et al., 2020). However, in the absence of such incentives, as is the case for healthy undergraduate research participants in non-clinical settings, PVT failures are more appropriately conceptualized as reflecting suboptimal effort and disengagement rather than deliberate deception. In this context, PVTs function as indicators of valid task engagement and effort (Green, 2007; McWhirter et al., 2020), with insufficient effort in research contexts significantly decreasing the accuracy of results (Green et al., 2001; Mittenberg et al., 2002). Reported base rates of invalid performance among undergraduate participants vary widely, ranging from 2% to 56% (An et al., 2012; DeRight & Jorgensen, 2014; Santos et al., 2014; Silk-Eglit et al., 2014). For instance, one study found that 55.6% of students failed at least one PVT at baseline and 30.8% at follow-up (An et al., 2012), whereas another reported more modest failure rates of 8.3% at baseline and 3.7% at 1-month retest (Ross et al., 2016). This variability challenges the assumption that credible responding occurs uniformly across non-clinical samples and suggests that performance invalidity can arise even in healthy student samples. As such, failure to assess performance validity introduces a risk of biased estimates and undermines confidence in the interpretation of experimental effects. Despite this concern, no prior work has directly examined whether screening for performance invalidity influences the outcomes of a direct replication study.
The present study
The present study is the first part of a larger replication project that aims to evaluate whether performance invalidity among undergraduate research participants influences study outcomes and, more broadly, contributes to challenges in the replicability of psychological research. Given that undergraduate students comprise a substantial proportion of research participants, and PVT failure is not negligible in this group, incorporating performance validity assessment may be necessary to ensure trustworthy data. As part of the larger project, a series of replications will be conducted, selected from the 100 studies in OSC’s Reproducibility Project based on their use of undergraduate participants and feasibility (i.e., time, funding, and the exclusion of studies requiring specialized equipment or cross-cultural samples).
Importantly, the use of undergraduate participants in the present study is not merely a matter of convenience, but a methodological choice. As a direct replication of Armor and coworkers (2008), which also recruited undergraduate participants, sample consistency was prioritized to maintain methodological fidelity; testing the same effect in a non-student sample would constitute a generalizability study rather than a replication. Moreover, undergraduate students represent the population most affected by the replication crisis, given their prevalence in the psychological literature. Examination of factors contributing to replication failures should therefore begin with the sample in which those failures most commonly occur.
PVTs were selected over Symptom Validity Tests (SVTs) because the construct of interest is effortful task engagement rather than symptom over-reporting (Sweet et al., 2021; Van Dyke et al., 2013). Specifically, nonverbal PVTs were chosen to minimize the risk that cultural or linguistic variables influenced failure rates in this demographically diverse sample, consistent with prior concerns in PVT research (An et al., 2012), and subsequent proposals that such factors may contribute to discrepant base rates in non-clinical samples (Ross et al., 2016).
The replicated study
The present study replicates Armor and coworkers (2008), who examined whether people preferred optimistically biased predictions rather than accurate ones. They found that participants were more likely to encourage optimistically biased predictions for vignette characters, t(124) = 10.36, prep > 0.99, p < .001, d = 0.93 (Armor et al., 2008), which countered the intuition that humans prefer accuracy in their predictions rather than undue optimism. This effect was subsequently successfully replicated in OSC’s Reproducibility Project, with an effect size larger than the original, t(175) = 15.64, p < .001, d = 1.18 (Lassetter et al., 2013). The prescribed optimism paradigm was selected because the effect is well-established, providing an ideal test case in which differences associated with performance validity status can be more confidently attributed to PVT screening rather than instability in the underlying effect.
Hypotheses
In this study, we examined whether performance invalidity among undergraduate participants influences the successful replication of a previously established psychological effect. We hypothesized that PVT performance would be associated with replication outcomes, such that participants who fail PVTs would differ on primary study variables compared to those who pass. We further predicted that excluding participants who demonstrate performance invalidity would meaningfully alter the observed effect. Finally, we hypothesized that a brief instructional prompt could reduce performance invalidity by enhancing task engagement during the experimental procedures.
Methods
Participants
Undergraduate students at a large Canadian university were recruited through an online research participation system in exchange for course credit. In addition to being a convenience sample, undergraduate students also constituted the population of interest. Individuals were pre-screened for English fluency, and only those who self-reported fluency were included in the sample. Of the 89 participants, one chose to withdraw their data, one was excluded due to experimenter error, and 87 participants were retained in the final sample.
Procedures
This study was approved by the institutional research ethics board. As with all 10 replications in the larger replication initiative, the present study’s methodology was consistent with the methodology of the original study, with the addition of (1) an experimental manipulation of task engagement, and (2) three commonly employed PVTs to assess performance validity.
Manipulation of task engagement
Upon arrival at the research laboratory, consenting participants were randomly assigned to one of three different conditions. In order to examine the effect of various instructions on the valid test engagement exhibited by undergraduate students, participants were randomized to one of three instructional conditions; two conditions received a brief prompt: (1) “You are contributing to the good of science if you do your best in the study” (n = 29), (2) “You will receive a gift card if you do your best in the study” (n = 29), and (3) one condition received no prompt (n = 29).
Replication of methods from 2008 study on prescribed optimism
For the replication portion, all materials and methods were obtained through OSF from the author’s original research on prescribed optimism. Consistent with the replication by Lassetter and coworkers (2013), the present study replicated only the between-subject “prescriptive” condition, which was the primary focus of the original study.
Armor and coworkers (2008) examined the value of optimism in personal predictions, challenging the assumption that accuracy is always the ideal standard. Through vignettes across four contexts, the study explored how people prescribed behaviors to vignette characters at varying levels of commitment, agency, and control. Results revealed that participants prescribed optimistic predictions over accuracy. Notably, these prescriptions were consistent across demographic groups and situational variables, highlighting the perceived utility of optimism.
Like in Armor and coworkers (2008), participants in the present study were further randomly assigned one of four different scenarios (i.e., award, investment, party, and surgery), which comprised eight vignettes presented to each participant in a counterbalanced manner. Each vignette had the same six questions, which were related to three distinct variables hypothesized to influence optimism, each of which was manipulated to produce the eight vignettes. The first variable was commitment, which refers to the degree to which a decision has been made (Armor & Taylor, 2003). The second variable was agency, which refers to the degree of power the individual of interest has in the outcome of the situation (Henry, 1994). The third variable was control, which refers to the degree of influence the individual of interest has over the situation (Klein & Helweg-Larsen, 2002).
The manipulation in the original study (and the subsequent replication) involved modifying the scenario details in each of the eight vignettes to produce varying levels of commitment, agency, and control. The first question asked participants what type of prediction the individual in the scenario should make; response options ranged from −4 (extremely pessimistic) to +4 (extremely optimistic). Other questions had response options ranging from 0 to 10 or 0% to 100%, depending on the question. Consistent with the original study, a Latin square was used to counterbalance scenarios and vignettes.
Following completion of the vignettes, participants were asked to complete the Life Orientation Test-Revised (LOT-R; Scheier et al., 1994). The LOT-R is a measure of dispositional optimism and has previously demonstrated strong predictive validity (Scheier et al., 1994). The LOT-R is a 10-item scale (six scored and four filler items) that is rated on a 5-point Likert scale, ranging from 0 (strongly disagree) to 5 (strongly agree). The order of the administered measures was consistent with that of the original replication.
Performance validity assessment
Participants were administered three PVTs to detect performance invalidity, following standardized procedures outlined test manuals. PVTs were counterbalanced across participants to account for order effects. In this study, because participants had no incentive to deliberately fail PVTs, failure on these measures was interpreted simply as reflecting “valid” and “invalid” performance, or “credible” and “non-credible” performance, rather than being operationalized as deception or malingering.
The Test of Memory Malingering (TOMM; Tombaugh, 1996). The TOMM is a 50-item recognition memory test that utilizes a forced-choice paradigm to detect performance invalidity. Participants are shown several pictures of objects and then asked to identify them in a series of two-choice recognition trials presented immediately afterward. This process is repeated in a second trial, and in some cases, a third trial is presented. All participants were administered Trial 1 and 2, with a cut-off score of ≤44 on Trial 2, as recommended by the manual, to classify failures. The TOMM has a substantial body of literature demonstrating its utility in discerning non-credible performance while being relatively insensitive to various neurological disorders and bonafide memory impairment (Lezak et al., 2012). Although several cutoff scores exist across populations, a cutoff score of ≤44 on Trial 2 is generally recommended in non-clinical samples to preserve specificity (Martin et al., 2020; Teichner & Wagner, 2004; Tombaugh, 1996).
The Dot-Counting Test (DCT; Boone et al., 2002). In the DCT, participants are shown a series of 12 cards with dots printed on them and instructed to count the number of dots as quickly as possible. Response time and accuracy (i.e., number of errors) are both utilized to calculate a composite score called the Effort Index Score (E-Score). Consistent with cross-validated recommendations for cognitively unimpaired samples (McCaul et al., 2018), an E-score cutoff of ≥14 was used to detect performance invalidity. A higher cutoff of ≥17, while cited in some literature, is validated for clinical populations with memory impairment (Hansen et al., 2023; McCaul et al., 2018) and would be less appropriate and lacking sensitivity for a healthy undergraduate sample.
The Reliable Digit Span Test (RDS; Greiffenstein et al., 1994). The RDS is a post-hoc measure used as an embedded PVT, derived from the Digit Span subtest of the Wechsler Adult Intelligence Scale (WAIS-IV; Wechsler, 2008). Participants are instructed to repeat strings of digits in forward and backward order. The RDS score is calculated by summing the longest successfully repeated forward and backward strings. RDS was first introduced by Greiffenstein and coworkers (1994), and research has since demonstrated its utility as an embedded PVT (Maiman et al., 2018). In healthy populations, a cutoff score of ≤7 is considered to reflect performance invalidity (Lezak et al., 2012; Spencer et al., 2013).
Other measures
Following the replicated method from Armor and coworkers (2008) and the administration of the PVTs, all participants completed a demographics questionnaire to assess sample characteristics, including participants’ sex, gender, race, educational status, marital status, and employment status, and a post-experimental questionnaire to assess their understanding of task instructions and the instructional prompt received. Following this, participants were debriefed on the study’s aims, compensated with a $10 gift card for their participation (regardless of their assigned prompt condition) and received their course credit.
Analysis plan
Data analyses were performed using the Statistical Package for Social Sciences (SPSS Version 28.0, IBM Corporation, 2021). Visualizations were created in R Version 4.5.3 (R Core Team, 2025) using the RStudio interface (Posit Team, 2024). An a priori power analysis indicated a minimum sample size of N = 84 (d = 0.40, α = 0.05, power = 0.95) (G*Power; Faul et al., 2007). Post-hoc power estimates are reported in the results. In accordance with the analyses performed by Armor and coworkers (2008), a t-test was conducted to compare the mean level of prescribed optimism with the midpoint of the scale. Similarly, t-tests were conducted to compare prescribed optimism levels across vignettes. PVT failure rates were computed based on manualized cutoffs. To assess the impact of task engagement on the detection of performance invalidity, the analyses were repeated in credible and non-credible subgroups. Chi-square tests were conducted to assess whether PVT failure rates differed by prompt received. An analysis of variance (ANOVA) was conducted to assess the effect of the scenario on prescribed optimism. Normality and homogeneity of variance were assessed using the Shapiro–Wilk test and Levene’s test with Welch’s correction applied as required. Parametric analyses were retained in all cases. Fisher’s exact test was reported for one chi-square test of independence, where one of four cell sizes had a count below 5.
Transparency and openness
Replicated materials from the OSF replication of the original study are publicly available at https://osf.io/ibhv7/. All measures, manipulations and exclusions for this study have been reported. Sample size was determined prior to data collection and analysis. Data is available by request directly from the authors. This study was not preregistered.
Results
The final sample consisted of 87 individuals (79.3% female, n = 69) with an average age of 18.5 years (SD = 2.21, range = 17 years). Of these individuals, 73.6% identified as female (n = 64), 21.8% (n = 19) identified as male, 1.1% (n = 1) identified as non-binary, and 3.4% (n = 3) preferred not to disclose their gender. 61.8% (n = 55) identified as Asian, 11.2% (n = 10) identified as Black, 7.9% (n = 7) identified as Caucasian, 4.5% (n = 4) self-identified as mixed-race, and 14.6% (n = 13) self-identified as another ethnicity. The demographic characteristics of the sample are summarized in Table 1.
| Characteristic . | n . | % . |
|---|---|---|
| Sex | ||
| Female | 69 | 79.3 |
| Male | 18 | 20.7 |
| Gender | ||
| Woman | 64 | 73.6 |
| Man | 19 | 21.8 |
| Undisclosed | 3 | 3.4 |
| Nonbinary | 1 | 1.1 |
| Ethnicitya | ||
| South Asian | 28 | 31.5 |
| East Asian | 18 | 20.2 |
| Black African/Caribbean | 10 | 11.2 |
| Southeast Asian | 9 | 10.1 |
| Other | 9 | 10.1 |
| White | 7 | 7.9 |
| Mixed | 4 | 4.5 |
| Middle Eastern | 4 | 4.5 |
| English as a First Language | ||
| No | 47 | 54.0 |
| Yes | 40 | 46.0 |
| First Languageb | ||
| English | 39 | 44.8 |
| Other | 19 | 21.8 |
| Urdu | 7 | 8.0 |
| Chinese | 6 | 6.9 |
| Mandarin | 5 | 5.7 |
| Tamil | 4 | 4.6 |
| Spanish | 3 | 3.4 |
| Bengali | 2 | 2.3 |
| Hindi | 2 | 2.3 |
| Employment Statusc | ||
| Student | 73 | 80.2 |
| Part-time employed | 13 | 14.3 |
| Unemployed | 3 | 3.3 |
| Full-time employed | 1 | 1.1 |
| Volunteer | 1 | 1.1 |
| Characteristic | n | % |
|---|---|---|
| Sex | ||
| Female | 69 | 79.3 |
| Male | 18 | 20.7 |
| Gender | ||
| Woman | 64 | 73.6 |
| Man | 19 | 21.8 |
| Undisclosed | 3 | 3.4 |
| Nonbinary | 1 | 1.1 |
| Ethnicity | ||
| South Asian | 28 | 31.5 |
| East Asian | 18 | 20.2 |
| Black African/Caribbean | 10 | 11.2 |
| Southeast Asian | 9 | 10.1 |
| Other | 9 | 10.1 |
| White | 7 | 7.9 |
| Mixed | 4 | 4.5 |
| Middle Eastern | 4 | 4.5 |
| English as a First Language | ||
| No | 47 | 54.0 |
| Yes | 40 | 46.0 |
| First Language | ||
| English | 39 | 44.8 |
| Other | 19 | 21.8 |
| Urdu | 7 | 8.0 |
| Chinese | 6 | 6.9 |
| Mandarin | 5 | 5.7 |
| Tamil | 4 | 4.6 |
| Spanish | 3 | 3.4 |
| Bengali | 2 | 2.3 |
| Hindi | 2 | 2.3 |
| Employment Status | ||
| Student | 73 | 80.2 |
| Part-time employed | 13 | 14.3 |
| Unemployed | 3 | 3.3 |
| Full-time employed | 1 | 1.1 |
| Volunteer | 1 | 1.1 |
Note. Total N = 87 unless otherwise indicated.
aOne participant selected two ethnicity categories: ethnicity total n = 89. “Black African/Caribbean combines ‘Black-African’ (n = 7) and ‘Black-Caribbean” (n = 3). ‘White’ combines ‘White European’ (n = 6) and ‘White-American’ (n = 1). ‘Other’ combines ‘Indo-Caribbean’ (n = 3), ‘Latin-American’ (n = 2), and ‘Other’(n = 4).
bOne participant reported both English and Bengali as first languages.
cTwo participants selected two employment categories: employment total n = 91.
The 87 participants were randomized to one of three prompts, with 29 (33.3%) participants in each group. Chi-squared tests were conducted to assess whether there were significant differences in demographic variables between the prompt-type groups following randomization. Between the three groups, there were no significant proportional differences in biological sex, X2 (2, N = 87) = 0.42, p = .81, or in English as a first language (EFL) speaker status, X2 (2, N = 87) = 2.87, p = .238. or ethnic group, X2 (22, N = 87) = 27.65, p = .187. This suggests a uniform distribution of demographics across prompt conditions. The average LOT-R score was 13.35 (SD = 3.84, range = 3–20; n = 86).
Order effects
A chi-square test of independence was conducted to evaluate whether the order of PVTs influenced PVT failure. The test revealed no significant differences in proportions, X2 (2, N = 87) = 0.42, p = .81, confirming that PVT order did not affect participants’ performance.
Effect of scenario on prescribed optimism
A one-way ANOVA was conducted to evaluate differences in prescribed optimism across the four vignette scenarios. No significant differences were observed, F(3, 83) = 0.56, p = .65.
PVT failure rates across the sample
All participants (n = 87) completed all three PVTs, administered in a counterbalanced order. For each measure, manualized cut-off scores were used to identify instances of PVT failure. 17.2% of participants (n = 15) failed at least one of the three PVTs they were administered. Across the sample, 16.1% of participants (n = 14) failed only one PVT, and 1.1% of participants (n = 1) failed two PVTs. No participants failed all three PVTs.
The mean score on Trial 1 of the TOMM was 47.92 (SD = 2.72, range = 38–50; n = 87), while the mean score on Trial 2 was 49.79 (SD = 1.61, range = 35–50). 1.1% of participants (n = 1) received a score below the cut-off on both Trial 1 and 2, resulting in a failure. The mean score obtained on Digit Span Forward was 6.45 (SD = 1.40, range = 3–9; n = 87), while the mean score obtained on Digit Span Backward was 5.22 (SD = 1.31, range = 0–8). The mean total score on the RDS was 13.36 (SD = 3.54, range = 4–24). 3.4% of participants (n = 3) received scores below the RDS cut-off, resulting in failure. The mean score obtained on the DCT was 11.18 (SD = 3.12, range = 6–23; n = 83). 14.5% of participants (n = 12) received scores below the E-score cut-off, resulting in failure. PVT failure rates are shown in Table 2.
| PVT . | n . | M . | SD . | Range . | % . |
|---|---|---|---|---|---|
| TOMM 1 | 87 | 47.92 | 2.72 | 12 | |
| TOMM 2 | 87 | 49.79 | 1.69 | 15 | |
| TOMM Fail | 1 | 1.1 | |||
| RDS Forward Span | 87 | 6.44 | 1.82 | 6 | |
| RDS Backwards Span | 87 | 5.21 | 1.69 | 8 | |
| RDS Total | 87 | 13.36 | 3.12 | 20 | |
| RDS Fail | 3 | 3.4 | |||
| DCT Ungrouped Mean | 83a | 6.63 | 1.82 | 11.13 | |
| DCT Grouped Mean | 83 | 3.22 | 1.69 | 12.11 | |
| DCT E-Score | 83 | 11.18 | 3.12 | 17 | |
| DCT Fail | 12 | 14.5 | |||
| Total PVT Fails | 15 | 17.2 |
| PVT | n | M | SD | Range | % |
|---|---|---|---|---|---|
| TOMM 1 | 87 | 47.92 | 2.72 | 12 | |
| TOMM 2 | 87 | 49.79 | 1.69 | 15 | |
| TOMM Fail | 1 | 1.1 | |||
| RDS Forward Span | 87 | 6.44 | 1.82 | 6 | |
| RDS Backwards Span | 87 | 5.21 | 1.69 | 8 | |
| RDS Total | 87 | 13.36 | 3.12 | 20 | |
| RDS Fail | 3 | 3.4 | |||
| DCT Ungrouped Mean | 83 | 6.63 | 1.82 | 11.13 | |
| DCT Grouped Mean | 83 | 3.22 | 1.69 | 12.11 | |
| DCT E-Score | 83 | 11.18 | 3.12 | 17 | |
| DCT Fail | 12 | 14.5 | |||
| Total PVT Fails | 15 | 17.2 |
aOf the 87 participants, four participants’ scores on the DCT could not be reliably calculated due to examiner error during administration.
Note. PVT = Performance Validity Test; TOMM = Test of Memory Malingering; RDS = Reliable Digit Span; DCT = Dot Counting Test. TOMM failure is based on Trial 2 only, consistent with manual scoring guidelines (Tombaugh, 1996); Trial 1 does not carry a formal failure designation.
Replication of Armor et al. (2008)
The present study aimed to replicate the findings of Armor and coworkers (2008), and the subsequent successful replication by Lassetter and coworkers (2013). Participant’s level of prescribed optimism, the primary outcome of interest, was assessed using a one-sample t-test that compared mean prescribed optimism to the midpoint of 0. Consistent with Armor and coworkers (2008) and Lassetter and coworkers (2013), participants prescribed significantly more optimistic predictions, t(86) = 9.14, p < .001, 95% CI [0.94, 1.46], d = 0.98, observed power = 1.00. The mean level of prescribed optimism was 1.20 (SD = 1.23), indicating that on average, participants’ responses were 1.2 points above the midpoint. The observed effect size was large (d = 0.98, 95% CI [0.72, 1.23]), exceeding the effect size reported in the original study by Armor and coworkers (2008) (d = 0.93) and smaller than that reported in the replication by Lassetter and coworkers (2013) (d = 1.18). Broadly, the present study’s findings replicate those of the original.
An examination of prescribed optimism for each of the eight vignette conditions individually further confirmed that participants tended to prescribe optimistic predictions over accurate ones. Mean levels of participants’ prescribed optimism across vignette-manipulated conditions of control, agency, and commitment are shown in Table 3. Participants prescribed the least optimistic predictions (M = 0.21) for conditions in which the character in the vignette had low control, less agency, and before they had committed to a given situation, t(86) = 0.88, p = 0.38, d = 0.10, 95% CI [−0.12, 0.31]. Conversely, participants prescribed the most optimistic predictions (M = 2.16) for conditions where the character in the vignette had high control, more agency and after they had committed to a given situation, t(86) = 10.8, p < .001, d = 1.16, 95% CI [0.88, 1.43]. Similarly, in other vignettes where the character had less control, the mean prescribed optimism was closer to 0 (i.e., accurate predictions).
Mean prescribed optimism across manipulations of commitment, agency, and control.
| Precommitment . | Post commitment . | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| External Agency . | Internal Agency . | External Agency . | Internal Agency . | ||||||
| Low control | High control | Low control | High control | Low control | High control | Low control | High control | Total | |
| Mean Prescribed Optimism | 0.21 | 1.43* | 0.56* | 1.52* | 0.74* | 1.85* | 1.13* | 2.16* | 1.20* |
| Standard Error of Mean | (0.23) | (0.19) | (0.21) | (0.20) | (0.23) | (0.18) | (0.23) | (0.20) | (0.13) |
| Cohen’s d | 0.10 | 0.80 | 0.29 | 0.81 | 0.35 | 1.12 | 0.53 | 1.16 | 0.98 |
| Precommitment | Post commitment | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| External Agency | Internal Agency | External Agency | Internal Agency | ||||||
| Low control | High control | Low control | High control | Low control | High control | Low control | High control | Total | |
| Mean Prescribed Optimism | 0.21 | 1.43* | 0.56* | 1.52* | 0.74* | 1.85* | 1.13* | 2.16* | 1.20* |
| Standard Error of Mean | (0.23) | (0.19) | (0.21) | (0.20) | (0.23) | (0.18) | (0.23) | (0.20) | (0.13) |
| Cohen’s d | 0.10 | 0.80 | 0.29 | 0.81 | 0.35 | 1.12 | 0.53 | 1.16 | 0.98 |
Note. Each cell reflects the mean level of prescribed optimism for the corresponding combination of three experimentally manipulated vignette variables: Commitment (pre- vs. post-commitment to a course of action), Agency (external vs. internal agency over the outcome), and Control (low vs. high control over the situation). Each mean was tested against the scale midpoint of 0 via a one-sample t-test. *p < .05.
An independent samples t-test was conducted to assess whether participants’ self-reported level of optimism (i.e., described optimism) was related to the level of optimism they prescribed for characters in the vignettes. Here, optimists were defined as participants with a LOT-R score greater than the midpoint of 2, while pessimists were defined as participants with a LOT-R score ˂2. While the entire sample skewed toward optimistic rather than pessimistic predictions for vignette characters, self-described optimistic participants prescribed significantly greater optimism to the characters (M = 1.38, SD = 1.10) than did pessimistic participants (M = 0.79, SD = 1.41), t(85) = −2.12, p = .04, 95% CI [−0.95, −0.03], d = −0.49, observed power = 0.55.
The role of performance invalidity in replication results
To evaluate whether the replication findings differed as a function of performance validity, primary and secondary analyses were repeated separately within credible and non-credible subgroups, with non-credible participants defined as those who failed at least one PVT. In the credible subgroup, prescribed optimism remained significantly above the midpoint, t(71) = 8.63, p < .001, 95% CI [1.00, 1.59], d = 1.02, observed power = 1.00. The mean level of prescribed optimism (M = 1.30, SD = 1.27, n = 72) was higher than that seen in the study’s entire sample (M = 1.20, SD = 1.23, n = 87), and the original study by Armor and coworkers (2008) (M = 1.12). In the non-credible subgroup, prescribed optimism was also greater than the midpoint, t(14) = 3.36, p = .005, 95% CI [0.27, 1.22], d = 0.87, observed power = 0.88, although the mean level of prescribed optimism was lower (M = 0.75, SD = 0.86, n = 15). Table 4 presents mean levels of prescribed optimism and corresponding effect sizes for the full sample and each PVT subgroup, each tested against the scale midpoint of 0 using one-sample t-tests. The credible subgroup yielded a numerically larger effect (d = 1.02) than the full sample (d = 0.98), whereas the non-credible subgroup showed a smaller effect (d = 0.85). Fig. 1 presents a forest plot of effect sizes with 95% confidence intervals across all replications, illustrating the consistency of the prescribed optimism effect across studies.
| Group . | N . | M . | SD . | t . | df . | p . | d . |
|---|---|---|---|---|---|---|---|
| Full sample | 87 | 1.20 | 1.23 | 9.14 | 86 | < .001 | 0.98 |
| PVT pass | 72 | 1.30 | 1.27 | 8.63 | 71 | < .001 | 1.02 |
| PVT fail | 15 | 0.75 | 0.86 | 3.36 | 14 | .005 | 0.87 |
| Group | N | M | SD | t | df | p | d |
|---|---|---|---|---|---|---|---|
| Full sample | 87 | 1.20 | 1.23 | 9.14 | 86 | < .001 | 0.98 |
| PVT pass | 72 | 1.30 | 1.27 | 8.63 | 71 | < .001 | 1.02 |
| PVT fail | 15 | 0.75 | 0.86 | 3.36 | 14 | .005 | 0.87 |
Note. Each row reflects an independent one-sample t-test comparing the group mean prescribed optimism to the scale midpoint of 0, consistent with the replicated study’s protocol and primary analysis. The full sample row reflects the primary replication analysis. The PVT pass and PVT fail rows reflect analyses stratifying by performance validity status. PVT = Performance Validity Test.

Plot of effect sizes (Cohen’s d) for primary analysis of prescribed optimism across studies.
To directly examine whether performance validity group membership was associated with differences in the primary outcome, an independent samples t-test compared prescribed optimism between the credible (M = 1.30, SD = 1.27, n = 72) and non-credible (M = 0.75, SD = 0.86, n = 15) subgroups. Credible participants prescribed significantly greater optimism than non-credible participants, t(28.61) = 2.055, p = .049, 95% CI [0.002, 1.097], d = 0.45, observed power = 0.35. Given unequal group sizes and heterogeneity in variance, Welch’s correction was applied.
To examine the role of performance invalidity in the secondary analysis, the relationship between participants’ described optimism and prescribed optimism was re-evaluated separately within each PVT subgroup using an independent-samples t-test. Within the credible subgroup, the findings mirrored those of the full sample; optimistic participants prescribed significantly greater optimism to vignette characters (M = 1.61, SD = 1.09) than pessimistic participants (M = 0.79, SD =1.40), t(70) = −2.78, p = .007, 95% CI [−1.41, −0.23], d = −0.67, observed power = 0.78. Conversely, in the non-credible subgroup, no significant difference was observed between optimistic participants (M = 0.63, SD = 0.93) and pessimistic participants (M = 1.08, SD = 0.57), t(13) = 0.89, p = .39, 95% CI [−0.64, 1.54], d = 0.52, observed power = 0.13. Interestingly, the direction of the mean difference was reversed in this subgroup, with pessimistic participants prescribing more optimism than optimistic ones.
Exploratory analysis using a conservative PVT failure threshold
As an exploratory analysis, and consistent with the commonly used clinical standard of ≥2 PVT failures to classify non-credible performance in neuropsychological evaluations (Sweet et al., 2021), the primary and secondary analyses were repeated after excluding the single participant who met this criterion, having failed the RDS and DCT. The resulting credible subgroup comprised 86 participants. The primary one-sample t-test evaluating prescribed optimism remained significant and virtually unchanged from the full-sample analysis, t(85) = 9.28, p < .001, 95% CI [0.96, 1.48], d = 0.99, observed power = 1.00. The secondary independent samples t-test comparing optimists (M = 1.45, SD = 1.10, n = 54) and pessimists (M = 0.83, SD = 1.32, n = 32) on prescribed optimism also remained significant, t(84) = −2.35, p = .021, 95% CI [−1.15, −0.09], d = −0.52. Notably, applying the ≥2 PVT failure threshold identified only one participant as non-credible, precluding meaningful subgroup analysis and underscoring the limited utility of forensic-derived thresholds in healthy student samples.
Effect of instructional prompts on performance invalidity
The effect of instructional prompt (i.e., type of instructions given to individuals at the beginning of the study) was analyzed using a chi-square test of independence. The analysis revealed no significant relationship between prompt type and PVT failure status, X2 (2, N = 87) = 0.48, p = .79, V = 0.08. Fig. 2 shows the proportion of participants passing and failing at least one PVT for each prompt condition, illustrating comparable pass/fail distributions across instructional prompts.

Proportion of participants who passed or failed Performance Validity Tests by prompt condition.
Demographic analyses
To assess whether demographic characteristics were associated with either the primary outcome or PVT failure, a series of independent samples t-tests and chi-square tests of independence were conducted.
Demographic variables and prescribed optimism
An independent samples t-test conducted to evaluate the effect of sex on prescribed optimism found no significant difference between males (M = 1.63, SD = 1.22) and females (M = 1.09, SD = 1.21), t(85) = 1.67, p = .099, 95% CI [−0.10, 1.17], d = 0.44. An independent samples t-test conducted to evaluate the effect of first language on prescribed optimism found no significant difference between those who spoke English as a first language (M = 1.19, SD = 1.08) and English as a second language (M = 1.20, SD = 1.35), t(85) = 0.02, p = .98, 95% CI [−0.52, 0.53], d = 0.004.
Demographic variables and PVT failure
Chi-square tests of independence were conducted to evaluate the relationship between PVT failure and demographic variables, including sex and English as a First Language (EFL) status. There were no significant differences in the proportions of males and females who failed a PVT, X2 (1, N = 87) = 0.01, Fisher’s exact p = 1.00, V = 0.01. There were no significant differences in proportions of EFL speakers and non-EFL speakers who failed a PVT, X2 (1, N = 87) = 0.26, p = .61, V = 0.06.
Discussion
The primary contribution of the present study is methodological: we examined whether incorporating performance validity screening into a psychological research protocol, using tools developed and validated in clinical neuropsychology, provides a practical mechanism to improve data quality and support accurate interpretation of research findings. A direct replication of Armor and coworkers (2008) prescribed optimism study served as the test scenario, providing a well-established effect against which the influence of performance invalidity could be evaluated, such that observed differences could be more confidently attributed to PVT screening procedure itself rather than instability in the underlying effect.
Replication of the prescribed optimism effect
The prescribed optimism effect was successfully replicated, with our sample’s results mirroring those of the original study. Participants prescribed significantly optimistic predictions rather than pessimistic or accurate predictions for the vignette characters, yielding a main effect size (d = 0.98) that was largely consistent with the original study’s (d = 0.93; Armor et al. (2008) and the subsequent replication (d = 1.18; Lassetter et al., 2013).
The pattern of effects, in which situational factors influenced these predictions, with greater optimism prescribed in conditions where the character had already committed to an action and had greater control and agency over that action, was also replicated. Finally, participants who described themselves more optimistically also prescribed greater optimism to others, consistent with prior findings. This successful replication established the validity of the present study’s paradigm and provided an empirical foundation for the subsequent performance validity analyses.
Rate of performance invalidity in undergraduate student samples
A total of 17.2% of participants failed at least one PVT, consistent with the range of base rates reported in prior undergraduate PVT research (An et al., 2012; DeRight & Jorgensen, 2014; Ross et al., 2016; Silk-Eglit et al., 2014). This rate is not negligible, and its presence in a sample that would conventionally be treated as uniformly credible underscores the value of systematic screening for performance validity. Another noteworthy finding was the differential failure rates observed across the three PVTs administered. Failure rates of the DCT (14.5%), RDS (3.4%), and TOMM (1.1%) likely reflect fundamental differences in the cognitive constructs and psychometric properties of each measure, a pattern consistent with published comparisons in clinical samples (Dandachi-FitzGerald et al., 2020; Martin et al., 2020; Teichner & Wagner, 2004). The TOMM is a forced-choice recognition memory paradigm designed to produce near-ceiling scores in cognitively intact individuals, for whom performance on Trial 2 is characteristically near-perfect (Martin et al., 2020; Teichner & Wagner, 2004). Accordingly, a 1.1% failure rate in our healthy student sample is consistent with established expectations. The RDS, as an embedded PVT, also operates near the ceiling in healthy young adults; notably, research has shown that standard RDS cutoffs can yield high false-positive rates in culturally and linguistically diverse, non-clinical samples (Gasquoine et al., 2017). The DCT’s composite E-score indexes both speed and accuracy, making it uniquely sensitive to individual variability in processing speed in ways that recognition memory measures are not. Indeed, when both are administered in the same sample, DCT failure rates exceed those of the TOMM (Dandachi-FitzGerald et al., 2020): 23.8% vs. 13.7%, mirroring our findings.
Taken together, these findings underscore the value of administering multiple heterogeneous PVTs: measures with differing sensitivity-specificity profiles and distinct cognitive demands collectively provide a more complete picture of engagement than any single instrument can achieve. While much of the literature comparing differential PVT failure rates has been conducted in clinical samples, our findings demonstrate that the same pattern of heterogeneous failure rates, with higher DCT failure rates, is also observable in cognitively healthy, non-clinical samples.
Influence of performance invalidity on research outcomes
The core question of this study was whether performance invalidity status was associated with meaningfully different responding on the primary outcome, thereby influencing the extent of the study’s replication. Accordingly, we hypothesized that removing participants who exhibited non-credible performance would affect the findings. Evidence for this was observed across multiple analyses.
Analysis of only the credible subgroup yielded a numerically larger effect size (d = 1.02) than the full sample (d = 0.98), while the non-credible subgroup yielded a smaller effect (d = 0.87). Although confidence intervals overlapped, and both estimates were large (Cohen, 1988), the pattern is informative: the primary effect remained robust regardless of the inclusion of invalid performers. Notably, excluding non-credible participants did not reduce the effect; rather, it was directionally larger, consistent with the hypothesis that performance invalidity introduces noise that can attenuate observed effects. Importantly, the present study examined a robust, large effect, for which this degree of attenuation did not meaningfully alter the overall outcome. However, in studies examining smaller, yet still theoretically or clinically meaningful, effects, comparable levels of noise may meaningfully influence effect size estimates, statistical significance, and ultimately the interpretation of findings. This pattern warrants further examination across a larger series of replications.
More compelling evidence comes from secondary analyses. When evaluating the relationship between described and prescribed optimism, the relationship was significant and directionally consistent within the credible subgroup (d = −0.67). Conversely, in the non-credible subgroup, the direction of the effect was reversed entirely; pessimistic participants prescribed more optimism than optimistic participants – a finding inconsistent with prior literature and suggestive of the kind of systematic noise that disengaged responding may introduce into research data.
Further support for the influence of performance invalidity on study outcomes comes from a direct comparison of the credible and non-credible subgroups on the primary outcome of prescribed optimism. These groups differed significantly in a theoretically predicted direction (p = .049, d = 0.45), despite being statistically equivalent on several tested demographic variables and the experimental condition. The emergence of a significant difference between otherwise equivalent groups provides further evidence that performance invalidity may be associated with differential responding on the primary outcome. Participants who demonstrated invalid performance prescribed systematically less optimism than those who performed credibly, again consistent with the notion that disengaged responding may attenuate observed effects. Taken together, the consistent directional pattern suggests that performance invalidity in undergraduate student research samples does not merely introduce random noise but may introduce systematic, theoretically predictable bias that could obscure or distort effects.
These findings carry a practical implication: PVTs represent a low-cost, standardized addition to psychological research protocols that can meaningfully improve data quality without requiring fundamental or high-barrier changes to existing paradigms. The present study demonstrates that incorporating PVTs can allow researchers to identify disengaged participants, characterize sample validity, stratify analyses by validity, and report findings with and without invalid performance – ultimately providing a more complete and transparent account of study outcomes.
The exploratory analysis applying the ≥2 PVT failure threshold, the most widely endorsed clinical criterion for non-credible performance (Sweet et al., 2021), further illustrates that population-appropriate criteria are essential. This threshold identified only one participant as non-credible, precluding meaningful subgroup comparisons and rendering it ineffectual as a screening tool in this context. This reflects a fundamental mismatch between criteria developed for forensic and clinical populations, in which external incentives to feign impairment are present, and healthy student samples, where the same incentives do not apply. The ≥1 threshold adopted in the present study is consistent with practice in prior undergraduate PVT research (An et al., 2012; DeRight & Jorgensen, 2015; Ross et al., 2016) and was deemed a population-appropriate choice. Notably, results under the stringent threshold were virtually identical to those of the full sample, further suggesting that even a single PVT failure carries a meaningful signal in this population.
Beyond its contribution to indexing engagement and validity in research with undergraduate student samples, the present study has implications for applied clinical and forensic neuropsychological practice. The observed 17.2% PVT failure rate in a healthy undergraduate sample provides a normative reference point for a population that is underrepresented in the base-rate literature. Because the positive predictive value of any PVT failure depends on base rates of invalidity in the population being assessed, population-specific data are essential for accurate clinical interpretation (Gaudet et al., 2022; Sweet et al., 2021). In lower base-rate populations, such as general clinical referrals or dementia evaluations, a single PVT failure carries a substantially elevated false-positive risk (Gaudet et al., 2022), underscoring that forensic-derived thresholds cannot be uncritically applied across settings. Our findings suggest that the 17.2% failure rate observed in healthy undergraduate student participants with no clinical incentive to feign impairment represents a meaningful signal that, if undetected, could bias clinical conclusions in settings where equivalent levels of participant engagement cannot be guaranteed. These findings reinforce the AACN’s call for universal validity assessment and highlight that population-appropriate base-rate data are a clinical and forensic necessity, not merely a methodological consideration.
Task engagement manipulation using instructional prompts
The present study also assessed whether valid task engagement could be experimentally manipulated. We hypothesized that participants would demonstrate greater performance validity when provided with either a monetary incentive prompt or an encouragement to perform well. No significant relationship was observed between prompt type and PVT failure status, suggesting that brief instructional prompts may not be sufficient to meaningfully alter performance validity in this population. While this null result is informative, it suggests that performance invalidity in undergraduate research is unlikely to be mitigated by brief or superficial manipulations and instead may require more substantive structural interventions, such as performance-contingent incentive structures or ecologically meaningful task designs, to reliably promote sustained engagement.
Limitations and future directions
While the present study has provided promising initial findings regarding the implementation of PVTs in psychological research, some limitations must be acknowledged. First, one cannot definitively conclude that PVT failure indicates invalid performance across all study components. Participants may have been validly engaged with the vignettes but not with the PVTs, although this is unlikely given evidence that PVT failure generalizes across cognitive tasks. Second, the non-credible subgroup was small (n = 15), and analyses involving this group were underpowered. These findings should be interpreted with appropriate caution and replicated in larger samples. Third, the study was conducted in a single Canadian university with a demographically specific undergraduate sample, limiting generalizability to other populations and research contexts.
Future research should explore whether performance validity plays a similar role in the reproducibility of study outcomes across other paradigms, populations, and research designs, particularly those involving more cognitively demanding tasks, where disengagement may have larger effects on study outcomes. While performance validity was not successfully manipulated with the chosen instructional prompts, future research might also consider integrating PVTs with task-intrinsic comprehension checks, commonly employed in behavioral economics (Oppenheimer et al., 2009), to simultaneously assess both generalized effort and instruction-following compliance. Similarly, performance-contingent incentive structures, in which monetary rewards scale with individual task performance rather than being distributed uniformly, may more effectively promote effortful engagement than the prompt manipulation employed here. Finally, incorporating (SVTs) to assess exaggerated symptom reporting and constructs such as positive impression management and social desirability alongside PVTs, will allow for examination of whether response bias contributes independently to variability in research outcomes – a question distinct from, but complementary to, the effort-based framework adopted here. Together, these approaches represent promising avenues for reducing performance invalidity in research contexts where PVT-failure-based exclusion is not feasible.
Conclusion
Our replication of Armor and coworkers (2008) study provides preliminary support for the influence of performance validity on the replicability of research. Our findings demonstrate that performance invalidity occurs at a non-trivial rate in undergraduate student research participants and may introduce systematic, directionally predictable noise into research data. Neuropsychological tools developed for clinical practice, PVTs, can be adapted for the research context as a practical, low-cost safeguard that improves data quality and supports more accurate interpretation of findings.
By incorporating PVTs into a direct replication of a well-established psychological effect, we show that performance validity screening could allow researchers to characterize, stratify, and report their data with greater precision and transparency. The methodological contribution of this work extends beyond its specific findings, positioning performance credibility as a measurable and addressable variable in research design and demonstrating that neuropsychology offers practical solutions to address psychological science’s reproducibility challenges. As part of a larger replication initiative, these findings will inform subsequent replications and contribute to an evidence base for when and how performance validity screening meaningfully influences research outcomes.
Funding
This work was supported by an insight Grant from the Social Sciences and Humanities Research Council (SSHRC) (Grant Number: 1037949).
Conflict of Interest
The authors have no known conflict of interest to disclose. Materials from the OSF replication are publicly available at https://osf.io/ibhv7/.
Author Contributions
Melanie Mascarenhas (Data curation, Formal analysis, Investigation, Methodology, Project administration, Supervision, Visualization, Writing—original draft) and Konstantine Zakzanis (Conceptualization, Funding acquisition, Investigation, Resources, Supervision, Writing—review & editing)
References
Maiman, M., Del Bene, V. A., MacAllister, W. S., Sheldon, S., Farrell, E., Arce Rentería, M., Slugh, M., Nadkarni, S. S., & Barr, W. B. (2018). Reliable Digit Span : Does it Adequately Measure Suboptimal Effort in an Adult Epilepsy Population?