Abstract

Objective

In this study we examined the temporal stability of the Immediate Post-Concussion Assessment and Cognitive Test (ImPACT) within NCAA Division I athletes across various timepoints using an exhaustive series of statistical models.

Methods

Within a cohort design, 48 athletes completed repeated baseline ImPACT assessments at various timepoints. Intraclass correlation coefficients (ICC) were calculated using a two-way mixed effects model with absolute agreement.

Results

Four ImPACT composite scores (Verbal Memory, Visual Memory, Visual Motor Speed, and Reaction Time) demonstrated moderate reliability (ICC = 0.51–0.66) across the span of a typical Division I athlete’s career, which is below previous reliability recommendations (0.90) for measures used in individual decision-making. No evidence of fixed bias was detected within Verbal Memory, Visual Motor Speed, or Reaction Time composite scores, and minimal detectable change values exceeded the limits of agreement.

Conclusions

The demonstrated temporal stability of the ImPACT falls below the published recommendations, and as such, fails to provide robust support for the NCAA’s recommendation to obtain a single preparticipation cognitive baseline for use in sports-related concussion management throughout an athlete’s career. Clinical interpretation guidelines are provided for clinicians who utilize baseline ImPACT scores for later performance comparisons.

Introduction

Research suggests that the concussion rate within the general population is approximately 600 cases per 100,000 people (Cassidy et al., 2004), with an estimated 1.6–3.8 million sports-related concussions (SRC) reported each year in the USA (Harmon et al., 2013; Langolis, Rutland-Brown, & Wald, 2006). SRCs are estimated to account for more than 6% of all sports-related injuries (Covassin, Swanik, & Sachs, 2003). Unfortunately, these numbers likely underrepresent the true prevalence of SRC because many athletes underreport postconcussive syndrome (PCS) (Echlin et al., 2010; Meier et al., 2015). Recent research estimates that up to 50% of SRCs go unreported (Harmon et al., 2013).

Research on the typical neurocognitive sequelae of concussions suggests that there is little uniformity in the presentation and course of PCS (Karr, Areshenkiff, & Garcia-Barrera, 2014). In their review of meta-analyses on the cognitive effects of concussion, Karr, Areshenkiff, and Garcia-Barrera (2014) found the influence of concussion on cognition to be highly variable, ranging from negligible to substantial across all cognitive domains (Cohen, 1988). They found that postconcussive cognitive sequelae often include decreased attentional abilities, executive dysfunction, reduced memory, visuospatial dysfunction, or decreased verbal ability. In a 2005 meta-analysis of only SRC research, Belanger and Vanderploeg (2005) analyzed 21 studies, with a total sample of 790 concussed athletes and 2,014 controls. They reported that, within the acute stage (i.e., within 24 hr of injury), concussion had a large effect on delayed memory, memory acquisition, and global cognitive abilities (see Table 1). Despite the heterogeneity in PCS, return to baseline typically ranges from 7 days (Belanger & Vanderploeg, 2005; McCrory et al., 2013) to 90 days (Binder, Rohling, & Larrabee, 1997) after first-time injuries in otherwise healthy adults.

Table 1

Postconcussion cognitive sequelae effect sizes

Reported effect size (Cohen’s d)
Cognitive domainKarr, Areshenkiff and Garcia-Barrera (2014)Belanger & Vanderpleog (2005)
Global cognitive abilities1.42
Attentional abilities0.05–0.63
Executive dysfunction−0.08 to –0.54
Memory0.13–0.781.03
Delayed memory0.13–0.691.00
Visuospatial dysfunction−0.25 to –0.57
Language skills−0.09 to –0.57

The accurate assessment of PCS via neurocognitive and/or physical measures (e.g., King-Devick Test, Vestibular/Ocular-Motor Screening) is crucial to athlete safety and is often at the center of return-to-play decisions for healthcare providers and sports medicine staff within professional and collegiate sports (Harmon et al., 2013). Concussion results in a cascade of metabolic changes, temporarily increasing the brain’s vulnerability to future insults during this acute recovery phase (Barkhoudarian, Hova, & Giza, 2016; MacFarlane & Glenn, 2015). Excessive cognitive and physical activity before the recovery of concussive injury is associated with prolonged vulnerability, often increasing the severity of and delaying the resolution of PCS (Harmon et al., 2013). Furthermore, athletes who sustain a second concussion shortly after the initial injury are vulnerable to a worsening of postconcussive neurological and metabolic dysfunction and can experience significantly poorer outcomes (Prins et al., 2010).

Assessment of SRC in Collegiate Athletes

Management of postconcussive cognitive sequelae relies heavily on the comparison of postconcussion test results to baseline data (Echemendia, Meeuwisse, Comper, Aubry, & Hutchison, 2016; Harmon et al., 2013; NCAA Sports Science Institute, 2017; Schatz, 2011). Though some research suggests that the comparison of postinjury data to individual baselines increases precision in diagnostic accuracy with specialty populations for whom current normative data may be inappropriate (e.g., above average intellect, those with learning disabilities or attention-deficit/hyperactivity disorder; Louey et al., 2014; Schatz & Robertshaw, 2014; Elbin et al., 2013), researchers have questioned the necessity of baseline assessment in collegiate sports concussion measurement with an emphasis instead on normative scores. Several studies have suggested that the use of normative data in postinjury assessment is equally as effective in the detection of postconcussive sequelae with both computerized cognitive screening measures (Echemendia et al., 2012; Schmidt et al., 2012) and measures of physical postconcussive sequelae (e.g., Balance Error Scoring System [BESS], Graded Symptom Checklist; Rebchuk et al., 2020). Other researchers have highlighted flaws with the individual baseline comparison paradigm in concussion management, arguing that baseline assessment does not adequately reduce risk associated with repeated concussion or severe/long-term postconcussive consequences (Randolph, 2011). In fact, the most recent consensus statement released by the 2017 Concussion in Sport Group does not require baseline neuropsychological assessment for all athletes (McCrory et al., 2018).

Findings that suggest that preparticipation baselines are unnecessary are particularly compelling given the inherent limitations with the baseline—postconcussive comparison paradigm. Practitioners generally attribute differences in test performance between those two administrations to the concussive injury (NCAA Sports Science Institute, 2017), though several factors, including practice effects (Bartels et al., 2010), inconsistencies in testing environment (Echemendia, Herring, & Bailes, 2009), athlete sandbagging efforts at baseline (Bailey, Echemendia, & Arnett, 2006; Schatz et al., 2017), and the psychometric properties of cognitive measures themselves (Lezak, Howieson, Bigler, & Tranel, 2012), can confound these changes.

Failure to appreciate all of the factors that contribute to performance differences across assessments may lead healthcare providers to prematurely assume recovery postinjury, granting dangerously early return-to-play and increasing an athlete’s vulnerability to poor outcomes or, alternately to prolonged removal from play in the case of a false positive result. As such, an understanding of the test–retest reliability of measures used in SRC is particularly crucial for effective concussion management when clinicians utilize baseline tests in the tracking of PCS.

Despite the controversy, the utilization of baseline data for postinjury comparison is a relatively common practice throughout collegiate sport programs. The NCAA Sports Science Institute released the Interassociation Consensus: Diagnosis and Management of Sports Related Concussion Best Practices, assisting participating universities in developing individual concussion protocols (NCAA Sports Science Institute, 2017). These guidelines suggest that each athlete undergo a “one-time, pre-participation baseline,” (NCAA Sports Science Institute, 2017, p. 8) consisting of baseline physical and cognitive assessment. Measures of physical indicators of postconcussive symptoms that are often utilized by collegiate athletic programs include the King-Devick Test (a measure of ocular-motor function) and the BESS (a balance and postural stability assessment; Kerr et al., 2015). If at any point during an athlete’s collegiate career, they sustain a concussive injury, they present to sports medicine for reassessment. Those results are compared to the original baseline data to inform return-to-play decisions, regardless of the time elapsed since the baseline assessment (NCAA Sports Science Institute, 2017), and clinicians attribute differences in performance between baseline and postinjury assessments to the concussive injury. Importantly, the interval between evaluations can be as long as 5 years, as NCAA eligibility rules allow collegiate athletes up to 5 years to complete four seasons of competition (NCAA, 2015, “Eligibility Timeline”). Presently, there are no clear recommendations for periodic baseline retesting or an indication that scores are no longer valid after a certain time.

The Immediate Post-Concussion Assessment and Cognitive Test

The development of computerized neurocognitive screening measures democratizes assessment in many ways and has allowed healthcare professionals from multiple disciplines to become involved in the evaluation of concussion and return-to-play decision making (De Marco & Broshek, 2016; Harmon et al., 2013). De Marco and Broshek (2016) reviewed the strengths and weaknesses of computerized cognitive assessment within the context of SRC. Among the many benefits to computerized assessment, the authors noted that computerized testing increases the ease with which sports medicine providers can assess for and monitor postconcussive sequelae, which may increase the likelihood that concussed athletes are assessed, identified, and receive appropriate care.

The Immediate Post-Concussion Assessment and Cognitive Test (ImPACT; Lovell et al., 2000) is one of the most widely used computerized neurocognitive assessment batteries in North America, and more than 77% of Division I athletic programs report using it in the management of SRC (Echemendia, et al., 2016; Kerr et al., 2015). The ImPACT is divided into six testing modules to examine five areas of cognitive function, including visual memory, verbal memory, processing speed, reaction time, and impulse control. The ImPACT can be administered individually or in a group setting and requires about 25 min to complete the six modules. After administration, the battery produces five composite scores: Visual-Motor Speed, Reaction Time, Verbal Memory, Visual Memory, and Impulse Control. In addition to the cognitive data, the battery includes a Symptom Inventory Scale to assess 22 self-reported postconcussive symptoms, and physical and emotional complaints, on a 7-point Likert scale. The psychometric properties of ImPACT have been previously examined with research supporting ImPACT’s concurrent and construct validity (Iverson, Lovell, & Collins, 2005; Maerlender et al., 2010). The reported sensitivity and specificity range from 82% to 94.6% and 69.1% to 97.3%, respectively, and internal consistency reliability (Cronbach’s α) estimates range from 0.82 to 0.84 (Higgins, Caze, & Maerlender, 2018; Lovell, 2016).

Despite the widespread use of ImPACT and the reliance on score changes for the detection and management of concussion and its sequelae, investigations of the temporal stability of ImPACT has yielded remarkably inconsistent results (Table 2). Researchers frequently measure temporal stability and test–retest reliability via intraclass correlation coefficients (ICC), a measure of the association between two or more statistics (Koo & Li, 2016; Vaz et al., 2013). ICCs equal or above 0.90 suggest excellent reliability, 0.75–0.90 suggest good reliability, 0.50–0.75 suggest moderate reliability, and values below 0.50 suggest poor reliability (Koo & Li, 2016).

Table 2

Summary of ImPACT ICC findings

ImPACT composite score temporal stability
CitationPopulationnAge mean (SD)Time intervalType of ICC analysisVerbal MemoryVisual MemoryVisual Motor SpeedReaction TimeImpulse Control
Brett and colleagues (2016)High school athletes25015 (1.85)b1 yearTwo-way, mixed-effects consistency0.460.670.870.67
11462 year0.530.660.830.59
1143 year0.360.630.900.70
Broglio and colleagues (2007)College students7321 (2.78)45 daysTwo-way, random effectse0.230.320.380.390.15
7321 (2.78)5 days (Day 45–50)0.400.390.610.510.54
Echemendia and colleagues (2016)Professional NHL players11925 (4.61)1 yearTwo-way, random effects, absolute agreement0.38  a0.54  a0.74a0.52a
18720.95 (3.05)2 years0.440.550.660.36
11821 (3.84)3 years0.350.460.570.54
11825 (4.38)4 years0.290.420.690.39
Maerlender and colleagues (2016)Collegiate athletes23119dVariableTwo-way, mixed effects, absolute agreement0.740.780.890.77
Schatz (2010)Nonfootball collegiate varsity athletes9518 (0.6)2 yearsTwo-way, mixed-effects consistency0.500.650.720.76
Schatz and Ferris (2013)Nonvarsity athletes undergraduate students25Age not reported1 monthTwo-way, mixed-effects consistency0.790.600.880.77
Tsuchima and colleagues (2016)High school athletes21215.3 (0.5)2 yearsTwo-way, random effectse0.210.490.720.46

aPearson’s correlation.

bMean/SD age reported for all participants and not broken up by interval.

cICCs calculated from the average of four test periods over a 4-year interval.

dNo SD reported.

eICC = intraclass correlation coefficients; ImPACT = Immediate Post-Concussion Assessment and Cognitive Test. ICC “definition” not reported.

Short-term (<1 year) reliability studies have yielded inconsistent results. Schatz and Ferris (2013) examined the 1-month test–retest reliability of ImPACT composite scores among a group of 25 undergraduates with no reported history of concussion. Interclass correlation coefficients indicated moderate to good reliability (see Table 2). Broglio and colleagues (2007) examined the temporal stability of ImPACT composite scores over two periods, baseline to 45 days and 45–50 days, among 73 healthy undergraduate students. Results generally indicated poor test–retest reliability (see Table 2). By focusing on shorter test–retest latency periods, these studies minimized any natural variability that may increase with time, rendering the utility and generalizability of these findings to be minimal.

Studies of ImPACT reliability across 2-year retest period have continued to yield relatively inconsistent results. Using a sample of 95 varsity collegiate athletes, Schatz (2010) reported ICC coefficients ranging from poor to moderate over 2 years (see Table 2), with ICCs for the Processing Speed Composite, Reaction Time Composite, and Visual Memory Composite suggesting at least marginal stability. In a study of NCAA Division 1 collegiate athletes, Maerlender and colleagues (2016) utilized a different statistical approach and averaged scores across four periods over 2 years. They found consistently strong intraclass correlations across four composite scores (see Table 2), suggesting good test–retest reliability for those composite scores. Tsuchima et al. (2016) reported ICC ranging from poor to moderate (Table 2) across a 2-year test–retest interval in high-school athletes. Brett and colleagues (2016) examined ImPACT reliability across multiple time points (1-, 2-, and 3-year intervals) among high-school athletes, and the ICCs were slightly improved and ranged from poor to good (see Table 2).

The research on ImPACT reliability extending beyond 3 years is extremely limited. Echemendia and colleagues (2016) did examine the reliability of ImPACT composite scores over longer periods of time in professional National Hockey League players. ICCs at 4 years between baseline administrations were reported to range from poor to moderate (Table 2), with Visual Motor Speed (ICC = 0.69) being the only composite score reaching acceptable levels of reliability. No research has examined the reliability of ImPACT across more than 3 years in collegiate athletes.

To conclude, the test–retest reliability for all ImPACT composite scores has ranged from poor to good with somewhat inconsistent results across a large range of time periods, suggesting the literature has not fully substantiated the temporal stability of this measure.

In addition to differences in test–retest procedure, testing environments, and sample characteristics, a notable explanation for variability in ICCs across studies may be differences in the approaches to calculating ICCs. ICCs can be calculated in a variety of ways, with each unique method resulting in different statistical outcomes when applied to the same sample (Koo & Li, 2016). In their recently published review of ICCs, Koo and Li (2016) discussed the variety of possible ICC analyses and examined their utility across different research questions. Noting that different ICC analytic approaches results in different values when applied to the same data, the authors indicate that using a two-way mixed effects model with absolute agreement should always be utilized with analyses of test–retest reliability (Koo & Li, 2016).

Within the literature summarized here, only one study (Maerlender et al., 2016) utilized an appropriate approach to test–retest ICC analysis. Three studies utilized a two-way, mixed effects, consistency ICC (Brett et al., 2016; Schatz, 2010; Schatz and Ferris, 2013), one study utilized a two-way, random effects, absolute agreement model (Echemendia et al., 2016), and two studies employed a two-way random effects model without indicating specific ICC definitions (Broglio et al., 2007; Tsuchima et al., 2016), and all these approaches have been deemed inappropriate for temporal stability analysis (Koo & Li, 2016).

To summarize, more research is needed examining temporal stability that utilized a two-way, mixed effects model with absolute agreement ICC (Koo & Li, 2016). Additionally, there are no studies that examine the utility of older retrospective ImPACT baselines, which can contribute to external validity by correctly accounting for the variability in time at which the postconcussion assessments are administered (Carlson & Morrison, 2009). The present study was designed to investigate the temporal stability of ImPACT in a cohort of male and female NCAA Division I athletes from a variety of sports. The recorded scores span the length of a collegiate athletic career, and test–retest intervals vary as is common practice in Division I programs. Because previous studies have suggested that Reaction Time and Visual Motor Speed Composite scores are at least moderately reliable across multiple time periods (i.e., ICC > 0.50), it is expected that these composite scores will generate ICC above 0.50 in the present study, reflecting moderate or good reliability. And, given the wider range in previously reported test–retest reliability coefficients for the Visual Memory and Verbal Memory scores, it is expected that Verbal Memory Composite, Visual Memory Composite, and Impulse Control Composite scores will be associated with poor reliability (ICC < 0.50). The aims of the current study were to more accurately characterize the temporal stability of ImPACT and assess clinical implications of its psychometric properties.

Methods

Participants

This cohort included 48 Division I athletes (Jager, Putnick, & Bornstein, 2017). Though small, other authors have provided meaningful insight into temporal stability using similarly small samples (Fukata et al., 2019; White et al., 2019). In addition to baseline ImPACT administered for the present study, athletes included in this analysis had undergone prior baseline ImPACT testing per usual NCAA practice and thus represented a unique subset of active Division I athletes with repeated ImPACT baseline scores at different intervals of elapsed time. Exclusion criteria included self-reported history of attention deficit/hyperactivity disorder (ADHD; Elbin et al., 2013) and/or concussion/head injury within 1 year prior to initial baseline or between ImPACT administrations, or invalid ImPACT performance, defined as Impulse Control composite score above 30 (Brett & Soloman, 2017). Of the 50 originally identified participants, two were excluded due to self-reported ADHD diagnoses. No participant sustained a concussion between baselines, and all ImPACT performances were valid.

The average age of male athletes was 19.94 years (SD = 1.44; n = 30), and the average age of female athletes was 19.62 years (SD = 1.52; n = 17). Participants predominantly self-identified as Caucasian (n = 38, > 81%). Average retest interval was 42.3 months (median = 33.9; SD = 30.8, [33.4, 51.3]; range = 2 days–3.99 years). Approximately 48% were retested within 12.5 months. About 71% within 25.0 months, 94% within 37.5 months (Fig. 1). Two-thirds of athletes participated in soccer (39.6%) and lacrosse (27.1%) with small numbers of athletes in hockey, basketball, volleyball, skiing, gymnastics, diving, and swimming.

Frequencies of latency periods between repeated baseline administration.
Fig. 1

Frequencies of latency periods between repeated baseline administration.

Materials and Procedure

Data were collected as part of a larger study examining biomarker and balance changes associated with SRC. The cohort comprised collegiate athletes from an NCAA Division I athletic program and included several high-impact sports except for football. The sample used in the current study was a subgroup of participants from the more than 400-person cohort that is currently under study. In addition to baseline data collected during the larger study, this subgroup of participants also had baseline ImPACT data collected as part of the university’s standard Sports Medicine preparticipation assessment of all incoming athletes.

All participants completed a standardized intake and demographic questionnaire to assess history of concussion. At admission into the study, these athletes were administered the ImPACT (Lovell et al., 2000), the BESS (Guskiewicz, Ross, & Marshall, 2001), and the King-Devick test (Galetta et al., 2011), which are commonly administered together in SRC management (NCAA Sports Science Institute, 2017). The order in which these measures were administered varied and depended predominantly on each participant’s schedule.

Certified athletic trainers administered ImPACT within the university’s Sports Medicine department and members of this research team. IRB approval allowed access to all ImPACT baseline scores completed by participating athletes before study enrollment, referred to here as “Baseline 1.” These scores were collected according to the usual Sports Medicine procedure based on NCAA preparticipation baseline assessment guidelines. Repeated baseline ImPACT scores, referred to as “Baseline 2,” were obtained from all athletes participating in this study. All participants completed both baseline assessments.

Upon presentation to sports medicine for assessment, ImPACT was administered following usual clinical practice; athletes were instructed to silence their phones, focus solely on the assessment, and follow the written directions in each section. Athletes completed the initial demographics/background/medical history section of ImPACT before the start of the cognitive testing. All assessments were conducted in a private office/computer lab and administered to athletes in small groups of up to three athletes.

Study data were collected and managed using REDCap electronic data capture tools. REDCap (Research Electronic Data Capture) is a secure, web-based application designed to support data capture for research studies, proving: an intuitive interface to validate data entry; audit trails for tracking data manipulation and export procedures; automated export procedures for seamless data downloads to common statistical packages; and procedures for importing data from external sources (Harris, Taylor, Thielke, Payne, Gonzalez, & Conde, 2009).

Data Analysis

Our analytic approach consists of multiple analyses, including the calculation of correlations (Pearson’s r and ICCs), Bland–Altman plots with limits of agreement (LOA), reliable change indices (RCIs), minimal detectable change (MDC) measures, and an examination of practice effects. Each analysis provides information about different aspects of the ImPACT’s temporal stability, improving the clinical application of the results.

Pearson’s Product Moment Correlations (r) and ICCs

Temporal stability, or test–retest reliability, can be assessed using both Pearson’s r and ICC, and both measures of association are frequently reported when studying temporal stability (Brett et al., 2016; Echemendia et al., 2016; Koo & Li, 2016; Vaz et al., 2013). Though Pearson’s r provides a general measure of overall linear association between baseline scores and follow-up scores, it does not reflect consistent changes in scores across trials (Vaz et al., 2013). In other words, a “perfect” correlation of +1 does not necessarily indicate perfect agreement but instead suggests a consistent and linear relationship between two sets of scores. The ICC is a more robust measure of test–retest reliability than Pearson’s r, as it considers both within-subject variability at baseline and retest sessions and any “systematic change” in the overall sample mean across trials (Vaz et al., 2013; Wilk et al., 2002). The ICC is frequently reported as a preferential indicator of temporal stability (Fernández-Marcos, de la Fuente, & Santacreu, 2018; Schatz, 2010; Wilk et al., 2002). ICC estimates with 95% confidence intervals were based on a mean rating (k = 1), absolute agreement, two-way mixed-effects model. ICCs equal or above 0.90 suggest excellent reliability, 0.75–0.9 suggest good reliability, 0.50–0.75 suggest moderate reliability, and values below 0.50 suggest poor reliability (Koo and Li, 2016). Here, both Pearson’s r and ICCs will be reported to serve as a comparison to previous findings which reported both measures of correlation.

Limits of Agreement

Agreement between Baseline 1 and Baseline 2 assessments was evaluated using 95% confidence interval (CI) of upper and lower LOAs representing measurement error (ME) boundaries as displayed by Bland–Altman plots (Bland & Altman, 1986, 1999, 2003; Carkeet, 2015; Sedgwick, 2013). Baseline means and differences were calculated, then one-sample t-tests were conducted to determine significant difference between means of the two measurements. Presence of fixed bias is indicated if the mean value of the difference differs statistically from zero, thus violating the assumption that the repeated scores are equivalent. A statistically significant p-value indicates the measurements are significantly different from one another, and therefore, LOAs are not meaningful in a test–retest application (Bland & Altman, 1999).

Bland–Altman plots were constructed by assigning the differences on the vertical axis and the means on the horizontal axis, LOAs and CIs were estimated which allowed for visual evaluation of the direction and magnitude of the relationship between score means and differences allowing for investigation of potential ME and assumed true value relationships. Heteroscedasticity was visually examined and tested using correlation between the differences and observed means were calculated and tested against H0: r = 0 for each of the five modules.

Reliable Change Index

RCIs were calculated to evaluate whether change from Baseline 1 to Baseline 2 was reliable and meaningful (Jacobson & Truax, 1991). Change is only meaningful for practical or clinical significance if it is reliable. RCIs provide probability estimates that a given difference score was not acquired due to ME (Iverson, Sawyer, McCracken, & Kozora, 2001). The efficacy of this method to assess reliable change with ImPACT has been established in identifying meaningful change assisting in sports concussion detection and management (Iverson, Lovell, & Collins, 2003). RCI cut-off scores for the five modules were calculated from the test–retest data (Chelune, Naugle, Lüders, Sedlak, & Awad, 1993).

Minimal Detectable Change and Practice Effects

To investigate random ME, MDCs were calculated at the 95% CI as recommended by Haley and Fragala-Pinkham (2006). We followed Huang and colleagues (2011) for MDC% computation and interpretation guidelines (acceptable MDC% < 30). Formulas are displayed subsequently. Cohen’s d was used to assess potential practice effects from paired samples t-test estimates (Cohen, 1988).

Bootstrapping

To improve our confidence in the statistical measures of temporal stability, a bootstrapping model was applied to all point estimates with 1,000 replications (Efron, 1979). Bootstrapping is a nonparametric statistical technique involving repeated sampling of study data with replacement (Miller et al., 2002). This approach allows for a better estimation of the theoretical variation of the estimate within the general population of NCAA Division I collegiate athletes, providing information about the accuracy and functional limitations the statistic (Davidson & Hinkley, 1997; Miller et al., 2002; Varian, 2005).

All statistical analyses were conducted using SPSS 25.

Results

Baseline 1 to Baseline 2 Comparisons

Pearson’s correlations ranged from 0.30 to 0.49 between baseline measures. As expected, mean ICC results indicated moderate reliability across all measures with 95% CI except Impulse Control which revealed poor reliability. Visual Motor Speed scores produced the highest stability level (ICC = 0.66; [0.40, 0.81]) followed by Verbal Memory (ICC = 0.65; [0.37, 0.80]), Visual Memory (ICC = 0.61; [0.32, 0.78]), Reaction Time (ICC = 0.51; [0.14, 0.73]), and Impulse Control (ICC = 0.45; [0.03, 0.69]). Table 3 displays summary statistics.

Table 3

Test–retest reliability

AssessmentBaseline 1Baseline 2raICCICC 95% CI
MSDMSDLowerUpper
Verbal memory8998810.480.650.370.80
Visual memory76108013.470.610.320.78
Visual motor speed39.16.840.37.2.490.660.400.81
Reaction time0.600.090.590.08.350.510.140.73
Impulse c31ontrol5465.300.450.030.69

Note: CI = confidence interval; ICC = intraclass correlation coefficient; LOA = limits of agreement; r = Pearson’s product moment correlation.

aBootstrapping model utilized based on 1,000 samples.

bStatistical significance indicates measurements are significantly different from one another, therefeore useful LOAs are not meaningful.

Bland–Altman measurement difference classification criteria were met for four modules. Initial t-test results indicated the presence of fixed bias for the Visual Memory module (t47 = −2.24; p = 0.03), therefore no plot or LOAs calculations were conducted. Examination of Bland–Altman plots revealed general random scatter for all four other modules; no evidence of heteroscedasticity was indicated. Only the Visual Motor Speed plot is provided for the sake of brevity (Fig. 2). Upper and lower LOA bounds and 95% CIs were spread on either side of zero, thereby indicating the difference between Baseline 1 and Baseline 2 measurements due to ME alone (Table 4).

Bland–Altman plot of ImPACT Visual Motor Speed Scores (Baseline 1 and Baseline 2). ImPACT = Immediate Post-Concussion Assessment and Cognitive Test.
Fig. 2

Bland–Altman plot of ImPACT Visual Motor Speed Scores (Baseline 1 and Baseline 2). ImPACT = Immediate Post-Concussion Assessment and Cognitive Test.

Table 4

Limits of agreement

ModuleMaSDatapcLOA LBLOA UBLOA 95% CIb
Verbal Memory0.839.520.610.55−17.819.5[−0.45, 0.10]
Visual Memory-3.8511.95-2.240.03---
Visual Motor Speed-1.177.06-1.150.26−15.012.7[−0.52, 0.37]
Reaction Time0.010.101.040.30−0.170.20[−0.66, 0.72]
Impulse Control-0.834.86-1.190.24−10.48.70[−0.84, 0.25]

aM and SDs for mean differences. One-sample t -test; α ≤ .05; df = 47. Bootstrapping model utilized based on 1,000.

bLOA = limits of agreement. CI = confidence interval. Statistically significant p value indicates measurements are significantly different from one another, therefeore useful LOAs are not meaningful. n = 48.

Reliable Change Index

The RCI provides a value that reflects the necessary difference between scores to reflect “statistically reliable” change (Chelune et al., 1993; Rai et al., 2015). Fig. 3 illustrates the percent of athletes with change scores within cut-off range at 80%, 90%, and 95% CIs. Interestingly, percent for each range increased as CI level decreased as expected, except for Reaction Time which decreased from 27.1% to a minimum 18.7% then increased to a 39.6% maximum (95% CI, 90% CI, 80% CI respectively), reinforcing the directive that clinicians should use the 95% confidence interval for Reaction Time Composite when making clinical decisions.

Percent of athletes with scores that reliably changed within cutoff at each range.
Fig. 3

Percent of athletes with scores that reliably changed within cutoff at each range.

Minimal Detectable Change and Practice Effects

Table 5 provides results of MDC and practice effect analyses. Similar to RCI, the MDC provides clinicians a guide to interpretation of change scores by providing a minimal amount of change between repeated scores required to reflect meaningful clinical change that extends beyond ME (Haley & Fragala-Pinkham, 2006; Rai et al., 2015). MDC95% estimates were less than 30% across all modules (16.0%–29.0%) except for Impulse Control (143.2%). Impulse Control estimates at the 80%–90% CIs met the criteria more closely (MDC90% = 33.4%; MDC85% = 29.3%; MDC80% = 26.0%). The Impulse Control Composite MCD% performed as expected, where the score reflects number of errors across all testing modules which should change over multiple administrations. Practice effects statistics revealed negligible to small effect sizes across all measures ranging from 0.09 to 0.32.

Table 5

Parameters of random measurement error and practice effect

AssessmentsRandom measurement errorPractice effect
MDC95aMDC%Difference (Mean ± SD)tbp  cd  d
Verbal memory14.1916.0%89 ± 90.61.5470.09
Visual memory17.2522.6%76 ± 10-2.24.0300.32
Visual motor speed11.0728.3%39.10 ± 6.84-1.15.257-0.17
Reaction time0.1729.0%0.60 ± 0.091.04.3020.15
Impulse control7.38143.2%5 ± 4-1.19.241-0.17

aMDC = minimal detectable change. Bootstrapping model utilized based on 1,000 samples.

bPaired-samples t -test; df = 47.

cp ≤ .05.

dCohen’s d effect size ranges: small (0.2), medium (0.5), and large (≥0.8).

Discussion

In this study we examined the temporal stability of the ImPACT test across retest timepoints of various lengths in a cohort study of NCAA Division I athletes. Our results failed to support the NCAA recommendations for obtaining a one-time preparticipation cognitive baseline within the context of SRC management. Overall, the results reflect low moderate temporal stability across four of the five composite scores, including Visual Motor Speed, Verbal Memory, Visual Memory, and Reaction Time with ICC ranging from 0.51 to 0.66. In contrast to previously reported findings (Brett et al., 2016; Maerlender et al., 2016; Schatz & Ferris, 2013), this study found that no composite score demonstrated good or excellent reliability (ICC > 0.75). As expected, the Impulse Control Composite score demonstrated particularly low stability, with an ICC < 0.50. The ImPACT manual instructs providers to utilize the Impulse Control Composite score as a measure of validity rather than a measure of cognitive ability (ImPACT Applications, Inc., 2016). Unlike cognitive measures, the value of a performance validity measure lies in the effective and accurate categorical identification of valid and invalid data, usually as determined by cut-off scores. The value of the actual score achieved on performance validity measures is less important than its position relative to the predetermined cutoff. The Impulse Control Composite’s efficacy as a performance validity indicator has been previously documented (Brett & Soloman, 2017).

This more exhaustive exploration of temporal stability is important within the larger context of ImPACT reliability research, because the majority of previous inquiries into ImPACT reliability utilized statistical analyses (i.e., ICC analyses inconsistent with two-way mixed effects model with absolute agreement) that are inappropriate for the demonstration of test–retest reliability. As previously discussed, there are a variety of statistical approaches to the calculation of ICC, and each approach is associated with unique statistical assumptions (Koo & Li, 2016).

As suggested by Koo and Li (2016), the analysis of test–retest reliability using ICCs should always include the two-way mixed-effect model with absolute agreement, noting the importance that the model acknowledges that the retest samples are not independent of each other and that reliability is dependent on the absolute agreement of the repeated measures. That is to say, when considering reliability, it is crucial for repeated baseline scores to be equivalent, with variability secondary only to ME, and not simply correlated. Alternative approaches to ICC analysis rely upon different statistical formulas and do not include similar assumptions. The variability in ICCs across the ImPACT temporal stability literature is likely due, at least in part, to the variability in approaches to ICC analysis.

Though the results indicate moderate (i.e., above “adequate”) reliability for four of the five composite scores, according to Koo & Li (2016), our reliability findings do not necessarily suggest that ImPACT is adequately reliable for use in clinical decision-making for individual student athletes. Nunnally and Bernstein (1994) recommended that psychometric measures used to make decisions about individuals must have reliability coefficients of at least 0.90. Even at this level, Nunnally and Bernstein note that ME is sufficiently large to undermine precision and accuracy of individual measurements. In other words, basing return-to-play decisions on measures with suboptimal reliability may either place athletes at undue risk for symptom exacerbation or may result in student athletes being withheld from play too long. As such, the reliability results in the present study do not indicate support for the utility and efficacy of baseline ImPACT testing, as suggested by the NCAA. Furthermore, the reliance on normative data alone will eliminate error secondary to the limitations discussed earlier (e.g., variability in testing environments, psychometric properties, etc.), the most important of which may be sandbagging efforts by athletes (Bailey, Echemendia, & Arnett, 2006; Schatz et al., 2017).

We utilized an analysis of Bland–Altman agreement for the purpose of further exploring the variability between repeated measures. This analysis revealed no evidence of heteroscedasticity or fixed bias within the Verbal Memory, Visual Motor Speed, Reaction Time, and Impulse Control Composite scores across repeated baselines in Division I athletes, suggesting that variation in scores across baselines was due to ME and supporting the reported findings of moderate reliability. However, results did reveal a violation of the equivalence assumption for the Visual Memory Composite score, suggesting that sports medicine clinicians who do utilize baseline comparisons should interpret changes on this composite score with more caution. Specifically, changes in Verbal Memory Composite, Visuomotor Speed Composite, and Reaction Time Composite are each more likely to reflect true change in SRC neurocognitive functioning relative to the Visual Memory Composite score.

Our findings reveal several other important considerations for return-to-play decision making when using the current NCAA recommendation of baseline comparison. Examination of MDC revealed MDC% were below 30% and exceeded LOAs for all four cognitive composite scores. The MDC provides a guideline for clinicians when interpreting repeated ImPACT results by quantifying the difference between scores that reflects true change in the measured variable (e.g., neurocognitive function) beyond simple ME. By understanding the requirements for meaningful change, clinicians can feel more confident when attributing a decline in post-SRC ImPACT scores to neurological dysfunction when making return-to-play decisions for concussed athletes.

Results suggest that practice effects were small to negligible across all composite scores, indicating that differences in ImPACT scores are unrelated to collegiate athletes simply “getting better” at the test. The observed temporal stability of the Impulse Control Composite supports the ImPACT manual’s instruction to interpret these scores as a measure of performance validity instead of cognitive performance (ImPACT Applications, Inc., 2016). Changes in the Impulse Control Composite are reflective of changes in the number of errors made across modules and are not inherently meaningful (ImPACT Applications, Inc., 2016).

Based on evidence of heteroscedasticity within the Visual Memory Composite, these results suggest that greater clinical emphasis should be placed on Verbal Memory, Visuomotor Speed, and Reaction Time Composite scores when comparing postconcussive performance to baseline when making return-to-play decisions. When using Verbal Memory, Visuomotor Speed, and Reaction Time Composite scores to determine meaningful change in cognitive status, clinicians are encouraged to consider MDC scores and RCIs for each composite score. If the absolute value of the score change exceeds the MDC or RCI, the clinician may conclude that the change they are observing is reflective of true cognitive change, increasing the confidence of return-to-play decisions (Haley & Fragala-Pinkham, 2006). Importantly, a change in ImPACT scores beyond expected ME may be associated with a variety of etiologies. The determination about the degree to which cognitive changes can be attributed to the concussive injury is a matter of professional clinical judgment and requires consideration of individual medical history, course of symptoms, and nature of the injury (McCrory et al., 2018).

There are several limitations to the present study. First, although a cohort study design provides several methodological and interpretive advantages including the representation of real-life variability that is eliminated in controlled studies, there are limitations to this approach (Carlson & Morrison, 2009; Mann, 2003). Specifically, by failing to eliminate those controllable sources of variability, another unidentified variable may have an additional, unaccounted for, influence on statistical outcomes, weakening the utility of these findings.

Second, the racial and ethnic homogeneity of this sample limits the generalizability of our findings to more diverse populations of student athletes. This sample was predominantly composed of self-identified Caucasian participants (>81%). Importantly, the NCAA has estimated that about 36% of student athletes identify as something other than Caucasian (NCAA, 2018). Future research should continue to examine the psychometric properties of ImPACT in more diverse samples to ensure that students of all racial and ethnic backgrounds receive the most effective concussion management.

In addition to demographic homogeneity, our small sample size represented an additional limitation. Statistical bootstrapping was applied to all point estimates, however allowing for more robust findings. Nevertheless, further research should attempt to replicate similar findings within a larger sample size. In future years, we will be able to expand this preliminary study significantly.

Third, this study relied on self-reported medical history, including brain injury history.

As has been previously reported, the incidence of concussion within collegiate and professional athletics is believed to be largely underreported, particularly when relying on athlete self-report (Echlin et al., 2010; Harmon et al., 2013; Meier et al., 2015). Without objective criteria for injury history, athletes with concussions may have been included in our analysis of noninjured athletes with repeated baselines. If this, indeed, occurred, these results may reflect inaccurate estimates of temporal stability and suggest repeated administration interpretation guidelines that do not correspond with true change in neurocognitive status. Future studies on concussion, particularly with athlete populations, may benefit from both consulting with personal medical records and regularly monitoring participants for symptoms of unreported injuries to ensure accurate medical classification. Extending the cohort to include high-school athletes in the future may also be one way in which an individual athlete’s medical history of concussions can be followed more closely and accurately precollege.

Finally, athletes with self-reported histories of ADHD were excluded in order to be consistent with previous ImPACT research. The exclusion of athletes with ADHD limits the generalizability of these results to the estimated 8% of college students with a diagnosis of ADHD (DuPaul, Weyandt, O’Dell, &Varejao, 2009). Research actually suggests that there is a higher incidence of concussive injuries among persons with ADHD (Alosco, Fedor, & Gunstad, 2014). There is a tremendous need for research on the relationship between ADHD and concussive injury course and management. Along a separate line of research, we will include athletes with ADHD and concussions in the future to delve further into this interesting aspect of concussions.

Conclusion

In the current study, our findings suggest that interpreting score changes from baseline to postconcussion injury may not be the best course of action even though this is currently directed by the NCAA recommendation for SRC management. An in-depth review of the temporal stability of ImPACT scores in our cohort of NCAA athletes suggests that the five domain scores are not adequately robust for use in individual decision-making regarding return-to-play. Given the widespread use of the NCAA recommendations, however this research suggests that clinicians who do rely on baseline comparisons should refer to the attached guidelines for identifying significant ImPACT score changes over time and interpret Visual Memory Composite scores with particular caution. These results also highlight the need for more research on the psychometric properties of the ImPACT with diverse athletes and athletes with a history of ADHD.

Funding

This work was supported in part by a Knoebel Institute for Health Aging Pilot Award.

References

Alosco
,
M. L.
,
Fedor
,
A. F.
, &
Gunstad
,
J.
(
2014
).
Attention deficit hyperactivity disorder as a risk factor for concussions in NCAA division-I athletes
.
Brain Injury
,
28
,
472
474
. doi: .

Bailey
,
C. M.
,
Echemendia
,
R. J.
, &
Arnett
,
P.
(
2006
).
The impact of motivation on neuropsychological performance in sports-related mild traumatic brain injury
.
Journal of the International Neuropsychological Society
,
12
,
475
484
. doi: .

Barkhoudarian
,
G.
,
Hovda
,
D. A.
, &
Giza
,
C. C.
(
2016
).
The molecular pathophysiology of concussive brain injury – Update
.
Physical Medicine and Rehabilitation Clinics of North America
,
27
,
373
393
. doi: .

Bartels
,
C.
,
Wegrzyn
,
M.
,
Wiedl
,
A.
,
Ackermann
,
V.
, &
Ehrenreich
,
H.
(
2010
).
Practice effects in health adults: A longitudinal study on frequent repetitive cognitive testing
.
BMC Neuroscience
,
11
,
1
12
. doi: .

Belanger
,
H. G.
, &
Vanderploeg
,
R. D.
(
2005
).
The neuropsychological impact of sports-related concussion: A meta-analysis
.
Journal of the International Neuropsychological Society
,
11
,
345
357
. doi: .

Binder
,
L. M.
,
Rohling
,
M. L.
, &
Larrabee
,
G. J.
(
1997
).
A review of mild head trauma. Part I: Meta-analytic review of neuropsychological studies
.
Journal of Clinical and Experimental Neuropsychology
,
19
(
3
),
421
431
. doi: .

Bland
,
J. M.
, &
Altman
,
D. G.
(
1986
).
Statistical methods for assessing agreement between two methods of clinical measurement
.
Lancet
,
327
(
8476
),
307
310
. doi: .

Bland
,
J. M.
, &
Altman
,
D. G.
(
1999
).
Measuring agreement in method comparison studies
.
Statistical Methods in Medical Research
,
8
(
2
),
135
160
. doi: .

Bland
,
J. M.
, &
Altman
,
D. G.
(
2003
).
Applying the right statistics: Analyses of measurement studies
.
Ultrasound in Obstetrics & Gynecology
,
22
,
85
93
. doi: .

Broglio
,
S. P.
,
Ferrara
,
M. S.
,
Macciocchi
,
S. N.
,
Baumgartner
,
T. A.
, &
Elliot
,
R.
(
2007
).
Test-retest reliability of computerized concussion assessment programs
.
Journal of Athletic Training
,
42
(
4
),
509
514
.

Brett
,
B. L.
,
Smyk
,
N.
,
Solomon
,
G.
,
Baughman
,
B. C.
, &
Schatz
,
P.
(
2016
).
Long-term stability and reliability of baseline cognitive assessment in high school athletes using ImPACT at 1-, 2-, and 3-year test-retest intervals
.
Archives of Clinical Neuropsychology
,
31
,
904
914
. doi: .

Brett
,
B. L.
, &
Solomon
,
G. S.
(
2017
).
The influence of validity criteria on immediate post concussion assessment and cognitive testing (ImPACT) test-retest reliability among high school athletes
.
Journal of Clinical and Experimental Neuropsychology
,
39
,
286
295
. doi: .

Carlson
,
M. D.
, &
Morrison
,
R. S.
(
2009
).
Study design, precision, and validity in observational studies
.
Journal of Palliative Medicine
,
12
,
77
82
. doi: .

Cassidy
,
J. D.
,
Carroll
,
L.
,
Peloso
,
P.
,
Borg
,
J.
,
Von Holst
,
H.
,
Holm
,
L.
, et al. (
2004
).
Incidence, risk factors and prevention of mild traumatic brain injury: Results of the WHO collaborating centre task force on mild traumatic brain injury
.
Journal of Rehabilitation Medicine
,
36
,
28
60
. doi: .

Chelune
,
G. J.
,
Naugle
,
R. I.
,
Lüders
,
H.
,
Sedlak
,
J.
, &
Awad
,
I. A.
(
1993
).
Individual change after epilepsy surgery: Practice effects and base-rate information
.
Neuropsychology
,
7
,
41
52
. doi: .

Cohen
,
J.
(
1988
).
Statistical power analysis for the behavioral sciences
( 2nd ed.).
Hillsdale, N.J
:
Lawrence Erlbaum
.

Covassin
,
T.
,
Swanik
,
C.
, &
Sachs
,
M. L.
(
2003
).
Epidemiological considerations of concussions among intercollegiate athletes
.
Applied Neuropsychology
,
10
,
12
22
. doi: .

Davidson
,
A. C.
, &
Hinkley
,
D. V.
(
1997
).
Bootstrap methods and their application
.
New York, NY
:
Cambridge University Press
.

De Marco
,
A. P.
, &
Broshek
,
D. K.
(
2016
).
Computerized cognitive testing in the management of youth sports-related concussion
.
Journal of Child Neurology
,
31
(
1
),
68
75
. doi: .

Dupaul
,
G. J.
,
Weyandt
,
L. L.
,
O’Dell
,
S. M.
, &
Varejao
,
M.
(
2009
).
College students with ADHD: Current status and future directions
.
Journal of Attention Disorders
,
13
,
230
250
. doi: .

Echemendia
,
R. J.
,
Bruce
,
J. M.
,
Meeuwisse
,
W.
,
Comper
,
P.
,
Aubry
,
M.
, &
Hutchison
,
M.
(
2016
).
Long-term reliability of ImPACT in professional ice hockey
.
The Clinical Neuropsychologist
,
30
,
311
320
. doi: .

Echemendia
,
R. J.
,
Bruce
,
J. M.
,
Bailey
,
C. M.
,
Sanders
,
J. F.
,
Arnett
,
P.
, &
Vargas
,
G.
(
2012
).
The utility of post-concussion neuropsychological data in identifying cognitive change following sports-related mTBI in the absence of baseline data
.
The Clinical Neuropsychologist
,
26
,
1077
1091
. doi: .

Echemendia
,
R. J.
,
Herring
,
S.
, &
Bailes
,
J.
(
2009
).
Who should conduct and interpret neuropsychological assessment in sports-related concussion?
 
British Journal of Sports Medicine
,
43
,
32
35
. doi: .

Echlin
,
P. S.
,
Tator
,
C. H.
,
Cusimano
,
M. D.
,
Cantu
,
R. C.
,
Taunton
,
J. E.
,
Upshur
,
R. E.
, et al. (
2010
).
A prospective study of physician-observed concussions during junior ice hockey: Implications for incidence rates
.
Neurosurgical Focus
,
29
(
5
),
E4
. doi: .

Efron
,
B.
(
1979
).
Bootstrap methods: Another look at the jackknife
.
Annals of Statistics
,
7
,
1
26
. doi: .

Elbin
,
R. J.
,
Kontos
,
A. P.
,
Kegel
,
N.
,
Johnson
,
E.
,
Burkhart
,
S.
, &
Schatz
,
P.
(
2013
).
Individual and combined effects of LD and ADHD on computerized concussion test performance: Evidence for separate norms
.
Archives of Clinical Neuropsychology
,
28
,
476
484
. doi: .

Fernández-Marcos
,
T.
,
de la
 
Fuente
,
C.
, &
Santacreu
,
J.
(
2018
).
Test-retest reliability and convergent validity of attention measures
.
Applied Neuropsychology: Adult
,
25
,
464
472
. doi: .

Fukata
,
K.
,
Amimoto
,
K.
,
Sekine
,
S.
,
Ikarashi
,
Y.
,
Fujino
,
Y.
,
Inoue
,
M.
, et al. (
2019
).
Test-retest reliability of and age-related changes in the subjective postural vertical on the diagonal plane in health subjects
.
Attention, Perception, & Psychophysics
,
81
,
590
597
. doi: .

Galetta
,
K. M.
,
Brandes
,
L.
E.
,
Maki
,
K.
,
Dziemianowicz
,
M. S.
,
Laudano
,
E.
,
Allen
,
M.
, et al. (
2011
).
The King-Devick test and sports related concussion: Study of a rapid visual screening tool in a collegiate cohort
.
Journal of the Neurological Sciences
,
309
,
34
39
. doi: .

Guskiewicz
,
K. M.
,
Ross
,
S. E.
, &
Marshall
,
S. W.
(
2001
).
Postural stability and neuropsychological deficits after concussion in collegiate athletes
.
Journal of Athletic Training
,
36
(
3
),
263
273
.
Retrieved from
. https://search-proquest-com.du.idm.oclc.org/docview/206653120?accountid=14608.

Haley
,
S. M.
, &
Fragala-Pinkham
,
M. A.
(
2006
).
Interpreting change scores of tests and measures used in physical therapy
.
Physical Therapy
,
86
,
735
743
. doi: .

Harmon
,
K. G.
,
Drezner
,
J. A.
,
Gammons
,
M.
,
Guskiewicz
,
K. M.
,
Halstead
,
M.
,
Herring
,
S.
, et al. (
2013
).
American medical Society for Sports Medicine position statement: Concussion in sport
.
British Journal of Sports Medicine
,
47
,
15
26
. doi: .

Harris
,
P. A.
,
Taylor
,
R.
,
Thielke
,
R.
,
Payne
,
J.
,
Gonzalez
,
N.
, &
Conde
,
J. G.
(
2009
).
Research electronic data capture (REDCap) – A metadata-driven methodology and workflow process for providing translational research informatics support
.
Journal of Biomedical Information
,
42
(
2
),
377
381
. doi: .

Higgins
,
K. L.
,
Caze
,
T.
, &
Maerlender
,
A. C.
(
2018
).
Validity and reliability in of baseline testing in a standardized environment
.
Archives of Clinical Neuropsychology
,
33
,
437
443
. doi: .

Huang
,
S. L.
,
Hsieh
,
C. L.
,
Wu
,
R. M.
,
Tai
,
C. H.
,
Lin
,
C. H.
, &
Lu
,
W. S.
(
2011
).
Minimal detectable change of the timed “up & go” test and the dynamic gait index in people with Parkinson disease
.
Physical Therapy
,
91
,
114
121
. doi: .

Iverson
,
G. L.
,
Sawyer
,
D. C.
,
McCracken
,
L. M.
, &
Kozora
,
E.
(
2001
).
Assessing depression in systemic lupus erythematosus: Determining reliable change
.
Lupus
,
10
(
4
),
266
271
. doi: .

Iverson
,
G. L.
,
Lovell
,
M. R.
, &
Collins
,
M. W.
(
2003
).
Interpreting change in ImPACT following sport concussion
.
The Clinical Neuropsychologist
,
17
,
460
467
. doi: .

Iverson
,
G. L.
,
Lovell
,
M. R.
, &
Collins
,
M. W.
(
2005
).
Validity of ImPACT for measuring processing speed following sports-related concussion
.
Journal of Clinical and Experimental Neuropsychology
,
27
,
683
689
. doi: .

Jacobson
,
N. S.
, &
Truax
,
P.
(
1991
).
Clinical significance: A statistical approach to defining meaningful change in psychotherapy research
.
Journal of Consulting and Clinical Psychology
,
59
(
1
),
12
19
. doi: .

Jager
,
J.
,
Putnick
,
D. L.
, &
Bornstein
,
M. H.
(
2017
).
MOre than just convenient: The scientific merits of homogeneous convenience samples
.
Monographs of the Society for Research in Child Development
,
82
,
13
30
. doi: .

Karr
,
J. E.
,
Areshenkoff
,
C. N.
, &
Garcia-Barrera
,
M. A.
(
2014
).
The neuropsychological outcomes of concussion: A systematic review of meta-analyses on the cognitive sequelae of mild traumatic brain injury
.
Neuropsychology
,
28
(
3
),
321
366
. doi: .

Kerr
,
Z. Y.
,
Snook
,
E. M.
,
Robert
,
L. C.
,
Dompier
,
T. P.
,
Sales
,
L.
,
Parsons
,
J. T.
, et al. (
2015
).
Concussion-related protocols and preparation assessments used for incoming student-athletes in National Collegiate Athletic Association Member Institutions
.
Journal of Athletic Training
,
50
,
1174
1181
. doi: .

Koo
,
T. K.
, &
Li
,
M. Y.
(
2016
).
A guideline of selecting and reporting intraclass correlation coefficients for reliability research
.
Journal of Chiropractic Medicine
,
15
,
155
163
.

Langlois
,
J. A.
,
Rutland-Brown
,
W.
, &
Wald
,
M. M.
(
2006
).
The epidemiology and impact of traumatic brain injury: A brief overview
.
Journal of Head Trauma Rehabilitation
,
21
(
5
),
375
378
. doi: .

Louey
,
A.
,
Cromer
,
J. A.
,
Schembri
,
A. J.
,
Darby
,
D. G.
,
Maruff
,
P.
,
Makdissi
,
M.
, et al. (
2014
).
Detective cognitive impairment after concussion: Sensitivity of change from baseline and normative data methods using the CogSport/axon cognitive test battery
.
Archives of Clinical Neuropsychology
,
29
,
432
411
. doi: .

Lovell
,
M. R.
(
2016
).
ImPACT: Administration and interpretation manual
.
San Diego, CA
:
ImPACT Applications Inc
.

Lovell
,
M. R.
,
Collins
,
M. W.
,
Podell
,
K.
,
Powell
,
J.
, &
Maroon
,
J.
(
2000
).
ImPACT: Immediate post-concussion assessment and cognitive testing
.
Pittsburgh, PA
:
NeuroHealth Systems, LLC
.

Lezak
,
M. D.
,
Howieson
,
D. B.
,
Bigler
,
E. D.
, &
Tranel
,
D.
(
2012
).
Neuropsychological assessment
( 5th ed.).
New York, NY
:
Oxford University Press, Inc
.

MacFarlane
,
M. P.
, &
Glenn
,
T. C.
(
2015
).
Neurochemical cascade of concussion
.
Brain Injury
,
29
,
139
153
. doi: .

Maerlender
,
A.
,
Flashman
,
L.
,
Kessler
,
A.
,
Kumbhani
,
S.
,
Greenwald
,
R.
,
Tosteson
,
T.
, et al. (
2010
).
Examination of the construct validity of ImPACT computerized test, traditional, and experimental neuropsychological measures
.
The Clinical Neuropsychologist
,
24
,
1309
1325
. doi: .

Maerlender
,
A. C.
,
Masterson
,
C. J.
,
James
,
T. D.
,
Beckwith
,
J.
,
Brolinson
,
P. G.
,
Crisco
,
J.
, et al. (
2016
).
Test-retest, retest, and retest: Growth curve models of repeat testing with immediate post-concussion assessment and cognitive testing (ImPACT)
.
Journal of Clinical and Experimental Neuropsychology
,
38
(
8
),
869
874
. doi: .

Mann
,
C. J.
(
2003
).
Observational research methods. Research design II: Cohort, cross-sectional, and case-control studies
.
Emergency Medicine Journal
,
20
,
54
60
. doi: .

McCrory
,
P.
,
Meeuwisse
,
W. H.
,
Aubry
,
M.
,
Cantu
,
R. C.
,
Dvoøak
,
J.
,
Echemendia
,
R. J.
, et al. (
2013
).
Consensus statement on concussion in sport: The 4th international conference on concussion in sport, Zurich, November 2012
.
Clinical Journal of Sports Medicine
,
23
,
89
117
. doi: .

McCrory
,
P.
,
Meeuwisse
,
W.
,
Kvorak
,
J.
,
Aubry
,
M.
,
Bailes
,
J.
,
Broglio
,
S.
, et al. (
2018
).
Consensus statement on concussion in sport—The 5th international conference on concussion in sport held in berlin, October 2016
.
British Journal of Sports Medicine
,
15
,
838
847
. doi: .

Meier
,
T. B.
,
Brummel
,
B. J.
,
Singh
,
R.
,
Nerio
,
C. J.
,
Polanski
,
D. W.
, &
Bellgowan
,
P. S. F.
(
2015
).
The underreporting of self-reported symptoms following sports-related concussion
.
Journal of Science and Medicine in Sport
,
18
,
507
511
. doi: .

Miller
,
E.
,
Neal
,
D.
,
Roberts
,
L.
,
Baer
,
J.
,
Cressler
,
S. O.
,
Metrik
,
J.
, et al. (
2002
).
Test-retest reliability of alcohol measures: Is there a difference between internet-based assessments and traditional methods?
 
Psychology of Addictive Behavior
,
16
,
56
63
. doi: .

National Collegiate Athletic Association (NCAA)
. (
2015
).
Transfer Terms: Eligibility timeline
.
Retrieved from
 http://www.ncaa.org/student-athletes/current/transfer-terms

National Collegiate Athletic Association (NCAA) Sports Science Institute (2017)
.
Inter-association consensus: Diagnosis and management of sports related concussion best practices
.
Retrieved on February 8, 2019 from
 http://www.ncaa.org/sites/default/files/SSI_ConcussionBestPractices_20170616.pdf

National Collegiate Athletic Association
. (
2018
).
NCAA Demographics Database [Data visualization dashboard]. Retrieved on June 1, 2019 from
 http://www.ncaa.org/about/resources/research/ncaa-demographics-database.

Prins
,
M. L.
,
Hales
,
A.
,
Reger
,
M.
,
Giza
,
C. C.
, &
Hovda
,
D. A.
(
2010
).
Repeat traumatic brain injury in the juvenile rat is associated with increased axonal injury and cognitive impairments
.
Developmental Neuroscience
,
32
,
510
518
. doi: .

Rai
,
S. K.
,
Yazdany
,
J.
,
Fortin
,
P.
, &
Aviña-Zubieta
,
J. A.
(
2015
).
Approaches for estimating minimally clinically important differences in systemic lupus erythematosus
.
Arthritis Research and Therapy
,
17
,
1
8
. doi: .

Randolph
,
C.
(
2011
).
Baseline neuropsychological testing in managing sports related concussion: Does it modify risk?
 
Current Sports Medicine Reports
,
10
,
21
26
. doi: .

Rebchuk
,
A. D.
,
Brown
,
H. J.
,
Koehle
,
M. S.
,
Blouin
,
J.
, &
Siegmund
,
G. P.
(
2020
).
Using variance to explore the diagnostic utility of baseline concussion testing
.
Journal of Neurotrauma. Advanced online publication.
. doi: .

Schatz
,
P.
(
2010
).
Long-term test-retest reliability of baseline cognitive assessment using ImPACT
.
The American Journal of Sports Medicine
,
38
,
47
53
. doi: .

Schatz
,
P.
(
2011
). Computerized neuropsychological assessment in sport. In
Webbe
,
F. M.
(Ed.),
The handbook of sport neuropsychology
(, pp.
173
186
).
New York, NY
:
Springer Publishing Company, LLC
.

Schatz
,
P.
,
Elbin
,
R. J.
,
Anderson
,
M. N.
,
Savage
,
J.
, &
Covassin
,
T.
(
2017
).
Exploring sandbagging behaviors, effort, and perceived utility of the ImPACT baseline assessment in college athletes
.
Sport, Exercise, and Performance Psychology
,
6
,
243
251
. https://doi.org/10.1037/spy0000100.

Schatz
,
P.
, &
Ferris
,
C. S.
(
2013
).
One-month test-retest reliability of the ImPACT test battery
.
Archives of Clinical Neuropsychology
,
28
,
499
504
. doi: .

Schatz
,
P.
, &
Robertshaw
,
S.
(
2014
).
Comparing post-concussive neurocognitive test data to normative data presents risks for under-classifying “above average” athletes
.
Archives of Clinical Neuropsychology
,
29
,
625
632
.

Schmidt
,
J. D.
,
Register-Mihalik
,
J. K.
,
Mihalik
,
J. P.
,
Kerr
,
Z. Y.
, &
Guskiewicz
,
K. M.
(
2012
).
Identifying impairments after concussion: Normative data versus individualized baselines
.
Medicine & Science in Sports & Exercise
,
44
,
1621
1628
. doi: .

Sedgwick
,
P.
(
2013
).
Limits of agreement (Bland-Altman method)
.
BMJ
,
346
. doi: .

Tsushima
,
W. T.
,
Siu
,
A. M.
,
Pearce
,
A. M.
,
Zhang
,
G.
, &
Oshiro
,
R. S.
(
2016
).
Two-year test-retest reliability in high school athletes
.
Archives of Clinical Neuropsychology
,
31
,
105
111
. doi: .

Varian
,
H.
(
2005
).
Bootstrap tutorial
.
The Mathematica Journal
,
9
,
768
775
.

Vaz
,
S.
,
Falkmer
,
T.
,
Passmore
,
A. E.
,
Parsons
,
R.
, &
Andreou
,
P.
(
2013
).
The case for using the repeatability coefficient when calculating test-retest reliability
.
PLoS ONE
,
8
, e73990. doi: .

White
,
N.
,
Flannery
,
L.
,
McClintock
,
A.
, &
Machado
,
L.
(
2019
).
Repeated computerized cognitive testing: Performance shifts and test-retest reliability in health older adults
.
Journal of Clinical and Experimental Neuropsychology
,
41
,
179
191
. doi: .

Wilk
,
C. M.
,
Gold
,
J. M.
,
Bartko
,
J. J.
,
Dickerson
,
F.
,
Fenton
,
W. S.
,
Knable
,
M.
, et al. (
2002
).
Test-retest stability of the repeatable battery for the assessment of neuropsychological status in schizophrenia
.
American Journal of Psychiatry
,
159
,
838
844
. doi: .

This article is published and distributed under the terms of the Oxford University Press, Standard Journals Publication Model (https://dbpia.nl.go.kr/journals/pages/open_access/funder_policies/chorus/standard_publication_model)