Abstract

Study Design

Using two observational methods and a within-subjects, counterbalanced design, this study aimed to determine if a computer’s hardware and software settings significantly affected reaction time (RT) on the Automated Neuropsychological Assessment Metrics (Version 4) Traumatic Brain Injury Military (ANAM4 TBI-MIL).

Methods

Three computer platforms were investigated: Platform 1—older computers recommended for ANAM4 TBI-MIL administration, Platform 2—newer computers with settings downgraded to run like the older computers, and Platform 3—newer computers with default settings. Two observational methods were used to compare measured RT to observed RT on all three platforms: 1, a high-speed video analysis to compare the timing of stimulus onset and response to the measured RT and 2, comparing a preset RT delivered by a robotic key actuator activated by optic detector to the measured RT. Additionally, healthy active duty service members (n = 169) were administered a brief version of the ANAM4 TBI-MIL battery on each of the three platforms.

Results

RT differences were observed with both the high-speed video and robotic arm analyses across all three computer platforms, with the smallest discrepancies between observed and measured RT on Platform 1, followed by Platform 2, then Platform 3. When simple reaction time (SRT) raw and standardized scores obtained from the participants were compared across platforms, statistically significant and clinically meaningful differences were seen, especially between Platforms 1 and 3.

Conclusions

A computer’s configurations have a meaningful impact on ANAM SRT scores. The difference in an individual’s performance across platforms could be misinterpreted as clinically meaningful change.

Introduction

Combat- and training-related traumatic brain injuries (TBI) are a common issue across the U.S. military. From 2000 to 2018, there were 383,947 cases of military service-related TBI, with 82.3% of those cases being mild TBI (mTBI; Defense and Veterans Brain Injury Center, 2018). In response to mounting concerns regarding military service-related TBI, the Department of Defense (DoD) mandated that all service members complete neurocognitive testing 12 months prior to deploying in order to establish baselines for post-injury assessments and assist with return to duty decision making (DoDi 6490.13). The instruction also designated the Automated Neuropsychological Assessment Metrics (Version 4) Traumatic Brain Injury Military (ANAM4 TBI-MIL) as the neurocognitive assessment tool (NCAT) to be used to meet these assessment requirements (Department of Defense, 2015).

The ANAM4 TBI-MIL is a brief, computer-based test battery, designed to assess cognitive functioning by measuring reaction time (RT) and accuracy on seven subtests purported to represent various cognitive domains that are sensitive to the effects of mTBI (Cognitive Science Research Center, 2014). Due to the automated nature of the test battery, ANAM4 TBI-MIL can be administered by a paraprofessional technician, simultaneously assess large groups of examinees, automatically generate alternate forms, and track multiple assessments over time (Arrieux, Cole, & Ahrens, 2017). The ANAM4 and ANAM4 TBI-MIL have also been widely used in clinical research in both civilian and military settings (Caplan et al., 2016; Cole, Arrieux, Ivins, Schwab, & Qashu, 2018; Nelson et al., 2016; Register-Mihalik et al., 2013). Investigations have generated a large body of literature on the ANAM4’s psychometric properties (i.e., reliability and validity), and studies have found that ANAM4 results are moderately stable over time, and are somewhat related to traditional neuropsychological measures (Cole et al., 2013; Cole et al., 2018; Nelson et al., 2016).

RT-based computerized cognitive tests like ANAM4 TBI-MIL are susceptible to measurement error, primarily because they are administered on Windows-based personal computers designed for multitasking. Myors (1999) compared RT measurements across early versions of Windows and data indicated that the operating system (OS) was less than ideal for psychological experiments. The use of external chronometry later demonstrated that different versions of Windows generated accurate RT measures when carefully configured (Chambers & Brown, 2003). Although ANAM4 TBI-MIL is verified to be compatible with various versions of the Windows OS, multiple factors such as display, peripherals, hardware settings, and multicore processors can still have an effect on RT measurement (Cernich, Brennana, Barker, & Bleiberg, 2007).

In consideration of hardware and software configuration issues, the Neurocognitive Assessment Branch (NCAB) of the DoD has standardized the computer platform on which ANAM4 TBI-MIL is administered and provides preconfigured laptops to military ANAM4 TBI-MIL testing centers. Using a standard platform also minimizes error by ensuring that data collected across different sites are consistent and use the same platform used to create the normative reference groups and baseline results. However, those without access to the standardized computer platforms might administer ANAM4 TBI-MIL on an off-the-shelf, Windows-based computer with default settings, and results might not be comparable to the normative reference groups or an individual’s baseline (i.e., predeployment).

Though the differences in RT measurement across various tablets have been published in peer reviewed psychological assessment literature (Schatz, Ybarra, & Leitner, 2015), there is no published research that compares ANAM4 TBI-MIL RT measurement across various computer platforms. This study aims to explore the issue of ANAM4 TBI-MIL RT measurement accuracy by comparing data generated from the DoD standard laptop and two other platform configurations. We anticipated that the platforms would generate statistically significant and clinically meaningful differences in RT scores. We also expected to demonstrate that a computer with default settings can be easily configured to generate ANAM4 TBI-MIL results comparable to data generated from the standard NCAB-approved platform.

Materials and Methods

ANAM4 TBI-MIL

As described earlier, the ANAM4 TBI-MIL is a computer-based cognitive assessment utilized by the U.S. military for predeployment and postconcussive injury cognitive screening assessments. The full battery consists of a questionnaire followed by seven subtests. Reaction time, accuracy, and throughput (i.e., an RT by accuracy score) are generated by the ANAM4 TBI-MIL software. The assessment may be modified to only administer a portion of the battery. For the purposes of our study, two ANAM4 TBI-MIL subtests were administered on each platform: simple reaction time (SRT) and procedural reaction time (PRO). These were chosen, as they are two of the more RT-dependent subtests of the ANAM4 TBI-MIL. However, as analyses were completed, since the focus was on RT, data across SRT and PRO were deemed redundant. Additionally, PRO includes an element of response accuracy as well as RT, and the potential influence of response accuracy on score differences was deemed beyond the scope of the current analyses. However, it should be noted that all analyses with PRO found similar results to SRT data. Thus, for parsimony we focus only on SRT data.

Computer Platforms

Three different computer platforms were evaluated to determine the accuracy of RT measurement: (a) the NCAB standardized ANAM4 TBI-MIL computer platform, a preconfigured Dell Latitude D630, which was in production in 2007; (b) Dell Latitude E6540 (in production circa 2013 and in use at Womack Army Medical Center when this study was conducted) with an adjusted basic input/out system (BIOS) intended to operate similar to the NCAB standard platform; and (c) Dell Latitude E6540 with default settings. The BIOS settings on Platform 2 were adjusted with the following parameters: multicore support (cpucore) set to “1,” disable HyperThread control (logicproc), select “not enabled” for IntelSpeedStep (speedstep), disable C-states control (cstatesctrl), disable Intel TurboBoost (turbomode), disable Windows Aero, and set the window and graphics to the best performance (Schlegel & Goss, n.d).

Reaction Time Measurement

Three prospective methods were used to evaluate RT measurement. Two observational analyses allowed comparisons of RT as measured by the ANAM4 TBI-MIL software to actual RT: (a) a high-speed video recording of stimulus display and response (described in Cernich et al., 2007) and (b) a robotic arm actuator programmed to respond to stimuli at a precise RT. The third method was a within subjects, quasi-experimental, counterbalanced design, where healthy active duty SMs were administered a brief version of the ANAM4 TBI-MIL on each platform.

Method 1: High-Speed Video

Eyak Services, LLC were subcontracted to conduct procedures described in a unpublished informational paper (Russell, Schlegel, Coldren, Brown, & Walker, n.d), where the authors assessed ANAM4 TBI-MIL RT accuracy on the DoD standard laptop (i.e., Dell D630) by comparing ANAM4 TBI-MIL data to external measurements of stimuli presentation and mouse response. Each platform was configured with a photodetector attached to the display monitor and response sensors wired into the mouse. When the monitor sensor detected stimuli, it was programmed to flash a small LED light mounted on the monitor, and when the mouse was clicked in response to the stimuli, a sensor was configured to flash a small LED light mounted on the mouse. A representative of Eyak Services, LLC completed the ANAM4 TBI-MIL SRT subtest, which consisted of 40 stimuli, whereas a high-speed camera mounted on a tripod recorded video of the generation of ANAM stimuli on the monitor, the stimuli presentation indicator light, the mouse clicks, and the mouse click indicator light. The video recording was used to determine an observed RT based on the frame number differential of the stimulus presentation indicator light and the mouse response indicator light. The measured RT data from the ANAM4 TBI-MIL software were generated for comparison to the observed RT data.

Method 2: Black Box Tool Kit and Robotic Arm Actuator

The Black Box Toolkit v2 (BBTK) is a proprietary hardware/software bundle developed to calibrate computer-based cognitive test protocols (Parsons, McMahan, & Kane, 2018). For this study, the BBTK was configured to respond to visual stimuli on the screen, detected with a photodetector attached to the monitor, at a preset interval using a programmed robotic key actuator to press the mouse (Plant, 2016). The ANAM4 TBI-MIL SRT subtest (40 trials on each run) was run four times on each computer platform while the BBTK responded to stimuli at a preset 286-millisecond RT. ANAM4 TBI-MIL software generated RTs were also captured for comparison to the BBTK preset RT.

Method 3: Healthy Participants

Data were collected from 186 healthy active duty SMs recruited from various briefings on Fort Bragg, NC. The SMs were required to be between the ages of 18–65, with no self-reported history of mTBI in the past year, no more than two mTBI in their lifetime, no neurological disorders that affect cognitive functioning, and no visual or physical limitations that affect their ability to take the assessment. Consent and data collection procedures were approved by the Womack Army Medical Center Institutional Review Board. Following consent, participants were administered a medical history questionnaire to further ensure eligibility. They were then administered three additional questionnaires and the two ANAM4 TBI-MIL subtests on the three computer platforms (i.e., taking each subtest three times). The participants completed one of the three questionnaires before taking the ANAM4 TBI-MIL subtests on each platform to allow for a brief break between administrations. The questionnaires administered were a demographic and military questionnaire to gather basic demographics and information related to military service, the Personal Health Questionnaire-15 (PHQ-15) to assess current health status, and the Neurobehavioral Symptom Inventory (NSI-22) to assess symptoms commonly associated with mTBI.

To account for potential order or practice effects, participants were randomly assigned to one of three platform testing orders in a quasi-experimental design described by Campbell and Stanley (1963; i.e., [1,2,3], [2,3,1], [3,1,2]). A priori power analyses indicated 56 participants per testing order (total n of 168) were needed to adequately power analyses, and additional participants were enrolled to account for an estimated 10% data loss. Following test administration, the ANAM4 TBI-MIL software was used to extract raw and standardized RT data for each subtest on each platform for comparison.

Data Analyses

Methods 1 and 2: Observed RT Measurements

The mean observed RT and standard deviation measured by both the high-speed video (Method 1) and BBTK (Method 2) were calculated for each platform and compared to the ANAM4 TBI-MIL software generated mean RT and standard deviation. The mean difference was calculated by subtracting the observed mean RT from the ANAM4 TBI-MIL software measured mean RT.

Method 3: Healthy Participants

Participant demographics, NSI-22, and PHQ-15 scores were calculated for the sample. ANAM4 TBI-MIL software was used to extract raw and standardized SRT data for each participant’s performance on each platform. The ANAM TBI-MIL software used the Military 2015 norms (age and gender based) to generate standardized Wechsler IQ scores (i.e., M = 100, SD = 15). Scores were excluded from the analysis if indicative of poor effort, which was defined by two or more instances of an ANAM4 TBI-MIL standardized subtest score < 70. SPSS 24 (IBM) was used to calculate means and standard deviations for demographics, NSI-22, PHQ-15, and performance on each platform.

To examine differences in SRT measurement across platform, a one-way analysis of variance (ANOVA) was performed with platform (three levels; Platforms 1, 2, or 3) as a between-subjects variable and SRT raw and standard scores as the dependent variables. To evaluate the potential impact of administration order, three one-way ANOVAs were completed for each platform (Platforms 1, 2, and 3) with administration order (three levels; administration 1, 2, or 3) as a between-subjects variable and SRT performance as the dependent variable. Post hoc pairwise comparisons were calculated to follow up significant differences. Effect sizes (ESs) were calculated using Cohen’s d and interpreted using the following criteria: small = 0.20, medium = 0.50, and large = 0.80 (Cohen, 1988). SPSS 24 was used for ANOVAs, and G-Power was used to calculate ES.

Two within-subjects analyses were completed to determine the difference in standardized scores (SS) across platforms (i.e., the difference between an individual’s score on Platforms 1 and 2, 2 and 3, and 1 and 3). First, the difference in SRT SS across platforms in standard deviation units of the normative reference group (e.g., 1 SD equals 15 SS points) was calculated to determine the number and percent of participants with SS that decreased or increased relative to preset ranges (i.e., <−2, −1.9 to 1.5, −1.4 to −1, −0.9 to −0.5, −0.4 to −0.1, etc.). Second, reliable change indices (RCI) were calculated for the difference in participants’ SS when compared across platforms (i.e., the difference between an individual’s score on Platforms 1 and 2, 2 and 3, and 1 and 3). The RCI’s were automatically calculated by the ANAM4-TBI MIL software, using the following formula described in the ANAM Military battery administration manual (Cognitive Science Research Center, 2014):
where Y1 and Y2 are the test scores at times 1 and 2, Xnormchg is the average norm group change, |${s}_{{\mathrm{X}}_1}^2$| and |${s}_{{\mathrm{X}}_2}^2$| are the time 1 and time 2 variance of the norm group, and rxx is the norm group test-retest reliability (Roebuck-Spencer, Vincent, Schlegel, & Gilliland, 2013; Cognitive Science Research Center, 2014). The number and percent of participants whose scores decreased beyond the ANAM4 recommended clinically relevant RCI threshold (−1.64 RCI) were calculated (Cognitive Science Research Center, 2014).

Results

Observed RT Measurements

Methods 1 and 2: Observed RT Measurements

Table 1 reports the observed mean RT (SD), ANAM4 TBI-MIL measured mean RT (SD), mean difference between observed and measured RT, and standard deviation for each platform for Methods 1 and 2. Though the actual observed differences were not the same between methods due to methodological differences, the pattern of scores was similar. Specifically, across both observational methods, Platform 1 had the smallest mean difference, followed by Platform 2, and Platform 3 generated considerably higher mean differences.

Table 1

Observed mean RT (SD), measured mean RT (SD) via ANAM4 TBI-MIL software, mean difference, and SD of trial by trial mean difference for Method 1 and Method 2a

PlatformObserved mean RT (SD) via high-speed videoMeasured mean RT (SD) via ANAM4 TBI-MIL softwareMean difference (SD)
1237.50 (41.50)268.50 (41.75)−31.00 (5.05)
2225.45 (59.25)262.45 (59.70)−37.00 (5.39)
3240.25 (61.25)311.00 (62.00)−70.75 (6.00)
PlatformObserved mean RT (SD) via BBTKMeasured mean RT (SD) via ANAM4 TBI-MIL softwareMean differencea
1286.00 (0.00)330.02 (4.92)−44.02
2286.00 (0.00)337.66 (5.43)−51.66
3286.00 (0.00)373.05 (6.03)−87.05

Note: ANAM4 TBI-MIL = Automated Neuropsychological Assessment Metrics (Version 4) Traumatic Brain Injury Military; BBTK = Black Box Toolkit; RT = reaction time.

aSD of mean difference not reported for BBTK because there is no variability in the observed mean RT in BBTK.

Method 3: Healthy Participants

Of the 186 healthy participants, data were removed from analyses for 8 participants due to test proctoring errors/data storage malfunction, and 9 participants for invalid data, as defined in the methods. This resulted in a final sample of 169 examinees generating SRT subtest data for each of the three platforms. Table 2 reports the demographics and self-report data, which demonstrates that participants were similar to active duty military demographics, as they were healthy predominantly young, white, men, with <16 years of education.

Table 2

Demographic characteristics and symptom scale scores for participants across platforms

nM (SD)Range
Age16930.09 (7.94)18–52
Sex
 Male151
 Female18
Race
 White83
 Black37
 Hispanic33
 Native American3
 Pacific Islander4
 Asian3
 Other6
Education
 General educational development5
 High school graduate35
 Some college61
 Associate degree26
 Bachelor’s degree25
 Postbachelor’s degree17
NSI-22 total scorea1696.65 (3.42)0–57
PHQ-15 total scoreb1693.12 (9.47)0–14

Note: NSI= neurobehavioral Symptom Inventory; PHQ = personal health questionnaire.

aTotal scores ≥24 may indicate clinically elevated symptoms.

bTotal scores ≥5, ≥10, and ≥15 represent mild, moderate, and severe levels of somatization, respectively.

See Table 3 for the mean, standard deviation, and 95th confidence interval for ANAM4 TBI-MIL SRT scores for each platform. Participants’ scores on Platform 1 generated the fastest raw RT measurements (raw M = 240.09; standardized M = 107.41), followed by Platform 2 (raw M = 250.00; standardized M = 104.35), and then Platform 3 (raw M = 282.65; standardized M = 93.77). Results of the ANOVA performed to examine between-subjects differences across Platforms 1, 2, and 3 indicated a statistically significant difference in scores across platforms (F  (2–504) = 98.32; p ≤ .001). Post hoc pairwise comparisons revealed that Platform 3 generated significantly slower data than Platforms 1 and 2, whereas Platform 2 generated significantly slower data than Platform 1. Cohen’s d ESs were large for comparisons of Platform 1 versus 3 (d = 1.41) and 2 versus 3 (d = 1.09), and small to medium for Platform 1 versus 2 (d = 0.35). The three one-way ANOVAs to examine differences in performance based on administration order revealed that SRT performance on each platform did not statistically differ based on order of administration (Platform 1, p = .082; Platform 2, p = .109; Platform 3, p = .141).

Table 3

SRT subtest raw and SS and ANOVA p value

GroupANOVA
Subtest scorePlatform 1 (n = 169)Platform 2 (n = 169)Platform 3 (n = 169)p value
SRT rawM240.09250.00282.65<.001
SD30.0533.1336.36
95% CI LL UL235.52 244.65244.97 255.03277.13 288.88
SRT SSM107.41104.3593.77<.001
SD8.768.8410.45
95% CI LL UL106.62 109.17103.09 105.8991.74 95.06

Note: SRT = simple reaction time; SS = standardized scores; ANOVA = analysis of variance; CI = confidence interval; LL = lower limit; UL = upper limit.

Table 4 reports the difference in participants’ performance across platforms in standard deviation units, which demonstrates that when comparing the difference between Platforms 1 and 3, 94.68% of the participants’ scores were lower on Platform 3, with 46.16% of the participants demonstrated a difference of 1 SD or more, and 15.39% a had a difference of 1.5 SD or more. A similar, though less pronounced, pattern was seen between Platforms 2 and 3. That is, 93.50% of the participants’ scores were lower on Platform 3, 29.00% with a 1 SD difference, and 7.70% 1.5 SD difference. Although most participants’ had lower scores on Platform 3 than on Platform 1, there were some exceptions as 5.33% of participants’ scores were better on Platform 3 than on Platform 1. Additionally, when comparing Platforms 2 and 3, 6.51% of participants’ scores were better on Platform 3 than on Platform 2.

Table 4

Within-subjects SS difference across platforms in standard deviation units

SRT subtest
Platforms 1 and 2 differencePlatforms 2 and 3 differencePlatforms 1 and 3 difference
SD unitn%n%n%
≤−221.1842.3795.33
−1.9 to −1.50095.331710.06
−1.4 to −1116.513621.305230.77
−0.9 to −0.54023.676538.465532.54
−0.4 to −0.15331.364426.042715.98
052.960021.18
0.1 to 0.44426.0463.5552.96
0.5 to 0.9127.1021.1810.59
1 to 1.421.1821.1810.59
≥1.50010.5900

Note: SS = standardized scores; SRT = simple reaction time.

Table 5 includes data from reliable change analyses. When evaluating the difference between a participant's performance on Platforms 1 and 3 relative to the RCI clinically significant cut point (RCI ≤ −1.64), 20.71% of the sample’s standardized RT scores were beyond the RCI threshold. When comparing Platforms 2 and 3, 14.20% of participants’ standardized RT scores exceed the cut point. Only 5.92% of the sample’s standardized RT scores were beyond the cut point when comparing Platforms 1 and 2.

Table 5

Number (and percent of total sample) of participants who demonstrated a reliable decrease in scores between platforms (per RCI threshold of −1.64)

Platforms 1 and 2 difference ≤ −1.64Platforms 2 and 3 difference ≤ −1.64Platforms 1–3 difference ≤ −1.64
n%n%n%
105.922414.203520.71

Note: RCI = reliable change indices.

Discussion

This study investigated the differences in RT measurement across computer platforms and to determine if any observed differences were clinically meaningful. We compared ANAM4 TBI-MIL SRT scores across three computer platforms: 1, the NCAB approved Dell D630; 2, a newer Dell E6540 with BIOS downgraded to perform more like the Dell D630; and 3, a newer Dell E6540 with default settings. RT was measured with two observational measures (a high-speed video analysis and BBTK), and a group of healthy participants took the ANAM4 TBI-MIL SRT and PRO subtests across the three computer platforms, though only SRT data was analyzed.

Per our expectations, Methods 1 and 2 consistently demonstrated differences in RT across the three platforms. In both Methods 1 and 2, the mean difference between observed and measured RT was largest on Platform 3, followed by Platform 2, and then Platform 1. Additionally in Methods 1 and 2, Platforms 1 and 2 demonstrated nearly half the mean difference in RT than Platform 3.

When healthy participants’ SRT performance was compared across each platform, statistically significant scores were identified. The difference between Platforms 1 and 3 were largest, followed by the difference between Platforms 2 and 3, and to a much lesser degree between Platforms 1 and 2. Because the order of administration was counterbalanced, the impact of administration order was investigated, with no indication of order effects. This is consistent with past research demonstrating little to no order effects when taking multiple computerized cognitive assessments in one session (Cole, Arrieux, Dennison, & Ivins, 2016). Additionally, within-subjects analyses demonstrated that 46.16% of the participants’ SRT performance was at least 1 SD lower on Platform 3 than Platform 1 and 20.71% of scores exceeded the ANAM4 TBI-MIL’s threshold for reliable change. This was observed, though to a less degree, when comparing Platforms 2–3, as 29% of the participants’ scores were at least 1 SD lower on Platform 3 than on Platform 2 and 14.20% of scores exceeded the RCI threshold. For comparison, a study comparing pre- and postdeployment ANAM4 performance in a sample of 8,002 military SMs’ without mTBI demonstrated that 3.70% of the participants demonstrated reliable change across sessions (Roebuck-Spensor, et al., 2013). The data generated from this study demonstrate that a participant could be misclassified as having a significant change in ANAM4 TBI-MIL SRT scores by using a different computer platform, and the number of participants whose scores decreased beyond RCI criteria was nearly five times higher than what was found by Roebuck-Spencer et al. (2013).

The differences identified across the three methods were due to the computer hardware and settings, especially the multicore processor and graphics card settings of the Dell E6540 when used with default settings. These differences demonstrate that the computer and its configuration could have a clinically meaningful impact on the SRT data. If an individual is taking the test on a computer configured differently than the computer the normative data were collected on and/or the computer upon which they took a baseline assessment, they may be deemed to have a clinically meaningful change in scores that is in fact due to error introduced by hardware and software settings. Though we focused only on ANAM4 TBI-MIL, it is reasonable to think other computerized cognitive assessments are vulnerable to similar issues. With examiners potentially using a multitude of different computer platforms that have not undergone the type of scrutiny we applied to our platforms, the impact on obtained scores is unknown.

It is important to acknowledge that most computerized cognitive testing companies, including the company responsible for ANAM4 (Vista Life Sciences), provide recommendations for computer configurations and/or scoring corrections. Additionally, there are plans to upgrade the DoD’s fleet of ANAM4 TBI-MIL computers and adjust the norms and subsequent scores accordingly to account for RT differences on newer equipment. Our data do suggest configuring a computer’s BIOS so it runs like the computers the normative data were collected on resulted in smaller RT differences and SS.

Limitations

The current study is not a comprehensive evaluation of the impact of computer hardware and software settings on ANAM4 scores or other computerized cognitive tests. Additionally, data were collected only in a healthy, educated, primarily male military population. Therefore, the generalizability to other cognitive tests, computers, computer configurations, and populations may be limited. Because the SRT subtest was the main focus of the study, other ANAM4 TBI-MIL subtests and its overall composite score may not be as affected. However, many computerized cognitive assessments, especially ANAM4, rely heavily on RT measurement and therefore we believe the findings in this study are noteworthy for consideration of the broader implications across other subtests and tests. We also acknowledge that our within-subjects analyses were largely descriptive in nature and do not account for order of administration. However, our methodology and study design did not lend itself to the utilization of a conventional repeated measures analysis. Our goal for this study was to demonstrate if (a) there were significant differences across computer platforms on RT measures and (b) if those differences were clinically meaningful. We believe that we have accomplished this and going beyond this would be preemptive and beyond the scope of our intent. Future analyses could incorporate more advanced statistics to obtain a more granular understanding of the impact of platform, order of administration, practice effects, demographic variables, and so forth on RT measurement. Finally, many cognitive testing companies provide guidelines for computer configurations and scoring corrections, which may negate many of the issues we identified. However, these guidelines are only effective at mitigating scoring errors if properly developed and applied. Regardless, we feel that this study provides an important demonstration as to why the type of computer and the configuration of that computer is important when doing computerized cognitive assessment of RT.

Conclusion

The type of computer and the subsequent hardware and software settings can have a significant and clinically meaningful impact on RT measurement from a computerized cognitive assessment. Those using computerized cognitive tests should be aware of the computer settings used for normative databases and should configure their computers to perform as closely to this as possible. Often this can be accomplished by following the guidelines put forth by the test development companies. However, it is important that the test developers provide guidelines that are current and user friendly. Alternatively, the awareness of specific RT differences can allow for post hoc scoring corrections that minimize the error introduced by hardware and software settings. Future studies may want to investigate differences in RT between a computer-based test and tablet-based test, as tablets are becoming increasingly more common in clinical settings. Clinical populations should also be investigated to determine if the RT differences observed in this study are different in the context of an mTBI or other medical conditions that affect RT.

Disclaimer

The views expressed in this manuscript are those of the authors and do not necessarily represent the official policy or position of the Defense Health Agency, Department of Defense, or any other U.S. government agency. This work was prepared under Contract HT0014-19-C-0004 with DHA Contracting Office (CO-NCR) HT0014 and, therefore, is defined as U.S. Government work under Title 17 U.S.C.§101.Per Title 17 U.S.C.§105, copyright protection is not available for any work of the U.S. Government. Formore information, please contact [email protected].

Funding

This research was supported by funding from Defense and Veterans Brain Injury Center, operated by General Dynamics Information Technology for the U.S. Defense Health Agency under Contract number W91YTZ-13-C-0015/ HT0014-19-C-0004 and a grant from the Army Advanced Medical Technology Initiative, grant number 5661.

Conflicts of Interest

The authors have no conflicts of interest to disclose.

Acknowledgments

The authors thank Angelica Ahrens and J. Wes McGee of the Defense and Veterans Brain Injury Center, Y. Sammy Choi, Chief of The Department of Clinical Investigations at Womack Army Medical Center, and Bob Schlegel with Eyak Services, LLC for their contributions to this study.

References

Arrieux
,
J. P.
,
Cole
,
W. R.
, &
Ahrens
,
A. P.
(
2017
).
A review of the validity of computerized neurocognitive assessment tools in mild traumatic brain injury assessment
.
Concussion
,
2
(
1
),
CNC31
.

Campbell
,
D. T.
, &
Stanely
,
J. C.
(
1963
).
Experimental and quasi-experimental designs for research
.
Dallas
:
Houghton Mifflin Company
.

Caplan
,
B.
,
Bogner
,
J.
,
Brenner
,
L.
,
Campbell
,
J. S.
,
Johnson
,
D.
,
Young
,
E.
, et al. (
2016
).
Reliable change estimates for assessing recovery from concussion using the ANAM4 TBI-MIL
.
Journal of Head Trauma Rehabilitation
,
31
(
5
),
329
338
.

Cernich
,
A. N.
,
Brennana
,
D. M.
,
Barker
,
L. M.
, &
Bleiberg
,
J.
(
2007
).
Sources of error in computerized neuropsychological assessment
.
Archives of Clinical Neuropsychology
,
22
,
39
48
.

Chambers
,
C. D.
, &
Brown
,
M.
(
2003
).
Timing accuracy under Microsoft windows revealed through external chronometry
.
Behavior Research Methods, Instruments, & Computers.
,
35
(
1
),
96
108
.

Cognitive Science Research Center
(
2014
).
ANAM military battery administration manual
.
Norman, OK, USA: University of Oklahoma
.

Cohen
,
J.
(
1988
).
Statistical power analysis for the behavioral sciences
.
Hillside, NJ
:
Lawrence Erlbaum Associates
.

Cole
,
W. R.
,
Arrieux
,
J. P.
,
Dennison
,
E. M.
, &
Ivins
,
B. J.
(
2016
).
The impact of administration order in studies of computerized neurocognitive assessment tools (NCATs)
.
Journal of Clinical and Experimental Neuropsychology
,
39
(
1
),
35
45
.

Cole
,
W. R.
,
Arrieux
,
J. P.
,
Ivins
,
B. J.
,
Schwab
,
K. A.
, &
Qashu
,
F. M.
(
2018
).
A comparison of four computerized neurocognitive assessment tools to a traditional neuropsychological test battery in service members with and without mild traumatic brain injury
.
Archives of Clinical Neuropsychology
,
33
(
1
),
102
119
. doi: .

Cole
,
W. R.
,
Arrieux
,
J. P.
,
Schwab
,
K.
,
Ivins
,
B. J.
,
Qashu
,
F. M.
,
Lewis
,
S. C.
, et al. (
2013
).
Test–retest reliability of four computerized neurocognitive assessment tools in an active duty military population
.
Archives of Clinical Neuropsychology
,
28
,
732
742
.

Defense and Veterans Brain Injury Center
. (
2018
).
DoD worldwide numbers for traumatic brain injury, worldwide - totals
.
Retrieved from
 https://dvbic.dcoe.mil/files/tbi-numbers/worldwide-totals-2000-2018Q1-total_jun-21-2018_v1.0_2018-07-26_0.pdf

Department of Defense
. (
2015
).
DoD policy guidance for management of mild traumatic brain injury/concussion in the deployed setting
.
Instruction 6490.13. Retrieved from
 https://www.hsdl.org/?view&did=800023.

Myors
,
B.
(
1999
).
Timing accuracy of PC programs running under DOS and windows
.
Behavior Research Methods, Instruments, & Computers
,
31
(
2
),
322
328
.

Nelson
,
L. D.
,
LaRoche
,
A. A.
,
Pfaller
,
A. Y.
,
Lerner
,
E. B.
,
Hammeke
,
T. A.
,
Randolph
,
C.
, et al. (
2016
).
Prospective, head-to-head study of three computerized neurocognitive assessment tools (CNTs): Reliability and validity for the assessment of sport-related concussion
.
Journal of the International Neuropsychological Society
,
22
,
24
37
.

Parsons
,
T. D.
,
McMahan
,
T.
, &
Kane
,
R.
(
2018
).
Practice parameters facilitating adoption of advanced technologies for enhancing neuropsychological assessment paradigms
.
The Clinical Neuropsychologist
,
32
(
1
),
16
41
.

Plant
,
R. R.
(
2016
).
The black box toolkit v2 user guide
.
Retrieved from
 www.blackboxtoolkit.com.

Register-Mihalik
,
J. K.
,
Guskiewicz
,
K. M.
,
Mihalik
,
J. P.
,
Schmidt
,
J. D.
,
Kerr
,
Z. Y.
, &
MCcrea
,
M. A.
(
2013
).
Reliable change, sensitivity, and specificity of a multidimensional concussion assessment battery: Implications for caution in clinical practice
.
The Journal of Head Trauma Rehabilitation
,
28
(
4
),
274
283
.

Roebuck-Spencer
,
T. M.
,
Vincent
,
A. S.
,
Schlegel
,
R. E.
, &
Gilliland
,
K.
(
2013
).
Evidence for added value of baseline testing in computer-based cognitive assessment
.
Journal of Athletic Training
,
48
(
4
),
499
505
.

Russell
,
M. L.
,
Schlegel
,
R. E.
,
Coldren
,
R.
,
Brow
,
T. A.
, &
Walker
,
D. E.
(n.d).
Accuracy of computer-based neuropsychological assessment timing: The need for external calibration
.
Unpublished manuscript
.

Schatz
,
P.
,
Ybarra
,
V.
, &
Leitner
,
D.
(
2015
).
Validating the accuracy of reaction time assessment on computer-based tablet devices
.
Assessment
,
22
(
4
),
405
410
.

Schlegel
,
R. E.
, &
Goss
,
J.
(n.d).
ANAM video response timing accuracy analysis final report
.
Anchorage, AK
:
Authors
.

This article is published and distributed under the terms of the Oxford University Press, Standard Journals Publication Model (https://dbpia.nl.go.kr/journals/pages/open_access/funder_policies/chorus/standard_publication_model)