-
PDF
- Split View
-
Views
-
Cite
Cite
Jacques P Arrieux, Brittney L Roberson, Katie N Russell, Brian J Ivins, Wesley R Cole, An Investigation of the Accuracy of Reaction Time Measurements on ANAM4 TBI-MIL Across Three Computer Platforms, Archives of Clinical Neuropsychology, Volume 35, Issue 7, October 2020, Pages 1145–1153, https://doi.org/10.1093/arclin/acaa032
Close - Share Icon Share
Abstract
Using two observational methods and a within-subjects, counterbalanced design, this study aimed to determine if a computer’s hardware and software settings significantly affected reaction time (RT) on the Automated Neuropsychological Assessment Metrics (Version 4) Traumatic Brain Injury Military (ANAM4 TBI-MIL).
Three computer platforms were investigated: Platform 1—older computers recommended for ANAM4 TBI-MIL administration, Platform 2—newer computers with settings downgraded to run like the older computers, and Platform 3—newer computers with default settings. Two observational methods were used to compare measured RT to observed RT on all three platforms: 1, a high-speed video analysis to compare the timing of stimulus onset and response to the measured RT and 2, comparing a preset RT delivered by a robotic key actuator activated by optic detector to the measured RT. Additionally, healthy active duty service members (n = 169) were administered a brief version of the ANAM4 TBI-MIL battery on each of the three platforms.
RT differences were observed with both the high-speed video and robotic arm analyses across all three computer platforms, with the smallest discrepancies between observed and measured RT on Platform 1, followed by Platform 2, then Platform 3. When simple reaction time (SRT) raw and standardized scores obtained from the participants were compared across platforms, statistically significant and clinically meaningful differences were seen, especially between Platforms 1 and 3.
A computer’s configurations have a meaningful impact on ANAM SRT scores. The difference in an individual’s performance across platforms could be misinterpreted as clinically meaningful change.
Introduction
Combat- and training-related traumatic brain injuries (TBI) are a common issue across the U.S. military. From 2000 to 2018, there were 383,947 cases of military service-related TBI, with 82.3% of those cases being mild TBI (mTBI; Defense and Veterans Brain Injury Center, 2018). In response to mounting concerns regarding military service-related TBI, the Department of Defense (DoD) mandated that all service members complete neurocognitive testing 12 months prior to deploying in order to establish baselines for post-injury assessments and assist with return to duty decision making (DoDi 6490.13). The instruction also designated the Automated Neuropsychological Assessment Metrics (Version 4) Traumatic Brain Injury Military (ANAM4 TBI-MIL) as the neurocognitive assessment tool (NCAT) to be used to meet these assessment requirements (Department of Defense, 2015).
The ANAM4 TBI-MIL is a brief, computer-based test battery, designed to assess cognitive functioning by measuring reaction time (RT) and accuracy on seven subtests purported to represent various cognitive domains that are sensitive to the effects of mTBI (Cognitive Science Research Center, 2014). Due to the automated nature of the test battery, ANAM4 TBI-MIL can be administered by a paraprofessional technician, simultaneously assess large groups of examinees, automatically generate alternate forms, and track multiple assessments over time (Arrieux, Cole, & Ahrens, 2017). The ANAM4 and ANAM4 TBI-MIL have also been widely used in clinical research in both civilian and military settings (Caplan et al., 2016; Cole, Arrieux, Ivins, Schwab, & Qashu, 2018; Nelson et al., 2016; Register-Mihalik et al., 2013). Investigations have generated a large body of literature on the ANAM4’s psychometric properties (i.e., reliability and validity), and studies have found that ANAM4 results are moderately stable over time, and are somewhat related to traditional neuropsychological measures (Cole et al., 2013; Cole et al., 2018; Nelson et al., 2016).
RT-based computerized cognitive tests like ANAM4 TBI-MIL are susceptible to measurement error, primarily because they are administered on Windows-based personal computers designed for multitasking. Myors (1999) compared RT measurements across early versions of Windows and data indicated that the operating system (OS) was less than ideal for psychological experiments. The use of external chronometry later demonstrated that different versions of Windows generated accurate RT measures when carefully configured (Chambers & Brown, 2003). Although ANAM4 TBI-MIL is verified to be compatible with various versions of the Windows OS, multiple factors such as display, peripherals, hardware settings, and multicore processors can still have an effect on RT measurement (Cernich, Brennana, Barker, & Bleiberg, 2007).
In consideration of hardware and software configuration issues, the Neurocognitive Assessment Branch (NCAB) of the DoD has standardized the computer platform on which ANAM4 TBI-MIL is administered and provides preconfigured laptops to military ANAM4 TBI-MIL testing centers. Using a standard platform also minimizes error by ensuring that data collected across different sites are consistent and use the same platform used to create the normative reference groups and baseline results. However, those without access to the standardized computer platforms might administer ANAM4 TBI-MIL on an off-the-shelf, Windows-based computer with default settings, and results might not be comparable to the normative reference groups or an individual’s baseline (i.e., predeployment).
Though the differences in RT measurement across various tablets have been published in peer reviewed psychological assessment literature (Schatz, Ybarra, & Leitner, 2015), there is no published research that compares ANAM4 TBI-MIL RT measurement across various computer platforms. This study aims to explore the issue of ANAM4 TBI-MIL RT measurement accuracy by comparing data generated from the DoD standard laptop and two other platform configurations. We anticipated that the platforms would generate statistically significant and clinically meaningful differences in RT scores. We also expected to demonstrate that a computer with default settings can be easily configured to generate ANAM4 TBI-MIL results comparable to data generated from the standard NCAB-approved platform.
Materials and Methods
ANAM4 TBI-MIL
As described earlier, the ANAM4 TBI-MIL is a computer-based cognitive assessment utilized by the U.S. military for predeployment and postconcussive injury cognitive screening assessments. The full battery consists of a questionnaire followed by seven subtests. Reaction time, accuracy, and throughput (i.e., an RT by accuracy score) are generated by the ANAM4 TBI-MIL software. The assessment may be modified to only administer a portion of the battery. For the purposes of our study, two ANAM4 TBI-MIL subtests were administered on each platform: simple reaction time (SRT) and procedural reaction time (PRO). These were chosen, as they are two of the more RT-dependent subtests of the ANAM4 TBI-MIL. However, as analyses were completed, since the focus was on RT, data across SRT and PRO were deemed redundant. Additionally, PRO includes an element of response accuracy as well as RT, and the potential influence of response accuracy on score differences was deemed beyond the scope of the current analyses. However, it should be noted that all analyses with PRO found similar results to SRT data. Thus, for parsimony we focus only on SRT data.
Computer Platforms
Three different computer platforms were evaluated to determine the accuracy of RT measurement: (a) the NCAB standardized ANAM4 TBI-MIL computer platform, a preconfigured Dell Latitude D630, which was in production in 2007; (b) Dell Latitude E6540 (in production circa 2013 and in use at Womack Army Medical Center when this study was conducted) with an adjusted basic input/out system (BIOS) intended to operate similar to the NCAB standard platform; and (c) Dell Latitude E6540 with default settings. The BIOS settings on Platform 2 were adjusted with the following parameters: multicore support (cpucore) set to “1,” disable HyperThread control (logicproc), select “not enabled” for IntelSpeedStep (speedstep), disable C-states control (cstatesctrl), disable Intel TurboBoost (turbomode), disable Windows Aero, and set the window and graphics to the best performance (Schlegel & Goss, n.d).
Reaction Time Measurement
Three prospective methods were used to evaluate RT measurement. Two observational analyses allowed comparisons of RT as measured by the ANAM4 TBI-MIL software to actual RT: (a) a high-speed video recording of stimulus display and response (described in Cernich et al., 2007) and (b) a robotic arm actuator programmed to respond to stimuli at a precise RT. The third method was a within subjects, quasi-experimental, counterbalanced design, where healthy active duty SMs were administered a brief version of the ANAM4 TBI-MIL on each platform.
Method 1: High-Speed Video
Eyak Services, LLC were subcontracted to conduct procedures described in a unpublished informational paper (Russell, Schlegel, Coldren, Brown, & Walker, n.d), where the authors assessed ANAM4 TBI-MIL RT accuracy on the DoD standard laptop (i.e., Dell D630) by comparing ANAM4 TBI-MIL data to external measurements of stimuli presentation and mouse response. Each platform was configured with a photodetector attached to the display monitor and response sensors wired into the mouse. When the monitor sensor detected stimuli, it was programmed to flash a small LED light mounted on the monitor, and when the mouse was clicked in response to the stimuli, a sensor was configured to flash a small LED light mounted on the mouse. A representative of Eyak Services, LLC completed the ANAM4 TBI-MIL SRT subtest, which consisted of 40 stimuli, whereas a high-speed camera mounted on a tripod recorded video of the generation of ANAM stimuli on the monitor, the stimuli presentation indicator light, the mouse clicks, and the mouse click indicator light. The video recording was used to determine an observed RT based on the frame number differential of the stimulus presentation indicator light and the mouse response indicator light. The measured RT data from the ANAM4 TBI-MIL software were generated for comparison to the observed RT data.
Method 2: Black Box Tool Kit and Robotic Arm Actuator
The Black Box Toolkit v2 (BBTK) is a proprietary hardware/software bundle developed to calibrate computer-based cognitive test protocols (Parsons, McMahan, & Kane, 2018). For this study, the BBTK was configured to respond to visual stimuli on the screen, detected with a photodetector attached to the monitor, at a preset interval using a programmed robotic key actuator to press the mouse (Plant, 2016). The ANAM4 TBI-MIL SRT subtest (40 trials on each run) was run four times on each computer platform while the BBTK responded to stimuli at a preset 286-millisecond RT. ANAM4 TBI-MIL software generated RTs were also captured for comparison to the BBTK preset RT.
Method 3: Healthy Participants
Data were collected from 186 healthy active duty SMs recruited from various briefings on Fort Bragg, NC. The SMs were required to be between the ages of 18–65, with no self-reported history of mTBI in the past year, no more than two mTBI in their lifetime, no neurological disorders that affect cognitive functioning, and no visual or physical limitations that affect their ability to take the assessment. Consent and data collection procedures were approved by the Womack Army Medical Center Institutional Review Board. Following consent, participants were administered a medical history questionnaire to further ensure eligibility. They were then administered three additional questionnaires and the two ANAM4 TBI-MIL subtests on the three computer platforms (i.e., taking each subtest three times). The participants completed one of the three questionnaires before taking the ANAM4 TBI-MIL subtests on each platform to allow for a brief break between administrations. The questionnaires administered were a demographic and military questionnaire to gather basic demographics and information related to military service, the Personal Health Questionnaire-15 (PHQ-15) to assess current health status, and the Neurobehavioral Symptom Inventory (NSI-22) to assess symptoms commonly associated with mTBI.
To account for potential order or practice effects, participants were randomly assigned to one of three platform testing orders in a quasi-experimental design described by Campbell and Stanley (1963; i.e., [1,2,3], [2,3,1], [3,1,2]). A priori power analyses indicated 56 participants per testing order (total n of 168) were needed to adequately power analyses, and additional participants were enrolled to account for an estimated 10% data loss. Following test administration, the ANAM4 TBI-MIL software was used to extract raw and standardized RT data for each subtest on each platform for comparison.
Data Analyses
Methods 1 and 2: Observed RT Measurements
The mean observed RT and standard deviation measured by both the high-speed video (Method 1) and BBTK (Method 2) were calculated for each platform and compared to the ANAM4 TBI-MIL software generated mean RT and standard deviation. The mean difference was calculated by subtracting the observed mean RT from the ANAM4 TBI-MIL software measured mean RT.
Method 3: Healthy Participants
Participant demographics, NSI-22, and PHQ-15 scores were calculated for the sample. ANAM4 TBI-MIL software was used to extract raw and standardized SRT data for each participant’s performance on each platform. The ANAM TBI-MIL software used the Military 2015 norms (age and gender based) to generate standardized Wechsler IQ scores (i.e., M = 100, SD = 15). Scores were excluded from the analysis if indicative of poor effort, which was defined by two or more instances of an ANAM4 TBI-MIL standardized subtest score < 70. SPSS 24 (IBM) was used to calculate means and standard deviations for demographics, NSI-22, PHQ-15, and performance on each platform.
To examine differences in SRT measurement across platform, a one-way analysis of variance (ANOVA) was performed with platform (three levels; Platforms 1, 2, or 3) as a between-subjects variable and SRT raw and standard scores as the dependent variables. To evaluate the potential impact of administration order, three one-way ANOVAs were completed for each platform (Platforms 1, 2, and 3) with administration order (three levels; administration 1, 2, or 3) as a between-subjects variable and SRT performance as the dependent variable. Post hoc pairwise comparisons were calculated to follow up significant differences. Effect sizes (ESs) were calculated using Cohen’s d and interpreted using the following criteria: small = 0.20, medium = 0.50, and large = 0.80 (Cohen, 1988). SPSS 24 was used for ANOVAs, and G-Power was used to calculate ES.
Results
Observed RT Measurements
Methods 1 and 2: Observed RT Measurements
Table 1 reports the observed mean RT (SD), ANAM4 TBI-MIL measured mean RT (SD), mean difference between observed and measured RT, and standard deviation for each platform for Methods 1 and 2. Though the actual observed differences were not the same between methods due to methodological differences, the pattern of scores was similar. Specifically, across both observational methods, Platform 1 had the smallest mean difference, followed by Platform 2, and Platform 3 generated considerably higher mean differences.
Observed mean RT (SD), measured mean RT (SD) via ANAM4 TBI-MIL software, mean difference, and SD of trial by trial mean difference for Method 1 and Method 2a
| Platform . | Observed mean RT (SD) via high-speed video . | Measured mean RT (SD) via ANAM4 TBI-MIL software . | Mean difference (SD) . |
|---|---|---|---|
| 1 | 237.50 (41.50) | 268.50 (41.75) | −31.00 (5.05) |
| 2 | 225.45 (59.25) | 262.45 (59.70) | −37.00 (5.39) |
| 3 | 240.25 (61.25) | 311.00 (62.00) | −70.75 (6.00) |
| Platform | Observed mean RT (SD) via BBTK | Measured mean RT (SD) via ANAM4 TBI-MIL software | Mean differencea |
| 1 | 286.00 (0.00) | 330.02 (4.92) | −44.02 |
| 2 | 286.00 (0.00) | 337.66 (5.43) | −51.66 |
| 3 | 286.00 (0.00) | 373.05 (6.03) | −87.05 |
| Platform | Observed mean RT (SD) via high-speed video | Measured mean RT (SD) via ANAM4 TBI-MIL software | Mean difference (SD) |
|---|---|---|---|
| 1 | 237.50 (41.50) | 268.50 (41.75) | −31.00 (5.05) |
| 2 | 225.45 (59.25) | 262.45 (59.70) | −37.00 (5.39) |
| 3 | 240.25 (61.25) | 311.00 (62.00) | −70.75 (6.00) |
| Platform | Observed mean RT (SD) via BBTK | Measured mean RT (SD) via ANAM4 TBI-MIL software | Mean difference |
| 1 | 286.00 (0.00) | 330.02 (4.92) | −44.02 |
| 2 | 286.00 (0.00) | 337.66 (5.43) | −51.66 |
| 3 | 286.00 (0.00) | 373.05 (6.03) | −87.05 |
Note: ANAM4 TBI-MIL = Automated Neuropsychological Assessment Metrics (Version 4) Traumatic Brain Injury Military; BBTK = Black Box Toolkit; RT = reaction time.
aSD of mean difference not reported for BBTK because there is no variability in the observed mean RT in BBTK.
Method 3: Healthy Participants
Of the 186 healthy participants, data were removed from analyses for 8 participants due to test proctoring errors/data storage malfunction, and 9 participants for invalid data, as defined in the methods. This resulted in a final sample of 169 examinees generating SRT subtest data for each of the three platforms. Table 2 reports the demographics and self-report data, which demonstrates that participants were similar to active duty military demographics, as they were healthy predominantly young, white, men, with <16 years of education.
Demographic characteristics and symptom scale scores for participants across platforms
| . | n . | M (SD) . | Range . |
|---|---|---|---|
| Age | 169 | 30.09 (7.94) | 18–52 |
| Sex | |||
| Male | 151 | ||
| Female | 18 | ||
| Race | |||
| White | 83 | ||
| Black | 37 | ||
| Hispanic | 33 | ||
| Native American | 3 | ||
| Pacific Islander | 4 | ||
| Asian | 3 | ||
| Other | 6 | ||
| Education | |||
| General educational development | 5 | ||
| High school graduate | 35 | ||
| Some college | 61 | ||
| Associate degree | 26 | ||
| Bachelor’s degree | 25 | ||
| Postbachelor’s degree | 17 | ||
| NSI-22 total scorea | 169 | 6.65 (3.42) | 0–57 |
| PHQ-15 total scoreb | 169 | 3.12 (9.47) | 0–14 |
| n | M (SD) | Range | |
|---|---|---|---|
| Age | 169 | 30.09 (7.94) | 18–52 |
| Sex | |||
| Male | 151 | ||
| Female | 18 | ||
| Race | |||
| White | 83 | ||
| Black | 37 | ||
| Hispanic | 33 | ||
| Native American | 3 | ||
| Pacific Islander | 4 | ||
| Asian | 3 | ||
| Other | 6 | ||
| Education | |||
| General educational development | 5 | ||
| High school graduate | 35 | ||
| Some college | 61 | ||
| Associate degree | 26 | ||
| Bachelor’s degree | 25 | ||
| Postbachelor’s degree | 17 | ||
| NSI-22 total score | 169 | 6.65 (3.42) | 0–57 |
| PHQ-15 total score | 169 | 3.12 (9.47) | 0–14 |
Note: NSI= neurobehavioral Symptom Inventory; PHQ = personal health questionnaire.
aTotal scores ≥24 may indicate clinically elevated symptoms.
bTotal scores ≥5, ≥10, and ≥15 represent mild, moderate, and severe levels of somatization, respectively.
See Table 3 for the mean, standard deviation, and 95th confidence interval for ANAM4 TBI-MIL SRT scores for each platform. Participants’ scores on Platform 1 generated the fastest raw RT measurements (raw M = 240.09; standardized M = 107.41), followed by Platform 2 (raw M = 250.00; standardized M = 104.35), and then Platform 3 (raw M = 282.65; standardized M = 93.77). Results of the ANOVA performed to examine between-subjects differences across Platforms 1, 2, and 3 indicated a statistically significant difference in scores across platforms (F (2–504) = 98.32; p ≤ .001). Post hoc pairwise comparisons revealed that Platform 3 generated significantly slower data than Platforms 1 and 2, whereas Platform 2 generated significantly slower data than Platform 1. Cohen’s d ESs were large for comparisons of Platform 1 versus 3 (d = 1.41) and 2 versus 3 (d = 1.09), and small to medium for Platform 1 versus 2 (d = 0.35). The three one-way ANOVAs to examine differences in performance based on administration order revealed that SRT performance on each platform did not statistically differ based on order of administration (Platform 1, p = .082; Platform 2, p = .109; Platform 3, p = .141).
| . | . | Group . | ANOVA . | ||
|---|---|---|---|---|---|
| Subtest score | Platform 1 (n = 169) | Platform 2 (n = 169) | Platform 3 (n = 169) | p value | |
| SRT raw | M | 240.09 | 250.00 | 282.65 | <.001 |
| SD | 30.05 | 33.13 | 36.36 | ||
| 95% CI LL UL | 235.52 244.65 | 244.97 255.03 | 277.13 288.88 | ||
| SRT SS | M | 107.41 | 104.35 | 93.77 | <.001 |
| SD | 8.76 | 8.84 | 10.45 | ||
| 95% CI LL UL | 106.62 109.17 | 103.09 105.89 | 91.74 95.06 | ||
| Group | ANOVA | ||||
|---|---|---|---|---|---|
| Subtest score | Platform 1 (n = 169) | Platform 2 (n = 169) | Platform 3 (n = 169) | p value | |
| SRT raw | M | 240.09 | 250.00 | 282.65 | <.001 |
| SD | 30.05 | 33.13 | 36.36 | ||
| 95% CI LL UL | 235.52 244.65 | 244.97 255.03 | 277.13 288.88 | ||
| SRT SS | M | 107.41 | 104.35 | 93.77 | <.001 |
| SD | 8.76 | 8.84 | 10.45 | ||
| 95% CI LL UL | 106.62 109.17 | 103.09 105.89 | 91.74 95.06 | ||
Note: SRT = simple reaction time; SS = standardized scores; ANOVA = analysis of variance; CI = confidence interval; LL = lower limit; UL = upper limit.
Table 4 reports the difference in participants’ performance across platforms in standard deviation units, which demonstrates that when comparing the difference between Platforms 1 and 3, 94.68% of the participants’ scores were lower on Platform 3, with 46.16% of the participants demonstrated a difference of 1 SD or more, and 15.39% a had a difference of 1.5 SD or more. A similar, though less pronounced, pattern was seen between Platforms 2 and 3. That is, 93.50% of the participants’ scores were lower on Platform 3, 29.00% with a 1 SD difference, and 7.70% 1.5 SD difference. Although most participants’ had lower scores on Platform 3 than on Platform 1, there were some exceptions as 5.33% of participants’ scores were better on Platform 3 than on Platform 1. Additionally, when comparing Platforms 2 and 3, 6.51% of participants’ scores were better on Platform 3 than on Platform 2.
| . | SRT subtest . | |||||
|---|---|---|---|---|---|---|
| Platforms 1 and 2 difference | Platforms 2 and 3 difference | Platforms 1 and 3 difference | ||||
| SD unit | n | % | n | % | n | % |
| ≤−2 | 2 | 1.18 | 4 | 2.37 | 9 | 5.33 |
| −1.9 to −1.5 | 0 | 0 | 9 | 5.33 | 17 | 10.06 |
| −1.4 to −1 | 11 | 6.51 | 36 | 21.30 | 52 | 30.77 |
| −0.9 to −0.5 | 40 | 23.67 | 65 | 38.46 | 55 | 32.54 |
| −0.4 to −0.1 | 53 | 31.36 | 44 | 26.04 | 27 | 15.98 |
| 0 | 5 | 2.96 | 0 | 0 | 2 | 1.18 |
| 0.1 to 0.4 | 44 | 26.04 | 6 | 3.55 | 5 | 2.96 |
| 0.5 to 0.9 | 12 | 7.10 | 2 | 1.18 | 1 | 0.59 |
| 1 to 1.4 | 2 | 1.18 | 2 | 1.18 | 1 | 0.59 |
| ≥1.5 | 0 | 0 | 1 | 0.59 | 0 | 0 |
| SRT subtest | ||||||
|---|---|---|---|---|---|---|
| Platforms 1 and 2 difference | Platforms 2 and 3 difference | Platforms 1 and 3 difference | ||||
| SD unit | n | % | n | % | n | % |
| ≤−2 | 2 | 1.18 | 4 | 2.37 | 9 | 5.33 |
| −1.9 to −1.5 | 0 | 0 | 9 | 5.33 | 17 | 10.06 |
| −1.4 to −1 | 11 | 6.51 | 36 | 21.30 | 52 | 30.77 |
| −0.9 to −0.5 | 40 | 23.67 | 65 | 38.46 | 55 | 32.54 |
| −0.4 to −0.1 | 53 | 31.36 | 44 | 26.04 | 27 | 15.98 |
| 0 | 5 | 2.96 | 0 | 0 | 2 | 1.18 |
| 0.1 to 0.4 | 44 | 26.04 | 6 | 3.55 | 5 | 2.96 |
| 0.5 to 0.9 | 12 | 7.10 | 2 | 1.18 | 1 | 0.59 |
| 1 to 1.4 | 2 | 1.18 | 2 | 1.18 | 1 | 0.59 |
| ≥1.5 | 0 | 0 | 1 | 0.59 | 0 | 0 |
Note: SS = standardized scores; SRT = simple reaction time.
Table 5 includes data from reliable change analyses. When evaluating the difference between a participant's performance on Platforms 1 and 3 relative to the RCI clinically significant cut point (RCI ≤ −1.64), 20.71% of the sample’s standardized RT scores were beyond the RCI threshold. When comparing Platforms 2 and 3, 14.20% of participants’ standardized RT scores exceed the cut point. Only 5.92% of the sample’s standardized RT scores were beyond the cut point when comparing Platforms 1 and 2.
Number (and percent of total sample) of participants who demonstrated a reliable decrease in scores between platforms (per RCI threshold of −1.64)
| Platforms 1 and 2 difference ≤ −1.64 . | Platforms 2 and 3 difference ≤ −1.64 . | Platforms 1–3 difference ≤ −1.64 . | |||
|---|---|---|---|---|---|
| n . | % . | n . | % . | n . | % . |
| 10 | 5.92 | 24 | 14.20 | 35 | 20.71 |
| Platforms 1 and 2 difference ≤ −1.64 | Platforms 2 and 3 difference ≤ −1.64 | Platforms 1–3 difference ≤ −1.64 | |||
|---|---|---|---|---|---|
| n | % | n | % | n | % |
| 10 | 5.92 | 24 | 14.20 | 35 | 20.71 |
Note: RCI = reliable change indices.
Discussion
This study investigated the differences in RT measurement across computer platforms and to determine if any observed differences were clinically meaningful. We compared ANAM4 TBI-MIL SRT scores across three computer platforms: 1, the NCAB approved Dell D630; 2, a newer Dell E6540 with BIOS downgraded to perform more like the Dell D630; and 3, a newer Dell E6540 with default settings. RT was measured with two observational measures (a high-speed video analysis and BBTK), and a group of healthy participants took the ANAM4 TBI-MIL SRT and PRO subtests across the three computer platforms, though only SRT data was analyzed.
Per our expectations, Methods 1 and 2 consistently demonstrated differences in RT across the three platforms. In both Methods 1 and 2, the mean difference between observed and measured RT was largest on Platform 3, followed by Platform 2, and then Platform 1. Additionally in Methods 1 and 2, Platforms 1 and 2 demonstrated nearly half the mean difference in RT than Platform 3.
When healthy participants’ SRT performance was compared across each platform, statistically significant scores were identified. The difference between Platforms 1 and 3 were largest, followed by the difference between Platforms 2 and 3, and to a much lesser degree between Platforms 1 and 2. Because the order of administration was counterbalanced, the impact of administration order was investigated, with no indication of order effects. This is consistent with past research demonstrating little to no order effects when taking multiple computerized cognitive assessments in one session (Cole, Arrieux, Dennison, & Ivins, 2016). Additionally, within-subjects analyses demonstrated that 46.16% of the participants’ SRT performance was at least 1 SD lower on Platform 3 than Platform 1 and 20.71% of scores exceeded the ANAM4 TBI-MIL’s threshold for reliable change. This was observed, though to a less degree, when comparing Platforms 2–3, as 29% of the participants’ scores were at least 1 SD lower on Platform 3 than on Platform 2 and 14.20% of scores exceeded the RCI threshold. For comparison, a study comparing pre- and postdeployment ANAM4 performance in a sample of 8,002 military SMs’ without mTBI demonstrated that 3.70% of the participants demonstrated reliable change across sessions (Roebuck-Spensor, et al., 2013). The data generated from this study demonstrate that a participant could be misclassified as having a significant change in ANAM4 TBI-MIL SRT scores by using a different computer platform, and the number of participants whose scores decreased beyond RCI criteria was nearly five times higher than what was found by Roebuck-Spencer et al. (2013).
The differences identified across the three methods were due to the computer hardware and settings, especially the multicore processor and graphics card settings of the Dell E6540 when used with default settings. These differences demonstrate that the computer and its configuration could have a clinically meaningful impact on the SRT data. If an individual is taking the test on a computer configured differently than the computer the normative data were collected on and/or the computer upon which they took a baseline assessment, they may be deemed to have a clinically meaningful change in scores that is in fact due to error introduced by hardware and software settings. Though we focused only on ANAM4 TBI-MIL, it is reasonable to think other computerized cognitive assessments are vulnerable to similar issues. With examiners potentially using a multitude of different computer platforms that have not undergone the type of scrutiny we applied to our platforms, the impact on obtained scores is unknown.
It is important to acknowledge that most computerized cognitive testing companies, including the company responsible for ANAM4 (Vista Life Sciences), provide recommendations for computer configurations and/or scoring corrections. Additionally, there are plans to upgrade the DoD’s fleet of ANAM4 TBI-MIL computers and adjust the norms and subsequent scores accordingly to account for RT differences on newer equipment. Our data do suggest configuring a computer’s BIOS so it runs like the computers the normative data were collected on resulted in smaller RT differences and SS.
Limitations
The current study is not a comprehensive evaluation of the impact of computer hardware and software settings on ANAM4 scores or other computerized cognitive tests. Additionally, data were collected only in a healthy, educated, primarily male military population. Therefore, the generalizability to other cognitive tests, computers, computer configurations, and populations may be limited. Because the SRT subtest was the main focus of the study, other ANAM4 TBI-MIL subtests and its overall composite score may not be as affected. However, many computerized cognitive assessments, especially ANAM4, rely heavily on RT measurement and therefore we believe the findings in this study are noteworthy for consideration of the broader implications across other subtests and tests. We also acknowledge that our within-subjects analyses were largely descriptive in nature and do not account for order of administration. However, our methodology and study design did not lend itself to the utilization of a conventional repeated measures analysis. Our goal for this study was to demonstrate if (a) there were significant differences across computer platforms on RT measures and (b) if those differences were clinically meaningful. We believe that we have accomplished this and going beyond this would be preemptive and beyond the scope of our intent. Future analyses could incorporate more advanced statistics to obtain a more granular understanding of the impact of platform, order of administration, practice effects, demographic variables, and so forth on RT measurement. Finally, many cognitive testing companies provide guidelines for computer configurations and scoring corrections, which may negate many of the issues we identified. However, these guidelines are only effective at mitigating scoring errors if properly developed and applied. Regardless, we feel that this study provides an important demonstration as to why the type of computer and the configuration of that computer is important when doing computerized cognitive assessment of RT.
Conclusion
The type of computer and the subsequent hardware and software settings can have a significant and clinically meaningful impact on RT measurement from a computerized cognitive assessment. Those using computerized cognitive tests should be aware of the computer settings used for normative databases and should configure their computers to perform as closely to this as possible. Often this can be accomplished by following the guidelines put forth by the test development companies. However, it is important that the test developers provide guidelines that are current and user friendly. Alternatively, the awareness of specific RT differences can allow for post hoc scoring corrections that minimize the error introduced by hardware and software settings. Future studies may want to investigate differences in RT between a computer-based test and tablet-based test, as tablets are becoming increasingly more common in clinical settings. Clinical populations should also be investigated to determine if the RT differences observed in this study are different in the context of an mTBI or other medical conditions that affect RT.
Disclaimer
The views expressed in this manuscript are those of the authors and do not necessarily represent the official policy or position of the Defense Health Agency, Department of Defense, or any other U.S. government agency. This work was prepared under Contract HT0014-19-C-0004 with DHA Contracting Office (CO-NCR) HT0014 and, therefore, is defined as U.S. Government work under Title 17 U.S.C.§101.Per Title 17 U.S.C.§105, copyright protection is not available for any work of the U.S. Government. Formore information, please contact [email protected].
Funding
This research was supported by funding from Defense and Veterans Brain Injury Center, operated by General Dynamics Information Technology for the U.S. Defense Health Agency under Contract number W91YTZ-13-C-0015/ HT0014-19-C-0004 and a grant from the Army Advanced Medical Technology Initiative, grant number 5661.
Conflicts of Interest
The authors have no conflicts of interest to disclose.
Acknowledgments
The authors thank Angelica Ahrens and J. Wes McGee of the Defense and Veterans Brain Injury Center, Y. Sammy Choi, Chief of The Department of Clinical Investigations at Womack Army Medical Center, and Bob Schlegel with Eyak Services, LLC for their contributions to this study.