Abstract

Objective

Assess the neuropsychology knowledge base of open-access large language models (LLMs) and inform potential applications in the field.

Method

We obtained 600 multiple-choice practice questions from the “Be Ready for ABPP in Neuropsychology” website of the American Academy of Clinical Neuropsychology. We tested OpenAI (GPT-3.5, GPT-4, o3-mini-high) and Google (Gemini 1.0, 2.0 Flash Thinking Experimental [FTE]) models in two trials (T1: used 20-question blocks, T2: single-item re-administration of incorrect questions during T1). We compared AI-estimated to actual accuracy using a paired-samples t-test, while binomial logit generalized linear mixed-effects models (GLMMs) with a random intercept for item compared LLMs and item-level predictors (domain of practice, word count, position-in-block, and higher/lower order question type). Finally, we thematically analyzed the questions missed by the top models.

Results

OpenAI o3-mini-high demonstrated the highest accuracy (T1:87.0%, T2:90.3%), followed by Gemini 2.0 FTE (T1:81.7%, T2:88.7%), GPT-4 (T1:74.0%, T2:85.5%), GPT-3.5 (T1:62.5%), and Gemini 1.0 (T1:52.3%). On average, LLMs overestimated their accuracy by 15.8% (87.3% vs. 71.5%, p < .001). In the GLMMs, specific LLM (p < .001) and practice domain (p = .045) were the only significant predictors of accuracy. Chain-of-thought “reasoning” models outperformed older models (p < .001) but displayed inaccuracies pertaining to aspects of neuropsychological testing and interpretation, diagnostic reasoning, clinical decision making, and neuroanatomy/neuroimaging.

Conclusions

Chain-of-thought “reasoning” models displayed the highest accuracy, suggesting they may have utility in neuropsychology education, research, and clinical practice. However, the LLMs displayed persistent neuropsychology content area weaknesses and tended to present inaccurate information with confidence, highlighting the need for caution when interpreting LLM output.

This article is published and distributed under the terms of the Oxford University Press, Standard Journals Publication Model (https://dbpia.nl.go.kr/pages/standard-publication-reuse-rights)
You do not currently have access to this article.