Abstract

Objective

We examined whether GPT-4o, a widely used large language model (LLM), could produce age- and education-appropriate versions of complex pediatric traumatic brain injury case descriptions, while preserving clinical accuracy and emotional tone.

Methods

Five cases were adapted into four audience scenarios. Text complexity was assessed via Flesch–Kincaid (FKS), Gunning Fog, and SMOG indices. Clinical human experts rated text fidelity and emotional appropriateness on a 3-point scale.

Results

Original texts showed very high complexity (FKS 18.2–20.5), equivalent to 18–20 years of education. Adaptations for parents with high school education were often over-simplified (FKS 4.75–7.1), while versions for 12-year-olds were well-matched (FKS ~5–6). Texts for 8-year-olds had FKS scores of 4.0–6.8 (above grade 2–3 targets) and reduced fidelity (scores 1–2). Emotional tone was consistently rated appropriate across all audiences.

Conclusion

Clinicians may use LLMs to draft explanations, but must carefully review and tailor them.

This article is published and distributed under the terms of the Oxford University Press, Standard Journals Publication Model (https://dbpia.nl.go.kr/pages/standard-publication-reuse-rights)
You do not currently have access to this article.