Skip to main navigation Skip to main content

CEEM : Clinical and Experimental Emergency Medicine

OPEN ACCESS
ABOUT
BROWSE ARTICLES
FOR CONTRIBUTORS

Articles

Review Article
AI & Digital Health

From non-agentic large language models to multi-agent systems in emergency medicine: a scoping review

Hyeseong Kim1orcid , Seunghoon Jo2orcid , Min Hyuk Lim1,2orcid , Dong Hyun Choi3,4orcid
Available online: April 8, 2026
1Artificial Intelligence Graduate School, Ulsan National Institute of Science and Technology (UNIST), Ulsan, Republic of Korea
2Graduate School of Health Science and Technology, Ulsan National Institute of Science and Technology (UNIST), Ulsan, Republic of Korea
3Department of Emergency Medicine, Seoul National University Hospital, Seoul, Republic of Korea
4Department of Emergency Medicine, Seoul National University College of Medicine, Seoul, Republic of Korea
Corresponding author:  Min Hyuk Lim,
Email: limmh@unist.ac.kr
Dong Hyun Choi,
Email: donghyun369@naver.com
Received: 19 March 2026   • Revised: 3 April 2026   • Accepted: 6 April 2026
Hyeseong Kim and Seunghoon Jo contributed equally to this study as co-first authors.
  • 1,271 Views
  • 112 Download
  • 1 Crossref
  • 0 Scopus

Objective
This study aimed to conduct a scoping review of studies on non-agentic large language models (LLMs), LLM-based agents, and multi-agent systems reported in emergency medicine, and to identify current research trends and major gaps by analyzing their clinical application scope, system structures, evaluation approaches, and input data characteristics.
Methods
The Web of Science, Scopus, PubMed, and CINAHL databases were searched for literature published from March 8, 2021, to March 7, 2026. Among English full-text articles, studies addressing the application, evaluation, or benchmarking of non-agentic LLMs, LLM-based agents, or multi-agent systems in emergency medicine or the emergency department (ED) were included. Through reference tracking, 35 studies were analyzed.
Results
Of the 35 included studies, 26 were application studies, 6 were framework studies, and 3 were benchmark studies. The studies were concentrated on a limited set of tasks, including triage, diagnostic and treatment decision support, and documentation. In terms of system type, non-agentic LLMs were the most common (n=25), followed by LLM-based agents (n=7) and multi-agent systems (n=3). Inputs were predominantly text-based, and evaluation mainly relied on expert comparison, retrospective record review, vignette-based comparison, and task-specific performance metrics. In contrast, workflow-level, prospective, and safety and trustworthiness-oriented evaluation were limited.
Conclusion
LLMs in emergency medicine have shown potential for task-level decision support and documentation. However, current literature remains focused on non-agentic LLM-based task support, while studies reflecting the dynamic workflow of real EDs remain limited. Future research should expand toward workflow-aware design, operational evaluation, multimodal data integration, multi-agent–based role coordination, and safety and trustworthiness validation.

Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:

Include:

From non-agentic large language models to multi-agent systems in emergency medicine: a scoping review
Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:
Include:
From non-agentic large language models to multi-agent systems in emergency medicine: a scoping review
Close