[Skip to Content]
Join AMIA
Menu
  • Register
  • Program Schedule
  • Speaker Search
  • My Account
  • Home
  • 2024 Annual Symposium Gallery
  • A systematic evaluation of large language models for biomedical natural language processing: benchmarks, baselines, and recommendations

Custom CSS

double-click to edit, do not edit in source

A systematic evaluation of large language models for biomedical natural language processing: benchmarks, baselines, and recommendations

Presentation Time: 04:45 PM - 05:00 PM

Abstract Keywords: Natural Language Processing, Large Language Models (LLMs), Deep Learning
Primary Track: Applications
Programmatic Theme: Clinical Informatics

Despite the potential of Large Language Models in biomedicine, they lack baseline performance, benchmarks, and recommendations for using LLMs in the biomedical domain. This study makes three contributions. First, it undertakes a comprehensive evaluation to establish the baseline performance of LLMs (GPT-3.5, GPT-4, and LLaMA) across 12 BioNLP datasets encompassing six distinct extractive and generative tasks. Second, we conducted thorough manual validation collectively over thousands of sample outputs in total. Third, the study offers valuable suggestions for the effective use of LLMs in BioNLP applications.

Speaker(s):
Qingyu Chen, PhD
Yale University

A systematic evaluation of large language models for biomedical natural language processing: benchmarks, baselines, and recommendations

Category

Podium Abstract

Description

Custom CSS

double-click to edit, do not edit in source



Date: Monday (11/11)
Time: 04:45 PM to 05:00 PM
Room: Franciscan A

Back to Speaker Gallery
11/11/2024 05:00 PM (Pacific Time (US & Canada))
Amia logo

Headquarters:
6218 Georgia Avenue NW, Suite #1
PMB 3077
Washington, DC 20011
Phone: 301.657.1291

© 2026 American Medical Informatics Association. All Rights Reserved.