[Skip to Content]
Join AMIA
Menu
  • Register
  • Program Schedule
  • Speaker Search
  • My Account
  • Home
  • 2026 Annual Symposium Gallery
  • Large Language Model-Based Evaluation of the Impact of Gender on Medical Research

Custom CSS

double-click to edit, do not edit in source


S48: Words Matter Out Here: Language, Bias, and Meaning (Oral Presentations)


11/9/2026 | 3:30 PM – 4:45 PM | Cortez B
Presentation Type: Oral Presentations

Large Language Models Generate Stigmatized Language During Reasoning Over Real World Clinical Data

Presentation Type: Podium Abstract
Presentation Time: 03:30 PM - 03:42 PM

Abstract Keywords: Natural Language Processing, Large Language Models (LLMs), Artificial Intelligence, Fairness and elimination of bias, Machine Learning, Health Equity, Evaluation, Natural Language Processing
Programmatic Theme: Clinical Informatics

Large language models are increasingly used in clinical workflows, yet their reasoning processes may reveal or amplify stigmatizing language. We evaluated 95 LLMs across 35 real world clinical tasks and a clinician validated NLP system. Stigmatizing language appeared in 84 percent of LLM outputs and was more frequent in reasoning models, highlighting the need for mitigation strategies before deploying LLMs in clinical environments.

Speaker(s):
Jie Yang, PhD, FACMI, FAMIA
Harvard Medical School

Author(s):
Yutong Yang, Master of Science Student - Harvard University; Bowen Gu, MS - Brigham and Women's Hospital; David Hathaway, MD - Brigham and Women’s Hospital; Richard Wyss, MD PhD - Brigham and Women’s Hospital; Qingyu Chen, PhD - Yale University; Nan Liu, PhD - National University of Singapore; Li Zhou, MD, PhD, FACMI, FIAHSI, FAMIA - Brigham and Women's Hospital, Harvard Medical School; Kueiyu Joshua Lin, MD MSc - Brigham and Women’s Hospital; Jie Yang, PhD, FACMI, FAMIA - Harvard Medical School;
Jie Yang, PhD, FACMI, FAMIA - Harvard Medical School
Large Language Model-Based Evaluation of the Impact of Gender on Medical Research

Presentation Type: Podium Abstract
2026 Symposium LIEAF Presentation

Presentation Time: 03:42 PM - 03:54 PM

Abstract Keywords: Information Extraction, Large Language Models (LLMs), Data Mining, Artificial Intelligence, Disability, Accessibility, and Human Function, Health Equity, Real-World Evidence Generation
Programmatic Theme: Academic Informatics / LIEAF

We conducted a retrospective bibliometric analysis using a large language model to predict the genders of academic researchers. Applying our method to ten million authors publishing over 1 million medical studies since 2015, we observed greater female representation among first authors in some medical specialties and increasing proportions of female principal investigators, underscoring recent progress towards gender equity in academic medical research.

Speaker(s):
Michael Yao


Author(s):
Michael Yao;
Michael Yao -
Evaluating Open-Source Large Language Models for Annotating Social and Behavioral Determinants of Health in Clinical Text

Presentation Type: Paper - Student
Presentation Time: 03:54 PM - 04:06 PM

Abstract Keywords: Artificial Intelligence, Large Language Models (LLMs), Natural Language Processing, Information Extraction, Health Equity, Documentation Burden, Workflow, Diversity, Equity, Inclusion, and Accessibility
Programmatic Theme: Clinical Informatics

Social and behavioral determinants of health (SBDHs) shape patient outcomes but are inconsistently documented in electronic health records. Large language models (LLMs) may enable scalable SBDH extraction from unstructured clinical text, though their performance and generalizability require further investigation. We evaluated three open-source LLMs (Llama 3.1, Phi-4, Gemma 2) for SBDH annotation using a meta-prompting approach on MIMIC-SBDH benchmark dataset. Models were tested under zero-shot, few-shot, and chain-of-thought prompting. Meta-prompting generated task-specific prompts programmatically. Model annotations were compared with MIMIC-SBDH gold-standards. Performance metrics included macro-F1, precision, recall, and hallucination rate. Models performed best on binary SBDHs (e.g. community presence; macro-F1>0.94) and lowest on behavioral determinants (e.g. alcohol use; macro-F1<0.46). Few-shot and chain-of-thought prompting yielded inconsistent improvements, while hallucination rates remained low (<1%). Overall, open-source LLMs matched the performance of traditional NLP models without task-specific fine-tuning, making them promising tools for scalable SBDH extraction. Further refinement is needed before clinical integration.

Speaker(s):
Ryan McConnell, B.S.
University of Pittsburgh School of Medicine

Author(s):
Nikita Kedia, MD - New York University; Ryan McConnell, B.S. - University of Pittsburgh School of Medicine; Gilles Clermont, MD, MSc - University of Pittsburgh; Andrew King, PhD, FAMIA - University of Pittsburgh;
Ryan McConnell, B.S. - University of Pittsburgh School of Medicine
Comparison of answers to questions about tobacco cessation services in English and Spanish from different large language models

Presentation Type: Paper - Student
Presentation Time: 04:06 PM - 04:18 PM

Abstract Keywords: Artificial Intelligence, Health Equity, Large Language Models (LLMs), Quantitative Methods
Programmatic Theme: Clinical Informatics

Our aim is to compare the performance of multilingual LLMs in answering user test questions related to a tobacco cessation quitline across English and Spanish within the context of a public health program. We conducted a 4×2×2 factorial design with three factors: four LLM models, two prompt languages, and two question languages. The probability of correct language generation and program accuracy was 15% higher (RR=1.15, 95% CI: 1.04–1.28; p=.006) and 0.14 points (on a 1 to 5 Likert scale) higher (β=0.14, 95% CI: 0.06–0.23; p=.001), respectively, when the language of the question matched the language of the prompt. English-prompted LLMs presented with questions in English achieved mean Likert scores 0.36 points higher than Spanish-prompted LLMs presented with questions in Spanish (β=–0.36, 95% CI: –0.48 to –0.24; p<0.001). Model-specific differences were observed, with Llama3.1-8B and ChatGPT-4o showing the best performances under congruent conditions.

Speaker(s):
David Villarreal-Zegarra, MPH
Department of Biomedical Informatics, University of Utah

Author(s):
David Villarreal-Zegarra, MPH - Department of Biomedical Informatics, University of Utah; Andy J. King, PhD - Huntsman Cancer Institute, University of Utah, Salt Lake City, Utah, USA; Mahony Reategui Rivera, MD - University of Utah; Adam Kotter, BS - University of Utah; Alana Woodbury, BSc - Department of Biomedical Informatics, School of Medicine, University of Utah, Salt Lake City, Utah, United States; Leticia Stevens - University of Utah; Anthony Banks, MS - Huntsman Cancer Institute, University of Utah, Salt Lake City, Utah, USA.; David Wetter, PhD - University of Utah and Huntsman Cancer Institute; Paul A. Estabrooks, PhD - Department of Health & Kinesiology, University of Utah, Salt Lake City, Utah, United States.; Lindsey Potter, PhD - Huntsman Cancer Institute, University of Utah, Salt Lake City, Utah, USA.; Chelsey Schlechter, MPH, PhD - University of Utah and Huntsman Cancer Institute; Guilherme Del Fiol, MD, PhD - University of Utah;
David Villarreal-Zegarra, MPH - Department of Biomedical Informatics, University of Utah
Structured Reasoning for Social Context: Schema-Guided Chain-of-Thought Large Language Model for Identification of Social Determinants of Health in Behavioral Health Session Conversation

Presentation Type: Podium Abstract
Presentation Time: 04:18 PM - 04:30 PM

Abstract Keywords: Natural Language Processing, Large Language Models (LLMs), Information Extraction, Artificial Intelligence, Privacy and Security, Machine Learning, Health Equity, Public Health
Programmatic Theme: Clinical Informatics

Social Determinants of Health (SDoH) are often underrepresented in clinical records but frequently discussed in behavioral health session conversations. We developed a privacy-preserving large language model pipeline to annotate SDoH in transcripts. A locally deployed model with schema-guided chain-of-thought reasoning performed multi-label classification across SDoH domains, subdomains, and valence. Evaluation against expert annotations showed domain-level performance (precision 0.956, F1 0.90), demonstrating feasibility of extracting socially contextualized information from behavioral health conversations while maintaining data privacy.

Speaker(s):
Shuxuan Li, master
University of Pennsylvania

Author(s):
Shuxuan Li, master - University of Pennsylvania; Aviv Landau, PhD - University of Pennsylvania; Elizabeth Matthews, PhD - Fordham University; Lauri Goldkind, PhD - Fordham University; Jiyoun Song, PhD - University of Pennsylvania School of Nursing;
Shuxuan Li, master - University of Pennsylvania
Can Large Language Models Generate Translations That Match Human Translators for an Online Participant Recruitment Platform?

Presentation Type: Paper - Regular
Presentation Time: 04:30 PM - 04:42 PM

Abstract Keywords: Large Language Models (LLMs), Patient Engagement and Preferences, Natural Language Processing, Health Equity
Programmatic Theme: Consumer Health Informatics

This study systematically evaluated the ability of four commercial large language models (LLMs) to generate
Spanish translations for use on a clinical research recruitment website, comparing their outputs to professional
human translations using multiple quantitative metrics. While LLMs produced translations that closely matched
human reference standards in form and meaning, exact matches were uncommon and performance varied by model, prompt, and metric. Despite strong semantic similarity scores, idiomatic and subtle differences were evident. A human evaluator generally favored professional human translations, highlighting challenges for using LLMs to
convey linguistic and cultural nuance, particularly for regionally specific language. Human oversight or
professional translation may be advised where precision and cultural sensitivity are critical; however, LLM-based
translation may suffice for less sensitive applications and could offer cost advantages. Future work should optimize
prompts and benchmark LLMs against other translation technologies.

Speaker(s):
David Hanauer, MD
University of Michigan

Author(s):
David Hanauer, MD - University of Michigan; Natalie Borrego, BA - University of Michigan; James Maszatics - MICHR/University of Michigan; Carl Schulman, MD., PhD - University of Miami; Megan Haymart, MD - University of Michigan;
David Hanauer, MD - University of Michigan

Large Language Model-Based Evaluation of the Impact of Gender on Medical Research

Category

Podium Abstract

Description

Custom CSS

double-click to edit, do not edit in source

Date: Monday (11/09)
Time: 3:30 PM to 4:45 PM
Room: Cortez B

Back to Speaker Gallery
11/9/2026 04:45 PM (Central Time (US &amp; Canada))


Amia logo

Headquarters:
6218 Georgia Avenue NW, Suite #1
PMB 3077
Washington, DC 20011
Phone: 301.657.1291

© 2026 American Medical Informatics Association. All Rights Reserved.