English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Evaluating LLMs for Personalized Health Answers Using Personal Health Records

Forum topic · 小凯 · 2026-05-21

Summary

This study evaluates how well Gemini 3.0 Flash answers patient health questions when given access to Personal Health Records (PHRs). Researchers tested 2,257 user queries drawn from three sources—short web search queries, longer chatbot-style questions, and questions patients ask their care teams—matched against 1,945 de-identified records. Responses were generated under three conditions: no PHR context, a basic summary of demographics, conditions, and medications, and full clinical notes. Evaluation used both automated raters on the full set and clinician ratings on a 95-query subset, employing the SHARP framework plus a new framework targeting PHR-specific errors. Results show significant improvements in helpfulness across all question types when PHR data is included (p < 0.001), along with gains in safety, accuracy, relevance, and personalization. The new PHR framework highlights LLM gaps including temporal disorientation and rare confabulations, offering a foundation for monitoring LLM performance on patient data and motivating further work to help users understand their health records.

Evaluating the Utility of Personal Health Records in Personalized Health AI

Research area: cs.AI Authors: Rory Sayres, Kejia Chen, Ayush Jain Published: 2026-05-21 arXiv: 2505.01252

Overview

Patient-managed Personal Health Records (PHRs) promise to empower patients to better understand their health. However, the information contained in these records is complex and can hinder insight. This study assesses the potential of large language models (specifically Gemini 3.0 Flash) to provide helpful answers to user health queries when given clinical data from PHRs as context.

Methodology

  • Queries: 2,257 user queries drawn from three distributions representing different patient question styles:
  • Short web search queries
  • Longer questions derived from chatbot conversation templates
  • Questions patients asked their healthcare team (patient calls)
  • Records: Queries were matched with de-identified PHRs drawn from a pool of 1,945 records.
  • Conditions: Gemini generated responses under three contexts:
  • 1. No PHR context 2. Basic summary of demographics, conditions, and medications 3. Full, extensive clinical notes

    Evaluation

  • Frameworks:
  • SHARP, an existing rating framework
  • A newly developed framework targeting specific error modes in PHR interpretation
  • Raters:
  • Automated raters evaluated the full query set
  • Clinician ratings were collected on a subset (n=95)
  • Both rater groups had access to the full PHR context
  • Key Findings

  • Significant improvements in answer helpfulness across all question types when PHR data was provided (p < 0.001, paired t-test).
  • Potential gains observed in safety, accuracy, relevance, and personalization of responses.
  • The new PHR evaluation framework identified gaps in LLM understanding, including:
  • Temporal disorientation when handling complex records
  • Rare but meaningful confabulations

Conclusion

The results suggest that PHR data has potential to help people across a wide range of user needs. The framework introduced in this work provides a foundation for monitoring gaps in LLM answers grounded in PHR context. The authors motivate further research to assess and realize the benefits of helping users understand their health records.

--- *Automatically collected on 2026-05-21*

Tags

#personal-health-records#large-language-models#healthcare-ai#gemini#clinical-nlp#patient-empowerment#evaluation-framework#arXiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620519