Close Menu
Healthtost
  • News
  • Mental Health
  • Men’s Health
  • Women’s Health
  • Skin Care
  • Sexual Health
  • Pregnancy
  • Nutrition
  • Fitness
  • Recommended Essentials
What's Hot

Indian Turmeric Rice (So Easy!)

July 20, 2026

Scientists discover hidden cellular network that leads to rapid gut renewal

July 20, 2026

7 Outdoorsman Personality Traits and What They Reveal

July 20, 2026
Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
Healthtost
SUBSCRIBE
  • News

    Scientists discover hidden cellular network that leads to rapid gut renewal

    July 20, 2026

    New study redefines high-risk criteria for multiple myeloma

    July 20, 2026

    Frequent TV viewing can damage brain health in the long term

    July 19, 2026

    Political polarization is causing an increase in anti-vaccine legislation across the US

    July 19, 2026

    New ctDNA blood test improves personalized prostate cancer treatment

    July 18, 2026
  • Mental Health

    I have spent the last 6 months reading hundreds of poems by young people – I was surprised to find hope, not despair

    July 17, 2026

    Is it okay to be imperfect and still be happy? 6 Challenges

    July 15, 2026

    How can you be tired but wired? Blame it on your stone age brain

    July 12, 2026

    Almost 20% of new mums have anxiety or depression, but a promising psychedelic treatment is on the horizon

    July 7, 2026

    How can ART help us improve our mental health? With 3 Ways

    July 5, 2026
  • Men’s Health

    7 Outdoorsman Personality Traits and What They Reveal

    July 20, 2026

    Considering Shockwave Therapy for ED? Here’s what you need to know

    July 18, 2026

    Does the timing of the blood test affect testosterone levels?

    July 17, 2026

    GLP-1 receptor activation is associated with lower odds of depression and bipolar disorder

    July 16, 2026

    The cost of neurophobia in Canadian medical education

    July 16, 2026
  • Women’s Health

    How 2025 Biogen Face of Fitness Jalencke Coetzee builds strength, serves others and protects her well-being

    July 20, 2026

    The O-Shot®: Beyond the Buzzwords, A Guide to Regenerative Medicine for Women’s Sexual Health

    July 19, 2026

    Understanding Breast Cancer – Life Among Women

    July 19, 2026

    5 Signs of an Unhealthy Relationship

    July 17, 2026

    Understanding withdrawal symptoms from common substances

    July 17, 2026
  • Skin Care

    Repêchage® wins AAEI’s 2026 Small Business Exporter of the Year award

    July 19, 2026

    K-Beauty for Celiac Disease and Allergic Skin: What Really Works and

    July 18, 2026

    Shea butter for hair: Benefits and uses

    July 17, 2026

    Your First Men’s Facial: What to Expect at Joanna Vargas

    July 16, 2026

    Summer skin care tips for sensitive skin – why your skin suddenly breaks out

    July 15, 2026
  • Sexual Health

    The Step-by-Step Reality of a No-Needle, No-Scalpel Vasectomy in the Labyrinth

    July 19, 2026

    Why more women are choosing hormone replacement therapy (HRT) at Maze Women’s Health

    July 19, 2026

    S*x in the Shadows of Big Tech

    July 18, 2026

    Do STD rates increase during major events like the World Cup?

    July 17, 2026

    How to Become a Sex Therapist — Sexual Health Alliance

    July 16, 2026
  • Pregnancy

    Free and Cheap Things to Do with a Baby (By Season)

    July 20, 2026

    Foods to avoid during pregnancy for a healthy mom and baby

    July 19, 2026

    What are the best multivitamins for women? – Pink stork

    July 18, 2026

    What are protein supplements during pregnancy and breastfeeding?

    July 17, 2026

    Exercise Wall Angels During Pregnancy: A Step-by-Step Guide

    July 15, 2026
  • Nutrition

    Indian Turmeric Rice (So Easy!)

    July 20, 2026

    Erythritol and Heart Health: Safe or Dangerous?

    July 19, 2026

    IM8 Review: Slip it on like Beckham?

    July 19, 2026

    5 Signs You’re Dealing With Burnout

    July 18, 2026

    Creamy tuna pasta salad with lemon and capers • Kath Eats

    July 17, 2026
  • Fitness

    Editor’s Pick: 10 brands we can’t recommend enough

    July 20, 2026

    The Best Glutathione Supplements | mindbodygreen

    July 18, 2026

    207: What Your Doctor Doesn’t Test | Thyroid, Hormones and Getting Real Answers with Ashley Cruz Arata

    July 17, 2026

    Getting stronger is corrective – Tony Gentilcore

    July 16, 2026

    7 Uplifting Emotional Benefits of Cooking

    July 16, 2026
  • Recommended Essentials
Healthtost
Home»News»GPT-4 demonstrates high accuracy in parsing multilingual medical notes
News

GPT-4 demonstrates high accuracy in parsing multilingual medical notes

healthtostBy healthtostJanuary 6, 2025No Comments6 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Reddit WhatsApp Email
Gpt 4 Demonstrates High Accuracy In Parsing Multilingual Medical Notes
Share
Facebook Twitter LinkedIn Pinterest WhatsApp Email

The study evaluates the GPT-4’s ability to process medical notes in English, Spanish and Italian, achieving physician agreement 79% of the time.

Study: The ability of Generative Pre-trained Transformer 4 (GPT-4) to analyze medical notes in three different languages: a retrospective model evaluation study. Image credit: SuPatMaN/Shutterstock.com

In a recent study published in Lancet Digital Healtha group of researchers evaluated the ability of Generative Pre-trained Transformer 4 (GPT-4) to answer predefined questions based on medical notes written in three languages ​​(English, Spanish and Italian).

Background

Medical notes contain valuable clinical knowledge, yet their unstructured narrative form poses challenges for automated analysis.

Large language models (LLMs) such as GPT-4 show promise in extracting explicit details such as medications, but often struggle with implicit understanding of contexts, crucial for nuanced medical decision making. Variability in documentation styles between providers adds to the complexity.

Existing research demonstrates the potential of LLMs for free-text medical data processing, including decoding abbreviations and extracting social determinants of health, however these studies mainly focus on English language notes.

Further research is vital to enhance the ability of LLMs to handle complex tasks, improve contextual reasoning, and assess performance in multiple languages ​​and settings.

About the study

The present retrospective model evaluation study involved eight university hospitals from four countries: the United States of America (USA), Colombia, Singapore, and Italy.

The participating institutions were part of the 4CE Consortium. They included Boston Children’s Hospital, the University of Michigan, the University of Wisconsin, the National University of Singapore, the University of Kansas Medical Center, the University of Pittsburgh Medical Center, the Universidad de Antioquia, and the Istituti Clinici Scientifici Maugeri.

The Department of Biomedical Informatics at Harvard University served as the coordinating center. Each site contributed seven de-identified medical notes written between February 1, 2020 and June 1, 2023, resulting in a total of 56 medical notes, with six sites submitting notes in English, one in Spanish, and one in Italian.

Participating sites selected notes based on proposed criteria, including patients aged 18–65 years with a diagnosis of obesity and coronavirus disease 2019 (COVID-19) at admission. Compliance with these criteria was optional.

Notes submitted included admission, progress and consultation notes, but not discharge summaries. The notes were removed in accordance with the guidelines of the US Health Insurance Portability and Accountability Act, regardless of country of origin.

The study used the GPT-4 API in Python to analyze medical notes through a predefined question-answer framework. Parameters such as temperature, top-p and frequency penalty were adjusted to optimize performance.

Physicians rated the free-text responses and indicated whether they agreed with the GPT-4 responses. They were masked in each other’s ratings but not in the GPT-4 responses.

Statistical analyzes were performed to assess agreement between the GPT-4 and physicians, exploring instances of disagreement and categorizing errors as issues of derivation, inference, or hallucinations.

Subgroup analyzes and sensitivity analyzes addressed variations in accuracy, such as differences in language and specific inclusion criteria.

The study highlighted the ability of GPT-4 to process medical notes in multiple languages, but noted challenges in inference based on context and variability in documentation styles. Data analyzes were performed in RStudio and no external funding supported the study.

Study results

A total of 56 medical records were collected from eight sites in four countries: USA, Colombia, Singapore and Italy. Of these, 42 (75%) notes were in English, seven (13%) in Italian and seven (13%) in Spanish. For each note, the GPT-4 generated responses to 14 predefined questions, resulting in 784 responses.

Among them, both physicians agreed with the GPT-4 in 622 (79%) responses, one physician agreed in 82 (11%) responses, and neither physician agreed in 80 (10%) responses. When the National University of Singapore data were excluded, agreement rates remained similar: 534 (78%) responses had double agreement, 82 (12%) had partial agreement, and 70 (10%) had no agreement.

Physicians were more likely to agree with the GPT-4 for Spanish (86/98, 88%) and Italian (82/98, 84%) notes than for English notes (454/588, 77%).

Note type or length did not affect agreement rates. In cases where only one physician agreed with the GPT-4 (82 responses), 59 (72%) disagreements arose from issues of inference, such as different interpretations of implicit information.

In one case, a physician concluded that a patient did not have COVID-19 based on a “recent infection with COVID-19” note, while the GPT-4 left the condition as undetermined. Extraction problems accounted for 8 (10%) of these disagreements, such as a physician overlooking a documented medical history that identified GPT-4. Differences in level of agreement accounted for the remaining 15 (18%) cases.

In responses where both clinicians disagreed with the GPT-4 (80 responses), inference issues were most frequent (47/80, 59%), followed by inference errors (23/80, 29%) and hallucinations (10/ 80, 13% ).

For example, GPT-4 sometimes failed to link complications such as multisystem inflammatory syndrome as related to COVID-19, a connection made by both doctors. Delusion issues included information about the construction of the GPT-4 that was not in the notes, such as falsely claiming that a patient had COVID-19 when it was not reported.

When evaluating the ability of the GPT-4 to select patients for hypothetical study enrollment based on four inclusion criteria (age, obesity, COVID-19 status, and admission note type), its sensitivity varied. GPT-4 showed high sensitivity for obesity (97%), COVID-19 (96%) and age (94%), but lower specificity for admission notes (22%).

When the acceptance note criterion was excluded, the GPT-4 accurately identified all three remaining criteria 90% of the time.

conclusions

In summary, the study showed that the GPT-4 accurately analyzed medical notes in English, Italian, and Spanish, even without a direct technique.

Surprisingly, it performed better on Italian and Spanish notes than on English, possibly due to the greater complexity of US medical notes, although note length did not affect performance. GPT-4 efficiently extracted explicit information, but its main limitation was extracting implicit details.

This aligns with previous findings that models optimized for medical tasks can overcome such challenges. While the GPT-4 excelled at identifying explicit study inclusion criteria such as age and obesity, it struggled to classify admission notes, likely due to reliance on indirect constructs.

accuracy demonstrates GPT4 high medical multilingual Notes parsing
bhanuprakash.cg
healthtost
  • Website

Related Posts

Scientists discover hidden cellular network that leads to rapid gut renewal

July 20, 2026

New study redefines high-risk criteria for multiple myeloma

July 20, 2026

Frequent TV viewing can damage brain health in the long term

July 19, 2026

Leave A Reply Cancel Reply

Don't Miss
Nutrition

Indian Turmeric Rice (So Easy!)

By healthtostJuly 20, 20260

I wasn’t looking to rediscover rice. I just wanted something with a little more personality…

Scientists discover hidden cellular network that leads to rapid gut renewal

July 20, 2026

7 Outdoorsman Personality Traits and What They Reveal

July 20, 2026

How 2025 Biogen Face of Fitness Jalencke Coetzee builds strength, serves others and protects her well-being

July 20, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
TAGS
Baby benefits body brain cancer care Day Diet disease exercise finds Fitness food Guide health healthy heart Improve Life Loss Men mental Natural Nutrition Patients People Pregnancy research reveals risk routine sex sexual Skin Skincare study Therapy Tips Top Training Treatment ways weight women Workout
About Us
About Us

Welcome to HealthTost, your trusted source for breaking health news, expert insights, and wellness inspiration. At HealthTost, we are committed to delivering accurate, timely, and empowering information to help you make informed decisions about your health and well-being.

Latest Articles

Indian Turmeric Rice (So Easy!)

July 20, 2026

Scientists discover hidden cellular network that leads to rapid gut renewal

July 20, 2026

7 Outdoorsman Personality Traits and What They Reveal

July 20, 2026
New Comments
    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Privacy Policy
    • Terms and Conditions
    • Disclaimer
    © 2026 HealthTost. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.