• Home
  • About US
  • Contact Us
  • Privacy Policy
  • Terms of Use

Archives

  • September 2026
  • August 2026
  • July 2026
  • June 2026
  • May 2026
  • April 2026
  • March 2026
  • February 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • November 2023
  • October 2023
  • September 2023
  • August 2023
  • July 2023
  • June 2023
  • May 2023
  • April 2023
  • March 2023
  • January 2023
  • December 2022
  • November 2022
  • October 2022
  • September 2022
  • August 2022
  • July 2022
  • June 2022
  • May 2022
  • March 2022
  • February 2022
  • January 2022
  • November 2019

Categories

  • Auto
  • Business
  • Climate & Earth
  • Coronavirus
  • Crypto
  • Entertainment
  • Happiness Basket
  • India
  • Learn | Unlearn | Relearn
  • Lifestyle
  • Politics
  • Research Discoveries
  • Science & Technology
  • Sports
  • Trending
  • Video
  • World
  • Home
  • About US
  • Contact Us
  • Privacy Policy
  • Terms of Use
Read Selective
  • World
  • Politics
  • India
  • Business
  • Entertainment
  • Lifestyle
  • Auto
  • Crypto
  • Coronavirus
  • Happiness Basket
  • Research Discoveries
  • Research Discoveries

Even the Best AI Chatbot Gets Health Questions Wrong 1 in 5 Times, Doctors Find

  • May 30, 2026
ChatGPT, Claude, and Gemini are among the most widely used AI apps used. (© prima91 – stock.adobe.com)

Board-Certified Physicians Put Popular LLMs Through Their Paces, and Found Real Problems

When people feel a strange pain or notice a worrying symptom, more and more of them are skipping the doctor’s office and heading straight to an AI chatbot. It’s fast, free, and available at 3 a.m. But a study suggests that convenience might come with a serious catch: even the best-performing AI gets medical questions wrong roughly one out of every five times.

In a preprint study (not yet peer-reviewed) posted online by researchers from Penn State, four popular AI chatbots were put to the test using real and imagined health concerns submitted by university students, staff, and faculty. A panel of nine board-certified physicians then graded the AI responses. Overall results were mixed: impressive enough to turn heads, but flawed enough to raise real concerns about what happens when someone acts on bad medical advice.

Nearly one in four adults under 30 already use AI monthly for health-related guidance, according to data cited in the paper. Understanding what these tools get right (and wrong) is essential.

How Researchers Tested AI Chatbots on Health Questions

Researchers organized a university-wide competition in fall 2024. A total of 34 participants were invited to query one of four AI chatbots — ChatGPT-4o, ChatGPT-3.5, Gemini-1.5 Pro, and Llama3-8b — with health-related questions they might genuinely want answered. Participants could approach the task from one of three angles: as a patient describing personal symptoms, as a medical professional seeking diagnostic help, or through an out-of-the-box track that allowed for alternative medical query scenarios, such as analyzing images of handwritten prescriptions.

Competition entries generated 212 AI responses in total. Those responses were then divided among a panel of nine board-certified physicians, each of whom graded them on four measures: how valid the information was, the quality of the information, how well the AI reasoned through the problem, and whether the response could cause harm.

Gemini-1.5 Pro produced the largest share of responses, 140 out of 212, while Llama3-8b generated only 6. That imbalance matters when comparing models directly, and the researchers acknowledged it as a limitation.

What Doctors Found When They Graded the AI Responses

Across all four AI models, about 76% of responses were rated as valid by physicians. That sounds reasonable until the math flips: nearly one in four responses didn’t make the cut. For ChatGPT-4o, the highest-performing model, validity hit 84.6%, still leaving more than 15% of answers falling short. Llama3-8b landed at the bottom, with only half its responses rated as valid.

Which type of medical question was asked also mattered. Questions about obstetrics and gynecology scored the highest for accuracy, while neurology, internal medicine, and dermatology consistently ranked lower. Neurology cases in the study often involved rare conditions that are hard to diagnose under any circumstances, while dermatology relies heavily on visual examination — something a text-based chatbot simply cannot replicate.

Prompt length turned out to be a factor, too. Very short questions and very long, detailed ones both produced weaker results. Best performance came from medium-length queries, somewhere between 60 and 250 characters. Medical professionals said in follow-up interviews that the more specific and focused the question, the better the AI tended to perform.

Adding a Medical Encyclopedia Didn’t Always Help AI Chatbots

One of the study’s more surprising results involved a technique called Retrieval-Augmented Generation, or RAG, essentially giving the AI access to a curated library of medical textbooks, clinical guidelines, and research articles from a university medical school before it generates a response. Grounding the AI in vetted medical sources should, in theory, make its answers more reliable.

Seven medical professionals were recruited to compare standard AI responses against RAG-enhanced ones, side by side. For Gemini-1.5 Pro and Llama3-8b, the medical professionals actually preferred the standard, unenhanced versions by a wide and statistically significant margin. For the ChatGPT models, there was no significant difference either way.

Researchers stopped short of declaring RAG unhelpful overall, noting that the results varied by model and that future research should explore the approach further.

Source : https://studyfinds.com/best-ai-chatbot-gets-health-questions-wrong-doctors-find/

Previous Article
  • Research Discoveries

Half Of Americans Say The Fun In Their Lives Has Disappeared

  • May 28, 2026
View Post
Next Article
  • Research Discoveries

Protein Isn’t Just For Gym-Goers: Study Links Low Intake To Physical Decline In Older Women

  • May 31, 2026
View Post
You May Also Like
View Post
  • Research Discoveries

Skipping the Gym? Walking More Controlled Asthma About as Well as Treadmill Workouts

  • September 4, 2026
View Post
  • Research Discoveries

Your Organs Are Aging On Different Timelines, Study Suggests

  • September 3, 2026
View Post
  • Research Discoveries

Legumes, Soy Linked To Lower Blood Pressure Risk

  • September 2, 2026
View Post
  • Research Discoveries

Men Resist Dieting Because It’s Too Feminine, Study Suggests

  • September 1, 2026
View Post
  • Research Discoveries

Toddlers Poisoned By Edible Drugs at Record Rates, Study Warns

  • August 30, 2026
View Post
  • Research Discoveries

Survey: 82% Of Americans Say They Live On ‘Autopilot’

  • August 28, 2026
View Post
  • Research Discoveries

Coconut Oil Jet Fuel Matches Kerosene’s Efficiency in Engine Tests

  • August 24, 2026
View Post
  • Research Discoveries

Gene Test Revealed Inherited Cancer Risk Across Three Generations of One Family

  • August 19, 2026

Recent Posts

  • Mohsin Naqvi’s Bizarre ‘India Remark’ When Asked About Pakistan’s Unending Crisis
  • Rohit Sharma’s Endgame: Gautam Gambhir, Ajit Agarkar, And A World Cup Dream
  • After backlash over ‘women shouldn’t leave home’ call, Kerala cleric clarifies
  • Wai Wai’s bhujia made from noodles lying on factory floor, FSSAI cracks down
  • BJP workers forcing BLOs to delete names: INDIA bloc’s big claim in Jharkhand
Categories
  • Auto (44)
  • Business (466)
  • Climate & Earth (16)
  • Coronavirus (18)
  • Crypto (28)
  • Entertainment (884)
  • Happiness Basket (7)
  • India (5,886)
  • Learn | Unlearn | Relearn (163)
  • Lifestyle (326)
  • Politics (98)
  • Research Discoveries (399)
  • Science & Technology (473)
  • Sports (1,000)
  • Trending (1,707)
  • Video (1)
  • World (9,823)
Read Selective

For Feedbacks, Advertisements or Any Other Concerns mail us at info@readselective.com

Pages
  • Home
  • About US
  • Contact Us
  • Privacy Policy
  • Terms of Use
Categories
  • Auto
  • Business
  • Climate & Earth
  • Coronavirus
  • Crypto
  • Entertainment
  • Happiness Basket
  • India
  • Learn | Unlearn | Relearn
  • Lifestyle
  • Politics
  • Research Discoveries
  • Science & Technology
  • Sports
  • Trending
  • Video
  • World
© 2024 Read Selective | Developed by SUGARA Technologies
  • Home
  • About US
  • Contact Us
  • Privacy Policy
  • Terms of Use

Input your search keywords and press Enter.

Go to mobile version