
An AI system just diagnosed sick patients better than the doctors treating them, inside a realistic hospital simulation. Another wrote a full scientific paper on its own, submitted it to an academic conference, and got it accepted by human reviewers. A third built data-analysis software that beat every method researchers had spent years perfecting. None of this is a thought experiment. It happened in 2026, described across three separate studies published in Nature.
A new analysis in Artificial Intelligence & Environment reviewed all three studies together, and its conclusion is blunt: science itself is changing in fundamental ways.
Still, the review is careful not to oversell the moment. These systems fail, often in strange ways, and the ethical questions they raise are far from resolved.
AI Systems Are Already Outperforming Human Scientists and Doctors
ERA, developed by Google researchers, writes the specialized code scientists use to analyze data. It reads research ideas, generates code to test them, and rewrites the code until it finds the best-performing version. Across six scientific tasks, including COVID-19 hospital admission forecasts, ERA outperformed the best human-developed methods, even beating the CDC’s own hospitalization model. Given two existing methods, it often creates a hybrid beating both, in one biological task by 14%.
Sakana AI’s system, called the AI Scientist and first released in 2024, goes further still. Given a topic, it forms a hypothesis, writes code, runs experiments, produces figures, and drafts a complete paper with references. It even conducts its own peer review, using an automated reviewer that slightly exceeded human reviewers in balanced accuracy, 69% versus 66%. An updated version scored 6.33 out of 10 at a 2025 academic workshop, placing it in the top 45% of 43 submissions, the first fully AI-generated paper to pass human peer review.
MIRA operates inside a simulated hospital, navigating electronic health records, ordering tests, and generating diagnoses. Tested across 574 emergency cases, it reached 88.9% diagnostic accuracy, versus 78.1% for board-certified physicians and 71.1% for a mixed-experience group, a statistically significant gap.
These AI Systems Hallucinate, Miscalculate, and Fail Often
ERA can write code that runs without errors, looks correct, and produces results that are physically impossible. A related system tested on astronomical data produced outputs that violated basic laws of physics without triggering any internal warning, what researchers call “silent errors,” a confident wrong answer more dangerous than an obvious crash.
Sakana’s system has flaws just as well documented. Of three submissions, only one was accepted; the other two were rejected for underdeveloped ideas, coding mistakes, or duplicated figures. It has invented citations that do not exist and, once, altered its own code to extend a time limit rather than solve the given problem. Its authors write plainly that “none met the higher bar for a main conference publication.”
MIRA’s diagnostic edge comes with real caveats. Accuracy varied widely by condition, reaching 98.6% for appendicitis but only 72.4% for pneumonia. It also ordered blood tests far more often than physicians, without clear evidence the extra testing made its process more efficient. And the dataset used to test it may have overlapped with the underlying model’s training data, inflating its apparent accuracy.
Research on multi-agent AI systems more broadly, cited in the review, found failure rates from 41% to 87%, with the most common problems being repetitive loops, actions mismatched to a system’s own reasoning, and an inability to know when to stop.
Human Scientists Are Not Being Replaced, Just Reassigned
Despite everything these systems can do, the review’s central argument is that human scientists are not obsolete. Every one of them worked inside limits a person set. A human decided what counted as a good result, filtered which outputs were worth pursuing, or built the simulation the system operated in.
That raises a deeper question the review keeps circling back to: what these systems actually understand, versus what they can produce. AlphaFold predicts protein shapes with extraordinary precision, but does it understand why proteins fold that way? The review argues AI-driven science is shifting from scientists forming theories and testing them toward patterns in massive datasets generating conclusions researchers interpret afterward.
Source: https://studyfinds.com/ai-writes-research-papers-make-diagnoses/

