Generative AI can summarize thousands of research papers in seconds. It can explain complex concepts, draft literature reviews, and help researchers explore unfamiliar topics. But science has a higher standard than most other fields.
A useful answer is not enough. Scientists need answers that can be traced back to evidence, reproduced through experiments, and challenged through peer review. This is where general-purpose Gen AI systems begin to struggle.
The problem is not that AI cannot process scientific information. It can process more information than any individual researcher. The problem is that science is not only about finding information. It is about understanding what information is reliable, what evidence is missing, and what conclusions can actually be supported.
Gen AI is becoming better at producing answers. But science needs systems that can produce defensible knowledge.
In this article, we will look at where generative AI is already failing scientific research and what R&D teams need to change in the way they use AI-assisted research. Let’s get started!

AI Hallucinations Are Creating a New Problem for Scientific Peer Review
Peer review has always been designed to catch human mistakes such as wrong calculation, an overextended conclusion, or a citation pointing to the wrong year of a real paper.
It was not built to catch a citation to a paper that was never written at all.
That is what makes AI-generated fabrication harder to spot. A model can invent a plausible title, attach familiar researcher names, add realistic publication details, and produce something that looks completely normal in a reference list.
NeurIPS 2025 showed how easily this can slip through.
A study by researcher Samar Ansari found that among 5,290 accepted papers, each reviewed by three to five expert researchers, at least 53 contained fabricated citations. Around two-thirds were classified as “total fabrications,” meaning the cited work did not exist at all rather than being a real source with incorrect details.
That matters because this happened at one of the world’s most competitive AI conferences, where reviewers are more likely than most researchers to understand how large language models fail.
If fabricated citations can pass through that process, the risk is likely harder to manage in disciplines where reviewers are not actively looking for AI-generated references.
A separate analysis of ICLR 2026 submissions found more than 50 hallucinated citations in papers under review, including some papers that reviewers had rated highly.
The problem is not simply that AI can make mistakes.
It is that those mistakes can look credible enough to survive expert review.
Fabricated Citations Are Already Spreading Through Scientific Research
This is not a future risk. It is already part of the published literature.
A team at Columbia University’s Data Science Institute, led by Maxim Topaz, examined 2.5 million biomedical papers and around 97 million references. Their results, published in The Lancet in May 2026, identified 4,046 fabricated citations across roughly 2,810 papers.
One example shows why the problem is difficult to catch.
In a urology paper, 18 of 30 checked references were fabricated. Yet each one matched the paper’s narrow surgical topic closely enough to look credible during a normal read.
The more concerning signal is how quickly the rate has increased. In 2023, roughly 1 in every 2,828 papers examined contained a fabricated reference. By 2025, that had risen to around 1 in 458. In the first seven weeks of 2026, it reached roughly 1 in 277.

That is about a twelvefold increase in two years, broadly overlapping with the period when generative AI tools became common in academic writing workflows.
And the Columbia study is not the only evidence.
A separate audit by researchers from Cornell, UCLA, UC Berkeley, and Tsinghua University examined 111 million references across 2.5 million papers from arXiv, bioRxiv, SSRN, and PubMed Central.
Their conservative estimate was that at least 146,932 hallucinated citations appeared across those four repositories in 2025 alone.
The researchers also noted that their method likely undercounted the true total.
The pattern is just as important as the number.
This is not only a handful of papers filled with fake references. The problem is often much harder to detect: otherwise legitimate papers containing one or two fabricated citations among dozens of real ones.
A reference list can therefore look normal even when part of the evidence base is not.
The same audit found that hallucinated citations appeared more often in fields adopting AI rapidly, were more common in papers from smaller or earlier-career teams, and disproportionately attributed credit to already prominent researchers.
So fabricated citations do more than weaken individual papers. They can also distort which researchers and ideas appear most authoritative in a field.

AI-Generated Citations Are Not Just an Academic Research Problem
If this sounds like something that only research journals need to worry about, Deloitte’s experience in Australia suggests otherwise.
In 2025, Deloitte Australia was commissioned by the country’s Department of Employment and Workplace Relations to conduct an independent review of its welfare compliance IT system. The contract was worth nearly AU$440,000.
After the report was published, University of Sydney researcher Chris Rudge identified multiple fabricated references.
The report cited academic papers that did not exist. It also misquoted a federal court judgment, including a statement attributed to a judge who had never made it.
Deloitte later confirmed that parts of the report had been produced using Azure OpenAI’s GPT-4o and agreed to repay the final installment of the contract.
The department maintained that the report’s main findings and recommendations remained unchanged.
For R&D teams, that is not really the important part.
The important part is that a major consulting firm, working on a paid government engagement and operating with its own internal review processes, still allowed fabricated citations and an invented legal quotation to reach the client.
The failure was not that AI had been used. The failure was treating AI-assisted research as if its supporting evidence had already been verified. The same risk exists inside R&D.
A literature scan, technology assessment, competitive landscape, or innovation report can look polished and complete while still containing evidence nobody has actually checked.
AI Can Find the Right Papers and Still Misread the Science
Suppose every citation generated by an AI system is correct. Every paper exists. Every author is real. Every source can be opened. There is still another problem.
Science is not simply a collection of facts waiting to be retrieved and summarized. It is a network of competing findings, incomplete evidence, conflicting methods, experimental conditions, and unanswered questions. And summarization tends to smooth those differences away.
Consider a common research situation. One study finds that a compound improves a biological marker. Another finds no meaningful effect. A third finds that the effect only appears under very specific conditions.
A researcher who has spent years working in that area will not simply average those findings into one conclusion. They will ask why the studies disagree.
Were the populations different? Were the doses different? Was the endpoint measured differently? Was the result statistically significant but biologically weak?
A general-purpose language model is designed to produce a coherent answer. That creates a natural pressure to resolve disagreement rather than preserve it. But in research, the disagreement is often the most important part. It may tell you where the mechanism breaks down, where the evidence is weakest, or where the next experiment should begin.
The Limits of Generative AI in Technology and Scientific Discovery
Scientific opportunity does not always sit inside the most obvious cluster of related papers. Some of the most valuable ideas come from connecting things that initially look unrelated. An imaging technique developed in one field may solve a measurement problem in another. A manufacturing process used in packaging may address a scale-up challenge in biotechnology. A material developed for aerospace may turn out to have useful properties in medical devices.
Finding those opportunities requires more than matching similar language across papers. Someone still has to judge whether the transfer makes technical sense.
Can it scale? Does the regulatory environment allow it? Does economics work? Is the mechanism genuinely transferable, or does it only look similar at the language level?
General-purpose AI can surface connections. It is much weaker at deciding which connections are worth acting on.
The R&D Knowledge Generative AI Cannot Find in Published Research
Published papers also represent only part of what researchers actually know. A paper may describe an experiment that worked. It may not tell you how many versions failed first. It may not explain what had to be adjusted in the lab to get the final result.
It may say little about whether the process is manufacturable, whether the material remains stable at scale, or whether the method becomes too expensive outside controlled conditions.
Those practical constraints often determine whether something becomes a product or remains an interesting paper. Much of that knowledge lives in lab experience, failed experiments, internal reports, supplier conversations, manufacturing trials, and years of domain judgment.
A language model cannot reliably recover knowledge that was never documented in the first place. That is why the risk is bigger than hallucinated references. Even when AI retrieves the literature correctly, the literature itself is not the whole decision.
5 Checks R&D Teams Need Before Trusting an AI Research Report
None of this is an argument against using AI in R&D. It’s genuinely faster for literature searches, exploring an unfamiliar area, or organizing a large pile of papers into something usable.
The risk shows up when speed gets mistaken for reliability. Five changes cut that risk down.
1. Check the Sources Before You Trust the Conclusion
A report with fifty references can still be wrong. Deloitte’s report had institutional review behind it, professional formatting, and a client who trusted the brand, and it still went out with invented court quotes and citations to papers that don’t exist. If an analysis is going to shape a real decision, someone needs to actually open the key sources and confirm they say what the report claims.
2. Review AI-Generated Research Before You Act on It
This goes for your own team’s use of the tools and for anything a consultant or vendor hands you. AI is good at narrowing down where to look. It shouldn’t be the thing that closes the question.
3. Ask Vendors What They Actually Verify
“We use AI to work faster” doesn’t tell you anything about quality. Ask directly: are references checked against the original source, are big claims traced back to where they came from, and does a person review the evidence before the conclusions get written up?
4. Review Articles Carry the Highest Fabrication Risk
The Lancet audit found review articles had a 57% higher fabrication rate than other paper types. That’s the opposite of reassuring, since review articles are exactly what R&D teams pull for a fast landscape read. A fabricated source in a review doesn’t stay contained. It gets cited forward into every analysis built on top of that review.
5. Keep a Person Responsible for the Judgment Calls
Confirming a paper exists is the easy part. Someone still has to decide whether two conflicting studies can both be right, whether a lab result will actually hold at scale, and whether a technique from a different industry genuinely transfers to yours. That’s not something a citation checker can do for you.
AI Made Research Faster. It Didn’t Make Research Trustworthy
Scientific publishing has not suddenly become unreliable. Most papers do not contain fabricated references. Most researchers are not inventing evidence. Peer review has not stopped working.
What has changed is the cost of creating something that looks credible.
Generative AI can produce convincing scientific language, plausible citations, polished analysis, and confident conclusions faster than previous research systems could. That means an old shortcut is becoming less safe.
For years, seeing that something was published, peer reviewed, professionally written, and heavily referenced was a reasonable first signal of credibility. Now each of those things can still be true while one of the supporting sources is completely invented.
The paper can be real. The authors can be real. The journal can be real. The reference behind the sentence your team is about to build a project around may not be.
And as AI becomes more deeply embedded in literature searches, technology assessments, scouting workflows, and research intelligence systems, those errors can travel much farther before anyone goes back to the original source. That is the bigger change generative AI is creating. The hallucination problem is no longer confined to the AI answer.
It is starting to enter the research record that future AI systems, researchers, and R&D teams will rely on. For R&D teams, source verification is no longer just good research hygiene. It is becoming part of decision quality.
Moving from AI Answers to Evidence-backed Research
Generative AI has made scientific information easier to access, but easier access does not always mean better decisions.
For R&D teams, the challenge is no longer only finding information. It is knowing which evidence to trust, understanding the relationships between different research findings, and building conclusions that can withstand scrutiny. This is where specialized research intelligence platforms become important.
Slate helps R&D teams explore technical questions through a connected knowledge base of patents, research papers, companies, and scientific data. Instead of relying on AI-generated summaries alone, teams can trace insights back to the underlying evidence and investigate the sources behind emerging technologies.
The future of AI-assisted research will not belong to systems that generate the most answers. It will belong to systems that help researchers build answers they can defend.
Explore how Slate helps R&D teams make evidence-backed technology decisions.
