Sponsored Content by BenchlingReviewed by Ify IsiborAug 7 2026
AI can produce hypotheses, but are they worthy of testing?
For the past several years, the standard AI scientist demo has begun with a literature review: read the papers, connect the data, and offer a hypothesis. It is strong, but the model receives barely a fraction of the scientific context it requires.
Point most AI technologies at the published literature, and users get nearly the same findings since a hypothesis derived from public data alone is a consensus hypothesis. Valuable ideas arise from mixing public data with data that no model was trained on: the experiments, the assays that failed, and the program decisions.
Benchling AI’s Hypothesis Generation uses web search to reason across both published literature and institutional memory, transforming the process of generating hypotheses from a literature exercise into a program-specific beginning point for the subsequent experiment.
Image Credit:Shutterstock.com/Krisana Antharith.com
Good hypotheses are more important than many hypotheses
Nicholas Larus-Stone is no stranger to this space. He previously worked at BenevolentAI, where he helped develop its target identification platform. The core idea closely mirrors what many companies are pursuing today: mine the biomedical literature at scale to uncover connections that would otherwise go unnoticed because the volume of published research is simply too great for any individual to process. To identify novel drug targets, BenevolentAI spent years building a knowledge graph from public biomedical data while training AI models to extract and interpret those relationships.
They identified a number of new targets. There were so many unique targets that they had to create an extra layer of technology to help their scientists sort through them all. Even after investing much in technology, they discovered that basing the models only on public data was insufficient. They finally set up their own laboratories to create proprietary data.
The technology was not the issue. It was the data.
Public data can only go so far
Every connection a model discovers in the published record is hidden in a corpus that thousands of other scientists and models are already reviewing. The literature is also structurally incomplete. Positive outcomes are publicized significantly more frequently than unfavorable ones. Research directions that have already been explored but considered ineffective are usually never publicized. A model reasoning over all of this inherits the gaps.
Ask a general-purpose chatbot for a hypothesis, and users will get a competent suggestion from someone who has read everything and understands nothing about the software.
A hypothesis worth running is original, specific, testable, and relevant to the specific scientific purpose, rather than merely biology in general. What users need is an AI Scientist who can produce high-quality hypotheses based on science.

Image Credit: Benchling
Why Benchling is different
Benchling has spent a decade developing a framework for capturing and structuring scientific data. Thousands of scientists use Benchling AI to create new experiments directly in their notebooks. They have now coupled that institutional memory to the public scientific record, allowing Benchling AI to reason about both at the same time. Benchling AI uses web search on AWS Bedrock AgentCore to safely combine internal and public data.
Hypothesis generation on Benchling begins with unique data and then searches public literature to generate fresh ideas. Put the correct context in front of the model before it reasons, and the result transforms from a bad guess to a high-quality hypothesis. It is like recruiting the same smart new person, only this time he’s already been with the company for years.
Multiple models for better performance
Getting the correct data in front of the model is critical. However, even under the best of circumstances, a single model has limitations. That is why Hypothesis Generation in Benchling AI runs many models from various providers in parallel to achieve better outcomes.
To develop hypotheses that users will use in the lab, users must push the boundaries of what today’s AI can achieve. When users ask an LLM the same question again, they receive variations on one point of view. Run numerous models from several sources, and users will see considerably better outcomes.
OpenRouter recently benchmarked this: on a challenging deep-research benchmark, a fused panel of models outperformed every individual frontier model, while a panel of cheaper models got within a point of the best single model at half the cost.
I have not changed my mind about the hard part
Creating hypotheses is the beginning of the scientific process.
The difficult aspect is turning it into an experiment design that can actually test it, carrying out the experiment in the real world, importing data that is messier than any AI predicted, and interpreting data that does not match what is expected. That is where programs fail, and AI tools found solely in literature break down.
All of this is part of the AI Scientist Benchling is developing, in which it generates ideas and conducts tests to validate them.
Try it
Customers of Benchling can now use web search to generate hypotheses. Launch Benchling AI, ask it about a goal, a mechanism, or a troubleshooting issue that users are currently working on, and observe how it uses both the data and the literature.
About Benchling
Benchling makes biotech research and development faster and more collaborative. Biotechnology has the potential to solve humanity’s most pressing challenges, such as disease, renewable energy, clean water, and hunger. The brightest minds are working on these problems but they are equipped with archaic tools. We aspire to fix this and increase the rate of scientific output with a web-based platform that allows researchers to design and run experiments, analyze data, and share results.
Hundreds of thousands of scientists all around the world use Benchling to do research. Whether they are at the world’s largest companies, the top research universities, or working on a startup in a garage, scientists use Benchling for the same reason: to be empowered, not encumbered, by their tools.
Sponsored Content Policy: News-Medical.net publishes articles and related content that may be derived from sources where we have existing commercial relationships, provided such content adds value to the core editorial ethos of News-Medical.net, which is to educate and inform site visitors interested in medical research, science, medical devices and treatments.