Skip to content
CampeloLabs
← Blog

AI in Biotech: The Bottleneck Is the Lab

Cicero Campelo

Cicero Campelo, CISSP
October 3, 2026 · 17 min read

Part of our guide to AI for startups.

A founder turning away from a screen full of AI-generated hypotheses toward a robotic arm running a single experiment at a lab bench
Table of contents

AI in biotech has split into two businesses. One builds models that predict: protein structures, binding affinity, candidate molecules, whole experimental plans. The other builds the physical capacity to test what those models predict. The first is crowded and well funded. The second is where the work stalls.

That is not a contrarian reading. It is what the people buying the models keep saying. Michelle Lee, founder and CEO of Medra, told Greylock partner Corinne Riley that the line she hears from a pharma executive is "I don't want more AI hypothesis, like I don't need more ideas. I don't need more things to test." Her own read on it: "That's actually that bottleneck layer that like everyone's hitting now."

That sentence is the whole strategic picture for anyone building here, and it is the opposite of where most of the capital and most of the demos are pointed. This article takes it apart: what AI in biotech actually does today, which step is genuinely slow, why a century of pharma data is not training data, where the defensible asset sits, and the strongest argument that the whole thesis is wrong.

What AI in biotech actually does today

AI in biotech is the use of machine learning to design, predict and run biological experiments, from proposing a protein structure or a candidate molecule to deciding which test to run next at the bench. Strip the press releases and it does four separable jobs.

  1. Prediction. Models that take a sequence or a structure and predict something useful: how a protein folds, whether a molecule binds a target, what a candidate drug might do. This is the line of work behind AlphaFold 2, published in 2021, and the reason Demis Hassabis and John Jumper shared half of the 2024 Nobel Prize in Chemistry for protein structure prediction, with the other half going to David Baker for computational protein design.
  2. High-throughput screening. Running a cheap, identical test across a very large library of candidates. This is the part people picture when they picture lab robots, and it is also the part that was largely automated a long time ago. In Lee's words, that step of screening "has been ongoing and has been like fairly automated actually for decades now. That's like easy."
  3. Experiment design. Turning a goal such as finding an antibody that blocks a given target into a sequence of specific tests with specific reagents, cells, timings and controls. The step-by-step outline of such a campaign is not the hard part, and Lee is blunt that a frontier model can produce it.
  4. Physical execution. Actually running the custom tests, reading the result, deciding what to run next.

Those four jobs describe AI in genomics, synthetic biology, diagnostics and bioprocessing just as well. This article follows the pharma and antibody case, because that is where the capacity constraint is sharpest and where the money has already arrived.

Jobs one and two drew most of the attention and a lot of the investment. Jobs three and four are where the clock is.

The distinction that matters is inside job two. Screening a million compounds for one cheap property is automated. Screening is not the same as the targeted work that follows it, which is what the rest of this article is about.

The bottleneck moved from ideas to experiments

Lee's argument against building a standalone AI scientist is a competitive one before it is a scientific one. Hypothesis generation is not differentiated, she argues, because the frontier labs are already building science-specific models of their own. Her conclusion is that "I don't think hypothesis is the bottleneck right now," and her evidence is the room: "you talk to scientists, none of them will say, 'Yeah, I need more ideas and hypothesis.'"

She is the founder of a lab company, so discount for interest. The useful check is whether a company on the other side of that trade agrees, and it does. Chai Discovery is a model company, founded in 2024 by Joshua Meier, Jack Dent, Matthew McPartlon and Jacques Boitreaud, and it has every incentive to argue that models are the hard part. Talking to Sequoia Capital, two of its co-founders described the readout as the slow, noisy step anyway. One of them put the comparison with software directly: "Maybe it takes a little bit longer to validate it, right? It's not like five seconds to like get a readout and run a unit test. You might have to spend a couple days, a couple weeks in the lab to get that readout, but at least you can be honest with yourself."

Two companies with opposed commercial interests, describing the same constraint. That is a stronger signal than either claim alone, and it is the one worth building a strategy on.

Why assay development is the hard part

An assay is a test: a specific recipe for measuring one property of one candidate. Does this antibody bind this target? That question, made concrete enough to run, is an assay. So is whether the antibody still works inside a living cell, and whether it survives being warmed up.

Lee walks the funnel. You start with maybe a million compounds and run the cheapest possible screen. A million becomes a thousand, a thousand becomes a hundred, a hundred becomes fifty. Then the character of the work changes completely: "Then suddenly the type of experience you want to run becomes super targeted, super custom. You have to really understand the science."

At that point the question stops being throughput and starts being whether the test works at all. Which cells, at what density, in what buffer, incubated how long, with which controls. Lee describes it as a multi-parameter optimization you have to run just to build the instrument you then use to get your answer. Her figures for how long that takes: outsource it to a contract research organization and "it can take them months to just develop that one assay"; keep it in-house and it is still weeks to months. Those are her figures, offered in conversation rather than drawn from an industry survey, so treat them as a practitioner's estimate.

The reason you cannot shortcut it with reading is the line a founder in any empirical domain should write down: "the paper can give you experimental steps that like don't actually yield the result you care about. The only way you can do it is run it."

Medra has published one concrete example of what compressing that loop looks like. In its own antibody screening workflow, generating binding data took three days, with a plasmid cloning step as a large part of the wait. Its AI Experimentalist redesigned the assay around linear DNA expression templates to remove the two-day cloning step, which cost protein yield, then tested multiple linear-template designs and protocol conditions in parallel, iterating toward a workflow that recovers comparable expression without giving back the time. The company reports the three-day workflow dropped to roughly 14 hours. Note the shape of the win: no single clever trick, a sequence of small changes each validated by a run.

Why a century of pharma data is not training data

Here is where the strategy gets interesting for founders outside biotech, because the constraint generalizes.

Lee's position is that AI for biology is held back by the supply of trainable data, and that many of the AI-first companies are training on the wrong data. "Many of the very AI focused companies are actually training on data that have just existed for decades now." Her line for it, adapted from Coleridge: "data data everywhere, nor a drop to train."

The distinction she draws is between having records and having training data. A pharma company can hold more than a century of experimental records and still have almost nothing a model can learn from, "it doesn't mean it's actually trainable data. So you actually need trainable data where you know exactly how the experiment is run." Known parameters, known conditions, low variance, reproducible. Most historical lab data fails that test not because anyone was careless but because nobody was recording for a consumer that did not exist yet.

By Lee's own order-of-magnitude estimate, offered in conversation rather than from a published benchmark, the largest AI models for biology train on something like a thousand times less data than frontier general models. Read that as a claim about the shape of the problem rather than a measurement.

Chai Discovery's co-founders put the quality half of the same problem in sharper terms. One of them, on why rigor matters more in biology than in code generation: "You can fool yourself so easily in biology. Like the error bars in the wet lab are actually quite large as well." If plus or minus 5 percent in your lab might all be the same result, then a model improvement smaller than your noise floor is not an improvement, it is a story. The same logic governs any startup measuring a fuzzy outcome, which is the argument in why data, not compute, is the real bottleneck.

In AI biotech the moat is how you run the experiment, not what it finds

This is the single most transferable idea in the interview, and most founders get the analogous question wrong in their own domain.

Medra runs experiments for pharma companies, frontier model labs and government partners, which means it sits on top of other people's most valuable results. It explicitly does not treat those results as its asset. Lee on the IP question customers ask first: "no experimental results are our customers. They own that completely." No training on customer outcomes.

What it does train on is the part nobody else bothers to keep. Run the same experiment three times and you get three slightly different answers. The variance, and the reason for the variance, is the data: "What we're really trying to be able to have our agent capture is what does it mean to run a good experiment that can produce high quality reproducible data."

That data is physical and absurdly fine-grained. The way you pipette, the angle of the pipette, the depth of the tip. Lee gives the example of stem cell work where pipetting too fast puts the cells into shock and kills them. None of that is in a paper, because humans running the bench do not write it down. A platform that watches every run does.

So the defensible asset is process knowledge rather than results, and that split is worth stealing whole. The outcome belongs to the customer and is the thing they will fight you over. The telemetry of how the outcome was produced is the thing nobody has claimed, nobody is recording, and nobody can reconstruct after the fact. The equivalent question for your company: what does your system observe on every run that your customer does not care about and your competitor cannot get? That is your training set.

That a closed experimental loop beats a model is the same conclusion AI drug discovery reaches from the pipeline economics, so this post does not re-argue it. What is new here is the sharper version: inside the loop, it is the process data rather than the results that nobody else can buy.

Price the outcome, not the motion

The commercial model follows from the same logic, and it is the part most infrastructure startups get backwards.

Lee contrasts how lab automation and outsourced research have traditionally been sold, as a quantity of robot motions or a number of staffed scientists for a period, with how Medra signs deals: "It's always what is the scientific outcome you care about?" The contract is written against the science, not the mechanism. "that is ultimately the ROI they care about, not like your robot can do step one, step two, step three, but here's the science we'll be able to do for you."

Palantir is credited in the conversation with championing this forward deployed model, and Lee says Medra has worked that way from the beginning. The bet is a real one: pricing on outcomes means absorbing the risk that your system needs five runs where you budgeted two. You take that risk precisely because the runs teach you something, which is the version of the argument in the forward deployed engineer.

Sell the platform or run the lab?

When Medra opened its own facility it did not replace the deployment business, it added a second one. The company sells the platform to run inside a customer's own space, and it also runs experiments for customers in Medra Lab 001, which it opened in April 2026 in San Francisco and describes as the largest autonomous lab in the United States, at 38,000 square feet and built in 77 days. The second model is what let it say yes to customers with no labs at all, including, it says, a DARPA-funded project.

Lee frames the choice with a computing analogy: having built the equivalent of a new chip, whether you then run the cloud or help customers run their own data centers is a business model question, not a product question. A pharma company with its own labs and its own IP concerns often will not outsource its most critical experiments, so the only way to serve it is to build the capability inside its walls. Owning a facility opened a different set of customers, mostly though not always the ones with no labs at all, which is how a government agency and frontier model companies became buyers.

The founder lesson is to stop treating the deployment model as a binary. If your capability is genuinely infrastructure, the same technology supports a service business and a build-it-for-you business with different customers, different sales cycles and different margins. Picking one early because it sounds cleaner can cost you a segment that was never going to buy the other.

Are autonomous labs actually autonomous?

Medra's platform has two halves: the AI Experimentalist, which does the reasoning, and what the company calls the Physical AI Lab, the robots that run the experiment. Together it calls them the Physical AI Scientist. The phrase autonomous lab still oversells what is shipping, and Lee is direct about the limit: "it's not fully autonomous. We actually want there to be a human in the loop" so a scientist signs off on the experimental plan.

The throughput claim is narrower and more credible than full autonomy. The platform runs when people do not: "We're able to run our platform at night over weekends." Results get analyzed and the next experiment proposed whether or not anyone is awake, so "something that could take a really good scientist months to optimize, we can shrink that time down," to weeks or days.

There is a hiring insight buried in how she describes the target. The best human experimentalists, in her account, are the ones who both run the experiment themselves and analyze the data, because they notice the odd film forming while it forms and let that change how they read the result. Splitting those two jobs across two people loses the observation. That is a useful test for any workflow you are automating: find the handoff where tacit context dies, and either keep one person on both sides of it or build the system that carries the context across.

Her end state is a change of shape rather than speed, replacing the discovery funnel with "many smaller loops that can run at scale, be able to validate itself." The constraint she keeps returning to is why the lab never disappears: "Science again is empirical." There is no simulation that proves a drug works in a body. If your product makes claims about the physical world, something has to go and check, which is the line between physical AI and software that only has to compile.

The counterargument: what if models need far less data?

The thesis has a real opponent, and it is worth stating at full strength rather than as a straw man.

Medra's bet is that the answer to data-poor science is orders of magnitude more data. Flapping Airplanes, an AI research lab founded in 2025 by brothers Ben and Asher Spector with Aidan Smith, bets the opposite: that the answer is models that need far less. Presenting at Sequoia Capital, Ben Spector made the case that the valuable domains ahead are precisely the ones where data is scarce and hard to manufacture, naming science among them: "In trading, there's only so much financial data. Scientific discoveries, obviously very little data." If a model can learn a domain from a thousandth of the data, then paying to industrialize data generation is a very expensive way to solve a problem that was about to get cheaper.

Both can be right for a while, which is the honest read. Data efficiency research and data generation infrastructure are complements until one of them wins decisively, and no founder knows which. The decision-relevant version: if you are building the data generation layer, your risk is not a competitor, it is a step change in sample efficiency that makes your capacity less necessary. If you are building on data efficiency, your risk is that the remaining gains stay stubbornly empirical and the labs with capacity eat the field. Price the risk you are actually taking.

What this means if you are building in a physical domain

Three things carry over whether or not you touch biology, and they are the ones the action list below does not already cover. The wider map of where this logic shows up across a startup is in AI for startups.

  • Whatever is upstream of the physical step gets commoditized. In any empirical domain there is a point where only running the thing produces the answer. Frontier models will reach everything before that point sooner than you expect, and they will not reach the step itself.
  • Records are not a data asset. Rows with unknown parameters and uncontrolled variance are a liability dressed as an asset, and no amount of volume fixes it.
  • Do not confuse removing the human with raising throughput. Running overnight is most of the compounding. Sign-off from a person is cheap, and it is what makes an expensive customer willing to try you at all.

What to do this week

  1. Write down the one step in your product's loop that cannot be replaced by better reasoning, only by running something real. If you cannot name it, you are probably selling experiment design into a market that needs physical execution.
  2. Take your largest internal dataset and grade a sample of 50 records for trainability: are the parameters recorded, the conditions known, the variance measurable? Count how many pass. That number, not the row count, is your data asset.
  3. List what your system observes on every run that you currently throw away. Pick the one a competitor could not reconstruct from outputs alone and start persisting it this week.
  4. Reread your largest contract and find the unit you are paid in. If it is a motion, a seat or a quantity of work, draft the version of the same deal priced on the outcome and model what happens if iterations double.
  5. Run the same operation three times under identical conditions and measure the spread. If the spread is bigger than the improvement you have been reporting, your last three wins were noise.
  6. Find the handoff in your workflow where tacit context dies, where one person runs the thing and a different person interprets it. Either merge the roles for a week or specify what the system has to carry across.
  7. Name the technical breakthrough that would make your core bet unnecessary, and write the trigger that would tell you it is happening. Review it monthly, not when a competitor's launch forces you to.

This kind of reasoning is what the AI Operating System for Startups course is built around: how to pick the bet, instrument the loop, and tell early whether the data you are collecting is worth anything. If the job question in point one is the one you are stuck on, will AI take my job is the same argument pointed at a career instead of a company.

Sources

Frequently asked questions

How is AI being used in biotechnology?

AI is used in biotechnology in four separable jobs. Prediction: models that take a sequence or structure and predict folding, binding or a candidate molecule, the line of work behind AlphaFold 2 and the 2024 Nobel Prize in Chemistry, half of which Demis Hassabis and John Jumper shared for protein structure prediction, with David Baker taking the other half for computational protein design. High-throughput screening: running one cheap identical test across a huge candidate library, which Medra founder Michelle Lee notes has been largely automated for decades. Experiment design: turning a goal into specific tests with specific cells, reagents, timings and controls. And physical execution: running the custom tests, reading the result and deciding what to run next. Most of the attention went to the first two jobs. The last two are where the calendar time actually goes.

What is the biggest bottleneck in AI for biotech?

The biggest bottleneck in AI for biotech is experimental validation, not idea generation. Michelle Lee, founder and CEO of the autonomous science company Medra, argues that hypothesis generation is no longer differentiated because the frontier labs are building science-specific models themselves, and that scientists are not short of ideas. The slow step is developing an assay, the specific recipe that measures one property of one candidate, which she says can take a contract research organization months and still takes weeks to months in-house. A company on the other side of that trade agrees: co-founders of the model company Chai Discovery describe a lab readout as taking days to weeks rather than the seconds a unit test takes, and note that wet-lab error bars are large enough to swallow a model improvement.

What is an autonomous lab?

A facility where robots execute biological experiments end to end, paired with a reasoning layer that designs the experiments, interprets the results and proposes the next run. Medra calls its version the Physical AI Scientist: a reasoning layer it calls the AI Experimentalist, paired with a robotic execution layer it calls the Physical AI Lab. The name oversells the autonomy, and Medra's own founder says so: the system is not fully autonomous, and a human scientist signs off on the experimental plan. The real claim is throughput rather than independence. The platform runs at night and over weekends, analyzing results and proposing the next experiment whether or not anyone is awake, which is how optimization work that would take a good scientist months compresses toward weeks or days.

Will AI take over biotech jobs?

The evidence from the companies automating lab work points at a change in the job rather than its removal. Medra keeps a human in the loop to approve experimental plans, and it does not train on customer results at all: those belong to the customer. What it captures instead is the process detail humans never write down, like the angle and depth of a pipette, or the fact that pipetting stem cells too fast puts them into shock. That shifts the value of a scientist toward the judgment calls, which goals to pursue, which results to trust, which anomaly is worth chasing, and away from manual execution. The scientists Lee rates most highly are the ones who both run the experiment and analyze the data, because the observation made at the bench changes how the result is read.

Build your AI Operating System

A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.