The Naked Emperors of AI
On the contribution of each individual drug discovery task to an approved drug — and the contribution of an approved drug to global health
By Alex Zhavoronkov
A few days ago I spoke with an investor friend who specializes in biotechnology and techbio. He sounded excited and a little shaken.
“Alex, you will not believe it. Company X is raising at a $6 billion valuation. Company Y just raised at over $2 billion. And I heard Company Z closed a mega-round above $15 billion.”
I know all three of these companies to some degree. I know this industry very well. And I still could not believe it.
I will keep the companies anonymous. The point is bigger than any single financing round, and I have no interest in attacking people who are trying to build useful technology. Many of the scientists in these companies are brilliant. Some of the investors are brilliant too. But those valuations tell me that a large part of the private market still cannot tell the difference between solving an impressive computational task and discovering a drug that changes human health.
Those are very different achievements.
Drug discovery is a chicken-and-egg problem: you have to discover and develop a drug in order to learn how to discover and develop a drug. Every time you push a program farther, you run into a new family of tasks, constraints, failure modes and decisions. You unlock new skills — and, if you are disciplined enough to write down what actually happened, new benchmarks.
Going from zero to a developmental candidate unlocks more than 1,000 of those skills and opens roughly the same number of benchmarks. Getting through Phase II unlocks more than 1,200. The exact count depends on how finely you slice the work, but the central fact does not move: a drug program is a giant, partially observable, adversarial sequence of biological, chemical, translational, operational, regulatory and clinical decisions. Nobody wins it by topping a single leaderboard.
A beautiful model can win one benchmark and contribute almost nothing to an approved drug. A homely little tool can save an entire program by flagging a metabolic liability, predicting a clinical interaction, picking the right salt, finding the right biomarker, catching a toxicology signal, or steering a team toward the one experiment that prevents six months of wasted work.
So the question investors should be asking is simple: how much does this particular computational skill raise the probability that a safe and effective drug reaches patients?
The next question is harder: how much will that approved drug actually contribute to global health?
Between those two questions sits almost the entire pharmaceutical industry.
The billion-dollar folding reflex
Many of the companies raising at billion-dollar-plus valuations specialize in protein structure prediction, co-folding, protein design and adjacent tasks. The story sells itself. AlphaFold was a spectacular scientific success. Demis Hassabis and John Jumper took half of the 2024 Nobel Prize in Chemistry for protein structure prediction; David Baker took the other half for computational protein design. The Nobel committee noted that AlphaFold2 could predict structures for virtually all of the roughly 200 million proteins then catalogued, and that the tool had already been used by more than two million people in 190 countries.
That is a monumental contribution to science. I use these tools. Our scientists use structure prediction, docking, co-folding, crystallography and structure-based design every day. We know their value firsthand.
But once you actually get your hands on some of the newer models, you see how narrow their reliable range is. Even inside that range, critical drug-discovery skills such as binding-pose prediction remain far from solved. PoseBench — a recent benchmark of deep-learning docking and co-folding systems, including AlphaFold 3, Chai-1 and Boltz-1 — found that these methods generally beat conventional baselines yet still stumble on new or unusual targets. They can struggle to get the geometry right and reproduce the molecular interaction fingerprints medicinal chemists actually care about at the same time. Out-of-distribution performance is still a problem, the same way it is across all of modern AI.
A predicted structure can be enormously useful. It can also be confidently wrong, biologically irrelevant, frozen in a single conformation, or simply mute on whether a molecule will be potent in cells, selective across the proteome, orally bioavailable, metabolically stable, safe in two species, manufacturable, patentable and effective in patients.
A protein structure does not dose a patient.
Why AlphaFold became AlphaFold
To understand today’s excitement, go back to 2018, when the first AlphaFold competed in CASP13. The idea and the groundwork came earlier. The famous leap arrived with AlphaFold2 at CASP14 in 2020, when the system reached a median score of 92.4 GDT across the targets DeepMind reported — accuracy competitive with experimental methods for many proteins. By 2022, predictions covering nearly the entire known protein universe were freely available. In 2024 came the Nobel Prize.
Why AlphaFold became this powerful and this famous is no mystery. Folding was an unlocked skill with a real benchmark, a blind competition and experimental ground truth. CASP was founded in 1994. Organizers picked recently solved structures the competitors had never seen, teams submitted predictions, and performance was scored against reality after the fact. There was a scoreboard. There was competition. There was a community willing to lose in public and get better.
That combination is rare in drug discovery.
AI beating the state of the art in folding was a genuinely big deal. A scientist without tens of thousands of dollars, months of time, specialized equipment, sufficient protein yield or a cooperative crystallization system can now use a computed structure to generate hypotheses and publish useful work. Solving a structure experimentally can take endless trial and error, expensive facilities and, for hard proteins, years of labor. Prediction lets a team examine far more hypotheses before it commits scarce laboratory resources.
Then the Google public-relations machine did what the Google public-relations machine does better than almost anyone. In the minds of many people in technology and finance, folding became the holy grail — practically the whole solution to drug discovery. Large Chinese technology companies chased the same skill and became very good at it. New companies formed around ever more sophisticated versions of structural modeling, co-folding and interaction prediction.
Eight years after the first AlphaFold appeared at CASP13 in 2018, where are the approved drugs it caused?
Some programs have certainly benefited. Many papers have benefited. Many scientists have benefited. Future drugs will benefit. But structure prediction is one family of skills out of the roughly 1,200 it takes to discover and develop a medicine. The contribution is real. The attribution is often absurd.
If a structure model helps a chemist generate twenty better ideas, and one of them eventually becomes part of a successful program, how much of the final drug belongs to the structure model? What about the assay scientist who discovered the target biology was wrong in the disease-relevant cell type? The DMPK scientist who fixed clearance? The toxicologist who caught the liability? The clinician who redesigned the trial? The patient who enrolled? The regulator who demanded the experiment that revealed the correct dose?
We have no serious accounting system for any of this. In its absence, capital flows toward the skills that are easiest to demonstrate, easiest to benchmark, easiest to drop into a beautiful demo, and easiest to explain to a software investor.
Drug discovery rewards a different set of skills: the ones that survive contact with organisms, regulators and patients.
“Give us ten times more data and we will solve biology”
Some multi-billion-dollar companies have spent close to a decade raising money to generate data at massive scale. They have delivered no drugs. Their new explanation is that they still have too little data.
Give us ten times more data, they say, and we will solve biology.
Their investors somehow accept this. Some of them repeat it in public as if it were a law of physics.
I love data. We generate it, we buy it, we clean it, we automate whole laboratories to produce data that models can learn from. But “more data” is not a universal solvent. Ten times more of the wrong assay buys you ten times more confidence in the wrong abstraction. A beautifully standardized dataset can perfect a surrogate endpoint that means nothing in a patient. Biology shifts across cell states, tissues, ages, sexes, disease stages, comorbidities and treatment histories. Clinical medicine piles on adherence, diagnosis, physician behavior, placebo effects, moving standards of care, geography and regulation.
Suppose one of these companies really does solve biology. Congratulations. It still needs chemistry, CMC, formulation, toxicology, pharmacology, pharmacokinetics, regulatory strategy, clinical operations, patient recruitment, statistics, manufacturing and commercialization. A drug’s journey is measured in a decade or more. Even after a mechanism is “solved,” clinical development can burn years and then fail for reasons no omics atlas ever saw coming.
And after all of that, the drug may still be no better than the standard of care.
Unless we also solve aging.
Aging is the common engine under a large fraction of chronic disease. A drug can hit its molecular target perfectly, produce a statistically significant endpoint, win approval — and still add very little to healthy longevity. I care about approved drugs, but approval is not the final score. The final score is years of healthy life created, suffering prevented, disability delayed, and the number of people who can actually get the medicine.
That is why you cannot judge the contribution of an individual AI task by benchmark accuracy alone. Its value has to be discounted by the probability that its output changes a decision, that the decision changes a program, that the program becomes an approved drug, and that the drug produces a meaningful global-health benefit.
That discount rate is brutal.
I was one of the preachers
I remember being one of the main adepts — probably one of the loudest preachers — of a computation-first approach to drug discovery. I never promoted a single skill. I argued for an end-to-end approach. We were probably the first company to put the term “end-to-end generative AI” on drug discovery.
I also remember the skeptics: Derek Lowe, Pat Walters and many others. They challenged the molecular-generation claims, questioned the novelty, asked whether the benchmarks reflected real medicinal chemistry, and kept reminding the field that a generated molecule is not a drug. Back then I often thought they were too conservative.
Now I am afraid of becoming one of them.
The difference is that I know exactly what AI can do and what it cannot do yet. We pushed the frontier. We pushed the limits of efficiency. We built the algorithms, ran them internally, sold them to experts, watched them fail, rebuilt them, wired them to experiments, and dragged programs through the entire developmental gauntlet.
People may accuse me of bragging. Why not brag about it?
Today the total computational work across a whole drug program can add up to less than a week of aggregated task time and a few thousand dollars in AI inference. I am talking about roughly 1,200 tasks, not one glamorous model call. Our algorithms have been used by thousands of experts. In our own hands they have produced 33 developmental candidates, more than a dozen molecules that reached clinical development, zero failures so far among the programs that entered IND-enabling development and beyond, and one Phase III program built on a novel, AI-identified target.
Rentosertib, our TNIK inhibitor for idiopathic pulmonary fibrosis, entered Phase III in July 2026. Its target was prioritized on our biology platform, its molecule was generated and optimized with our generative chemistry, and its Phase IIa results were published in Nature Medicine. The 60 mg arm showed a mean improvement in forced vital capacity of 98.4 mL at 12 weeks in that study. Now it faces the only judge that ultimately counts: a larger, longer, randomized Phase III in patients.
None of this proves every one of our algorithms is brilliant. It proves the system keeps producing things that survive enough tests to become real programs.
And the most exciting discoveries did not always come from a single superhuman model. Thanks to massive scale — and, honestly, less to AI than some people would like to hear — we stumbled onto genuinely novel mechanisms in pain and in aging, mechanisms that had not been floated even in perspective papers or as serious hypotheses. Scale gives prepared minds more chances to run into the unexpected. AI helps us search. Experiments decide whether what we found is real.
So who is going to challenge me for the drug-R&D UFC title?
Step into the octagon with developmental candidates, IND clearances, human data, clinical progression and audited financials. Do not bring a benchmark slide and expect the judges to score it a knockout.
The strange punishment for profitability
We also became profitable. According to some technology investors, this is apparently a bad thing. Profitability means you are no longer a limitless story. Revenue creates expectations. Financial discipline makes it harder to pretend that every dollar burned is proof of a larger future market.
Insilico listed on the Hong Kong Stock Exchange on December 30, 2025. For the first half of 2026 we issued a positive profit alert: expected revenue of roughly $102.5 million to $106.5 million and expected net profit of roughly $33.5 million to $39.5 million, subject to the usual finalization and review. In the same half-year we announced collaborations with Servier, Hygtia, Qilu, CMS, Tenacia, Eli Lilly and SK Biopharmaceuticals, among others. Our platform now serves 13 of the top 20 global pharmaceutical companies.
The only reason a private company valued at two or three times our public-market value should bother me at all is not wounded pride. That company may be better than us at one or two low-value skills out of 1,200 — and it may not even be perfect at those. The danger is what happens on judgment day.
A correction could crash the entire AI drug-discovery sector. Very smart computational scientists with a lot of venture money behind them may simply take a vacation. They are already too expensive for me to hire. Then, after the memories fade and the next buzzword arrives, some of the same people will do the same thing all over again.
And I will be the one left standing here, trying to solve aging.
Clothes for the emperors
Take this as a note, not an attack. Private investors should do better research and make the effort to understand this industry. Purely computational companies should produce high-quality drugs as proof of concept. If and when their bubbles burst, there ought to be something worth carrying forward: a molecule, a dataset connected to real outcomes, a validated workflow, a clinical insight, a drug that can still help patients.
When a naked emperor realizes he is naked, he needs clothing. We are here to help with that.
Our benchmarks, MMAI Gym and MMAI models were built for exactly this. If you lack certain skills, we can help you acquire them. Science MMAI Gym integrates more than 1,000 drug-discovery benchmarks and roughly 120 billion tokens of specialized pharmaceutical data, and it is designed to train and test models across chemical, biological and drug-discovery workflows. We have already put it to work in collaborations with Liquid AI and Human Longevity.
The point of all this is diagnostic. A serious benchmark bank should expose which skills a model has, which it lacks, where it generalizes, where it memorizes, and how much work remains before its output has earned the right to influence an expensive experiment or a human trial. We are not trying to turn every frontier model into a make-believe pharmaceutical company.
A serious benchmark bank should also look like the real value chain: target identification, target validation, disease linkage, assay design, hit finding, hit-to-lead, structure–activity reasoning, selectivity, ADME, PK, toxicology, formulation, translational biomarkers, indication selection, clinical protocol design, recruitment, endpoint choice, competitive intelligence, regulatory reasoning, and a great deal more. And it should tell low-value competence apart from high-value competence.
Structure prediction and co-folding are great. We use both. But I would trade another incremental gain in structure prediction for reliable in-vivo efficacy prediction any day of the week. I might even trade it for a solid improvement in PK prediction. Those are higher-value skills because they sit closer to the decisions that kill programs, consume animals and time, set human exposure, and determine whether a molecule can ever become a medicine.
The market pays more for visible intelligence than for causal contribution. Benchmarks can nudge that. Drugs will settle it.
One candidate a month
I am happy to report that this month we broke our own internal record in drug discovery. We announced our 33rd developmental candidate — the ninth this year alone. That works out to roughly one developmental candidate a month.
It is about the same number as our pharmaceutical deals this year, and we are not planning to slow down.
I will also admit that I care about our stock price. Every CEO who claims not to care about it at all is either lying or should not be CEO. We owe solid returns to the investors who take biotechnology risk. Our early investors gave us the blood to exist; they have now been unlocked for a while, and they made solid returns. I am very happy about that.
To the more speculative investors, a word of caution: please learn at least the basic biotechnology fundamentals and timelines. If the private company you are backing is valued at two or three times a publicly traded company that built AI to deliver roughly one developmental candidate a month, reached profitability, advanced a novel-target drug into Phase III, and now helps frontier AI laboratories train pharmaceutical models and acquire drug-discovery skills as a business, then that private company ought to be producing two or three times the productivity with its AI.
Maybe it can. Show the work.
The rate-limiting step will still be IND-enabling studies and clinical development. Computation has become absurdly cheap next to the time it takes to establish safety, exposure and efficacy in living systems. The physical world refuses to run at GPU speed.
I would give up two of my toes and a lot of money to test our novel-mechanism pain drug even on myself before the IND-enabling studies are complete, simply to learn how well it works for me. If you know of a place where we could obtain a fully compliant IRB approval and pursue a fully lawful regulatory pathway for such a study, let me know. I would take the personal risk to save at least two years of waiting, because the potential public benefit is immense.
But I cannot do it and stay compliant. And I intend to stay compliant. A convincing model output is not a substitute for the evidence that protects patients, and I will not pretend otherwise. That frustration is the clearest picture of the bottleneck I can give you. We can generate hypotheses faster. We can design molecules faster. We can run computational tasks faster. We cannot skip the evidence that keeps patients safe.
An approved drug is already a miracle of coordinated skill. A drug that materially improves global health is a rarer miracle still. It has to work, be safe enough, reach the right patients, be manufactured reliably, earn reimbursement, stay affordable enough to access, and move an outcome that matters.
So celebrate structure prediction. Celebrate co-folding. Celebrate protein design. Give the Nobel Prize when the achievement earns it. Use every good model you can get.
Then ask the impolite question: how much did this skill contribute to the drug, and how much did the drug contribute to global health?
Be careful in this complicated AI drug-discovery world. The majority of the emperors are still naked.
Live long, peak long, and prosper.
Selected sources and fact notes
- Royal Swedish Academy of Sciences, “The Nobel Prize in Chemistry 2024,” October 9, 2024.
- Google DeepMind, “AlphaFold: a solution to a 50-year-old grand challenge in biology,” November 30, 2020.
- Morehead et al., PoseBench evaluation of deep-learning protein–ligand docking and co-folding (AlphaFold 3, Chai-1, Boltz-1), Nature Machine Intelligence, 2025/2026 publication record.
- Insilico Medicine, “Initiates Phase III Clinical Trial for Rentosertib,” July 7, 2026.
- Insilico Medicine, “Positive Profit Alert for the First Half of 2026,” July 9, 2026.
- Insilico Medicine, 2025 annual results, March 29, 2026.
- Insilico Medicine, developmental-candidate announcements and internal operating figures current as of August 2026.
Disclosure: Alex Zhavoronkov is Founder, Chairman, Co-CEO and CBO of Insilico Medicine (HKEX: 3696). This essay is his personal point of view. References to company operating figures should be read together with the company’s formal public disclosures.