<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Carlos Outeiral]]></title><description><![CDATA[I write about frontier AI in biology and life sciences. Previously Eric Schmidt AI in Science Fellow, and Lecturer in Biochemistry, at the University of Oxford.]]></description><link>https://couteiral.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!sEgI!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb357d710-d998-4626-bea4-47c730117f1b_400x400.jpeg</url><title>Carlos Outeiral</title><link>https://couteiral.substack.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 07 Aug 2026 16:36:48 GMT</lastBuildDate><atom:link href="https://couteiral.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Carlos Outeiral]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[couteiral@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[couteiral@substack.com]]></itunes:email><itunes:name><![CDATA[Carlos Outeiral]]></itunes:name></itunes:owner><itunes:author><![CDATA[Carlos Outeiral]]></itunes:author><googleplay:owner><![CDATA[couteiral@substack.com]]></googleplay:owner><googleplay:email><![CDATA[couteiral@substack.com]]></googleplay:email><googleplay:author><![CDATA[Carlos Outeiral]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Coevolution took us here, but it's not enough]]></title><description><![CDATA[8.4k words, 33-41 minutes reading time]]></description><link>https://couteiral.substack.com/p/coevolution-took-us-here-but-its</link><guid isPermaLink="false">https://couteiral.substack.com/p/coevolution-took-us-here-but-its</guid><dc:creator><![CDATA[Carlos Outeiral]]></dc:creator><pubDate>Wed, 05 Aug 2026 13:10:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6qt2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Five years since <a href="https://www.nature.com/articles/s41586-021-03819-2">AlphaFold 2 was published</a>, it remains <strong>the single most significant advance in computational biology this century</strong>. Before it, <a href="https://en.wikipedia.org/wiki/Structural_biology">solving the structure of a protein</a>, a crucial task in molecular biology,<strong> often took years and millions of dollars</strong>. After its release, it could (mostly)<strong> be answered in a few hours, for a handful of dollars</strong>.</p><p>You would expect a breakthrough of this magnitude to be the beginning of a curve. Such is generally the case with <a href="https://www.nobelprize.org/prizes/chemistry/2024/press-release/">Nobel-worthy discoveries</a> (think of CRISPR, or the transistor): the seminal work unlocks a field that then flourishes independently, with manifold improvements that far outrun the original contribution. Unfortunately, <strong>that is not what happened in AI for proteins</strong>. This essay is my attempt to explain why.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6qt2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6qt2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6qt2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6qt2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6qt2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6qt2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg" width="745" height="412" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:412,&quot;width&quot;:745,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Protein designer and structure solvers win chemistry Nobel | Science | AAAS&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Protein designer and structure solvers win chemistry Nobel | Science | AAAS" title="Protein designer and structure solvers win chemistry Nobel | Science | AAAS" srcset="https://substackcdn.com/image/fetch/$s_!6qt2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 424w, https://substackcdn.com/image/fetch/$s_!6qt2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 848w, https://substackcdn.com/image/fetch/$s_!6qt2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!6qt2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82b0a954-0c0b-430f-9e1e-4bebe2a45bf4_745x412.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Traditional slide from the Royal Swedish Academy of Sciences announcing the Nobel Prize in Chemistry 2024, half of which went to Demis Hassabis and John Jumper. With half of the Physics prize that year going to Geoffrey Hinton, it was a good year for Google!</figcaption></figure></div><p>You can see the problem in the models themselves. After AlphaFold 2 answered the question <em>&#8220;what is the shape of this protein?&#8221;</em>, its successor <a href="https://www.nature.com/articles/s41586-024-07487-w">AlphaFold 3</a> set out to answer a harder one: &#8220;<em>how do proteins bind other biological molecules, from drugs to DNA?</em>&#8221;. The model is a solid piece of architectural work and does push the frontier. But, unlike AlphaFold 2, <strong>you cannot trust its predictions</strong>. Independent analyses show that it <a href="https://www.biorxiv.org/content/10.1101/2025.02.03.636309v3.abstract">has a tendency to memorise and regurgitate its training set</a>, and that it <a href="https://www.nature.com/articles/s41467-025-63947-5">makes egregious mistakes that show a lack of understanding of basic protein physics</a>, such as clear clashes between atoms. The recently released <a href="https://storage.googleapis.com/isomorphiclabs-website-public-artifacts/isodde_technical_report.pdf">IsoDDE</a> reports real gains on some of the hardest cases, but from the limited information released it feels more like AlphaFold 3.5 than a true step change.</p><p>In essence, <strong>progress seems to have slowed down.</strong></p><p>This is frustrating because <strong>far more resources went into these later models</strong>. While AlphaFold 2 was a side project at Google DeepMind (its team <a href="https://www.ft.com/content/61b2953d-ee0d-45de-af6e-a9c1cf524b33?syn-25a6b1a6=1">reportedly wound down</a>), its successors are the whole <em>raison d&#8217;&#234;tre</em> of Isomorphic Labs (amongst many other AI for Bio companies), which has raised <a href="https://www.isomorphiclabs.com/articles/isomorphic-labs-announces_-600m-external-investment-round">$600M</a> and then a further <a href="https://www.isomorphiclabs.com/articles/isomorphic-labs-announces-series-b-investment-round">$2B</a> in outside money. Yet, <strong>despite orders of magnitude growth in investment, we are getting only marginal improvement</strong>.</p><p>What went wrong? A lot of people like to point their finger at training data availability. The <a href="https://www.rcsb.org/stats">Protein Data Bank</a>, used to train every AlphaFold, held around 140,000 structures in 2018 and roughly 240,000 today, a rise that shrinks further once you account for how many are near-duplicates of proteins already in the set. The slow growth in structural data is a possible reason for slowdown, but considering the amount of cash in their hands, it is hard to believe data is the true bottleneck. At least, Demis Hassabis has been quite vocal that <a href="https://endpoints.news/isomorphic-labs-ceo-demis-hassabis-bets-on-biotechs-ai-future/">he doesn&#8217;t think data is the problem</a>. If you instead follow the AI world, you might pin the marginal improvements on <a href="https://en.wikipedia.org/wiki/Neural_scaling_law">scaling laws</a>: perhaps training even more powerful models requires inaccessible amounts of compute. And yet, if you <a href="https://www.nature.com/articles/s41592-024-02272-z">estimate the cost of training AlphaFold 2 at about 50,000 A100 40GB hours</a>, or about $80k in 2026 AWS prices, it&#8217;s clear that compute is not quite what&#8217;s holding everything back. Don&#8217;t get me wrong, <strong>these are </strong><em><strong>reasons</strong></em><strong>. But they are not </strong><em><strong>the</strong></em><strong> reason.</strong></p><p>The thesis of this essay is that <strong>the root cause of the slowdown is reliance on <a href="https://www.nature.com/articles/nrg3414">coevolution</a></strong>. The fundamental breakthrough in AlphaFold 2 was its ability to mine the evolutionary history of a protein to rapidly explore conformational space. Within a single protein, residues that touch in the folded structure (<em>&#8220;contacts&#8221;</em>) are constrained to evolve in concert: if one mutates, its partner must often mutate to compensate for it, lest the protein fall apart. The same holds across proteins that must bind each other. So if you align a protein&#8217;s sequence against millions of its relatives in genomic databases, <strong>these correlated changes capture a faint statistical shadow, left by evolution, of which residues sit close in space</strong>. AlphaFold models (and in particular their Evoformer/Pairformer modules) are extremely effective at doing this. The reliance on coevolution, while a powerful prior, also imposes a hard ceiling: because all life shares common ancestors, most proteins are variations on ones the model has already seen, and what looks like understanding is often closer to interpolation among evolutionary relatives. Put another way: <strong>the model appears to have learned the physics of biomolecular interactions, when it has largely learned protein genealogy</strong>.</p><p>I would like to make the case for going beyond coevolution as the main frontier in artificial intelligence for biomolecular modelling. I will do this through three objectives:</p><ul><li><p>map the main ideas that gave rise to the AlphaFold models, which will borrow heavily from my personal memory</p></li><li><p>discuss where the frontier is today, and what the likely hurdles are</p></li><li><p>make a few predictions about the future, without fear of any potential embarrassment a few years from now.</p></li></ul><p>Let&#8217;s start from the beginning.</p><h3>The pre-AlphaFold era (~2010-2021)</h3><p>AlphaFold has become so vital to life sciences that many struggle to remember the world before it, or where its ideas actually came from. But, trust me, they came from somewhere. The breakthrough that felt so sudden in 2020 owed much to a decade of slow, unglamorous progress toward a single idea.</p><p>I want to tell that story from the inside, because I happened to be there for part of it. I came to the field in 2017; at the time, I was working with <a href="https://www.matter.toronto.edu/basic-content-page/about-alan">Al&#225;n Aspuru-Guzik</a> at Harvard, on generative models that could design tailored molecules according to some properties (some of that work is <a href="https://arxiv.org/abs/1705.10843">here</a> and <a href="https://chemrxiv.org/doi/full/10.26434/chemrxiv.5309668.v3">here</a>). We were using <a href="https://en.wikipedia.org/wiki/Generative_adversarial_network">generative adversarial networks</a> (GANs, an ancestor of today&#8217;s <a href="https://en.wikipedia.org/wiki/Diffusion_model">diffusion models</a>, now relatively out of use) combined with reinforcement learning to steer molecules towards useful properties. If you plugged in a solubility predictor, you&#8217;d get small, very polar molecules. If you used a melting point predictor, the network would learn to string together long molecules with lots of hydrogen bond donors and acceptors.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!20Nk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!20Nk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 424w, https://substackcdn.com/image/fetch/$s_!20Nk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 848w, https://substackcdn.com/image/fetch/$s_!20Nk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 1272w, https://substackcdn.com/image/fetch/$s_!20Nk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!20Nk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png" width="1456" height="413" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:413,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:107009,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!20Nk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 424w, https://substackcdn.com/image/fetch/$s_!20Nk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 848w, https://substackcdn.com/image/fetch/$s_!20Nk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 1272w, https://substackcdn.com/image/fetch/$s_!20Nk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb771507b-1079-4e21-a88e-0eabe924f19e_1926x546.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Example of the molecules that our model <a href="https://chemrxiv.org/doi/pdf/10.26434/chemrxiv.5309668.v3">ORGANIC</a> would generate when paired with a melting point predictor: extremely long molecules with tons of hydrogen bond donors and acceptors and a tendency to stick together.</figcaption></figure></div><p>We chose those properties not because they were hugely interesting, but because they were easy to calculate and plug into the models. But what I really wanted to generate were <em>drugs</em>. And drug potency, unlike solubility or melting point, is not a property of just the molecule: it depends on the three-dimensional structure of the protein the molecule is meant to bind. My original assumption was that drug design could be solved by looking up the PDB, running a docking program, and using the predicted binding energy to prompt our system. But, of course, as I explored the problem, it became obvious that for a lot of drug discovery problems the structure that would be the starting point was not even available! For a twenty-something researcher, unburdened by experience, the conclusion was obvious enough: the interesting problem was not the drug. The really interesting question, the one that would occupy me for the best part of the following decade, was: <em>how do you find out the structure of the protein</em>?</p><p>At the time, there were two broad ways to predict protein structure. The first, <a href="https://carlos.outeiral.net/ai4proteins">template-based modelling</a>, was not really a prediction at all. One of the central observations in protein structure is that it is extremely <em>conserved</em> across organisms: because the structure determines the function, evolution tries to conserve it at all costs. If there exists a protein structure that was even distantly related to the target (as a rule of thumb, this means ~25-30% of amino acids are identical) you can use the reference as a scaffold and refine it in place. This method works beautifully, of course, because you aren&#8217;t as much <em>predicting </em>as copying what already existed.</p><p>Far less successful was the other current, <a href="https://carlos.outeiral.net/ai4proteins">template-free modelling</a>, which tried to understand the structure from first principles. These algorithms were inspired by <a href="https://en.wikipedia.org/wiki/Anfinsen%27s_dogma">Anfinsen&#8217;s hypothesis</a>, which holds that proteins fold to their lowest energy state, transforming &#8220;prediction&#8221; into an optimisation problem (see the video below for an example of how this was done in practice, with <a href="https://www.pnas.org/doi/abs/10.1073/pnas.91.10.4436">fragment optimisation</a>). Unfortunately, because the energy functions were crudely approximate and the landscape was astronomically large and riddled with false minima, the process was rarely successful. In other words: if you wanted to know a protein&#8217;s structure, your only choice was to <em>already know</em> (i.e. have a good template for) the structure of a well-related one!</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;2cff46e2-f98f-4f34-9757-073653abf28f&quot;,&quot;duration&quot;:null}"></div><p>However, in the 2010s an idea started to arise that made template-free structure prediction much more likely to succeed. The central idea is that <em>pairs of amino acids tend to coevolve</em> in protein sequences. When two residues interact in the folded structure, they are constrained to evolve in concert: <strong>if one suffers a mutation, its partner must often mutate to compensate, lest the protein fall apart</strong>. These correlated mutations leave a <strong>faint statistical shadow of which residues are close in space</strong>, and if you can use the right statistical methods to extract it, then you can get a rough blueprint of which parts of the protein are close to one another, and use it to make a strong template-free model. The idea was <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/prot.340180402">first proposed in 1994</a>, but the groups that really put it forward were those of <a href="https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0028766">Chris Sander and Debbie Marks</a> at Sloan-Kettering, and <a href="https://academic.oup.com/bioinformatics/article/28/2/184/198108">David Jones</a> at UCL.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9eZt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9eZt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!9eZt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!9eZt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!9eZt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9eZt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1526653,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9eZt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!9eZt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!9eZt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!9eZt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb8ab2e52-0925-4793-8ac4-549e390d592c_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Infographic explaining how coevolution leads to an imprint in the sequence record. The bottom line is that organisms with ill-functioning proteins will be in evolutionary disadvantage and die, and therefore their proteins&#8217; sequences will not appear in the sequence registers. Thus, we only see protein sequences where any random mutation that disrupted structure also had a compensatory, neutralising mutation. Image generated with GPT Image-2, after multiple prompts.</figcaption></figure></div><p>A &#8220;contact&#8221; is the simplest possible fact you can state about a folded protein: that two residues touch<strong>.</strong> In the language of structural biology, this means that their <a href="https://en.wikipedia.org/wiki/Amino_acid">beta carbons</a> (or alpha, in the case of <a href="https://en.wikipedia.org/wiki/Glycine">glycine</a>) sit within 8 &#197; of one another. It sounds almost trivially small, but even a handful of long-range contacts is enough to make template-free prediction much more effective, because rather than finding a needle in a haystack you are now choosing between a few plausible ones. The catch, of course, is that you need to predict these contacts with sufficient accuracy.</p><p>For a long time, unfortunately, this was barely possible. The trap is that, when you examine a list of sequences, two amino acids can look correlated not because they directly touch, but because both of them interact with a third. The first coevolutionary methods, like <a href="https://en.wikipedia.org/wiki/Direct_coupling_analysis">direct-coupling analysis</a> (DCA) or <a href="https://academic.oup.com/bioinformatics/article/28/2/184/198108">inverse covariance analysis</a> (ICA), used heavy statistical machinery (a Potts model, or a Gaussian graphical model) with a complicated inference process to distil true interactions from the noise. While they had a <a href="https://onlinelibrary.wiley.com/doi/full/10.1002/prot.25064">huge impact in the field</a>, they were held back on nearly every front: they were highly dependent on the number of sequences available to build multiple sequence alignments (in fact, some authors have attributed most of the contact prediction in the 2010s to <a href="https://onlinelibrary.wiley.com/doi/full/10.1002/prot.25407">the growth of sequence databases</a>), they were computationally really expensive, and they were underpowered.</p><p>Here is when, as someone who lived through the first half of the 2020s, you might be guessing that the solution is <em>deep learning</em>. In fact, since it was a good number of years after <a href="https://en.wikipedia.org/wiki/AlexNet">AlexNet </a>you may reasonably be wondering <em>&#8220;wait, why weren&#8217;t they using deep learning already?&#8221;</em>.</p><p>The truth is that <em>nobody was doing deep learning in structural biology</em>. Ask someone today to name a pioneer of AI for biology and they&#8217;ll probably say <a href="https://en.wikipedia.org/wiki/David_Baker_(biochemist)">David Baker</a> &#8212; yet the Baker lab&#8217;s first <em>bona fide</em> deep learning paper, <a href="https://www.pnas.org/doi/10.1073/pnas.1914677117">trRosetta</a>, didn&#8217;t arrive until 2019, and could fairly be read as a follow-up to the first AlphaFold. My own supervisor, <a href="https://www.stats.ox.ac.uk/people/charlotte-deane">Charlotte Deane</a>, only picked up deep learning around the time I joined her group in 2018. A few folks were using some sort of neural networks, most notably David Jones&#8217;s <a href="https://academic.oup.com/bioinformatics/article/31/7/999/181290">MetaPSICOV</a> or <a href="https://pubmed.ncbi.nlm.nih.gov/31298436/">DeepMetaPSICOV</a>, and Jinbo Xu&#8217;s ultra-deep <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/prot.25377">RaptorX</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a>, which set some of the best contact-prediction records of its day. But these networks were feeding <em>hand-built</em> coevolutionary features into a network rather than letting the neural network figure out what to extract from the representation of the model, which is what we would today consider <em>deep learning</em>. The one system that genuinely pointed at the future with deep learning was the <a href="https://linkinghub.elsevier.com/retrieve/pii/S2405471219300766">Recurrent Geometric Network</a> by <a href="https://systemsbiology.columbia.edu/faculty/mohammed-alquraishi">Mohammed AlQuraishi</a> which, notably, did not use coevolution at all.</p><p>In 2018, a team at DeepMind entered the fray. While the company was well-known from its stints such as <a href="https://www.nature.com/articles/nature14236">training a neural network to play Atari games</a> or <a href="https://www.bbc.co.uk/news/technology-35785875">defeating Go&#8217;s world-champion with a reinforcement learning system</a>, they had a tendency to keep their projects very secret. Years later, over dinner, <a href="https://en.wikipedia.org/wiki/John_M._Jumper">John Jumper</a> told me that when he interviewed, the group&#8217;s leader <a href="https://research.google/people/author37792/">Andrew Senior</a> wouldn&#8217;t even confirm they worked on protein structure prediction. He would only admit, vaguely, that they did &#8220;<em>coevolutionary analysis</em>&#8221;.</p><div class="callout-block" data-callout="true"><p><strong>Addendum: 6th August, 2026</strong></p><p>John Jumper reached out to tell me the anecdote was not quite as I remembered it. The secrecy part was real: during the interview nobody would admit that the group worked on protein structure prediction, or that they were recruiting anyone to work on it. The coevolution part was slightly different. At his interview, John was describing <a href="https://arxiv.org/abs/1610.07277">a piece of his PhD work</a>, using a statistical potential combined with the coevolutionary contact probabilities published by Sheng Wang. Knowing that Andrew Senior came from an audio processing background, he tried to explain it from first principles: how the spatial proximity of two residues leaves a correlation in their evolutionary record. Andrew interrupted him. &#8220;<em>Oh, do you mean coevolution</em>?&#8221; The term was niche enough to convince John that they were, as the rumours suggested, working on protein structure prediction.</p></div><p>The AlphaFold team came out of stealth at the 13th edition of <a href="https://en.wikipedia.org/wiki/CASP">CASP</a>, the field&#8217;s blind assessment of structure prediction, which you can think of as the <a href="https://arcprize.org/arc-agi">ARC-AGI</a> of proteins, where methods are scored on structures nobody has seen. Their method used a deep convolutional neural network, acting directly on a multiple sequence alignment, to predict not just the <em>binary</em> <em>contacts</em> between amino acids but the <em>distances</em> between every pair of residues. The result left no doubts: by fusing coevolution with genuine deep learning, that first AlphaFold pushed template-free prediction further in one CASP than the field had managed in a decade<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a>.</p><p>The zeitgeist is best captured in this <a href="https://moalquraishi.wordpress.com/2018/12/09/alphafold-casp13-what-just-happened/">beautiful blog on CASP13</a> by Mohammed AlQuraishi, one passage of which has stayed with me for years.</p><blockquote><p><em>Taken together the above suggests substantial progress, more so than usual, and hence not only did AlphaFold &#8220;win&#8221; CASP13, but did so by an unusual margin. Great! Does this mean the problem is solved, or nearly so? The answer, right now, is no. We are not there yet. However, if the (AlphaFold-adjusted) trend in the above figure were to continue, then perhaps in two CASPs, i.e. four years, we&#8217;ll actually get to a point where the problem can be called solved, in terms of gross topology (mean GDT_TS ~ 85% or so). Of course, this presupposes that the trendline will continue, and we have no real reason to believe that it will, at least not without new conceptual breakthroughs. Keep in mind that unlike other areas of machine learning, new protein structures are not appearing at an increasing rate, and so waiting things out will not help.</em></p></blockquote><p>I remember the months after CASP13 as some of the most vibrant the field has experienced. Perhaps that&#8217;s the nostalgia of a junior PhD student at beautiful Oxford, right before a global pandemic kept me locked up for some of the best years of my life. But something had shifted. Teams everywhere were building on AlphaFold&#8217;s ideas. And above all, deep learning had become <em>cool</em>: everyone wanted to apply it to biology. What none of us quite grasped was how far this single idea would carry; or where, in the end, it would run out.</p><h3>The AlphaFold 2 era (2021-2024)</h3><p>When I talk to non-experts about &#8220;AlphaFold <em>2</em>&#8221;, they often get confused. For most people there is only <em>one</em> AlphaFold, the one John and Demis received the Nobel for. The first AlphaFold has become a technical footnote of interest only to field experts, and understandably so, because AlphaFold 2 changed everything.</p><p>And I mean that literally. Do you remember the jolt when models like Claude Opus 4.5 first became good enough to hand real coding work to? Take that feeling and multiply it by a hundred. And unlike a model that was merely fast or fluent, back in 2020 AlphaFold 2 did something that <em>should have been impossible</em>: producing the structure of a protein, with high confidence, in a matter of minutes.</p><p>AlphaFold 2 was announced in November 2020, at CASP14, the same blind assessment that its predecessor had shaken two years earlier. Except this time, rather than being a competitor&#8230; it made <a href="https://moalquraishi.wordpress.com/2020/12/08/alphafold2-casp14-it-feels-like-ones-child-has-left-home/">organisers question whether the assessment needed to be carried out ever again</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jVEY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jVEY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 424w, https://substackcdn.com/image/fetch/$s_!jVEY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 848w, https://substackcdn.com/image/fetch/$s_!jVEY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 1272w, https://substackcdn.com/image/fetch/$s_!jVEY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jVEY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png" width="768" height="315" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:315,&quot;width&quot;:768,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jVEY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 424w, https://substackcdn.com/image/fetch/$s_!jVEY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 848w, https://substackcdn.com/image/fetch/$s_!jVEY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 1272w, https://substackcdn.com/image/fetch/$s_!jVEY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3291319e-38b5-4f50-8514-8bf946db45eb_768x315.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ranking of participants in CASP14, as per the sum of the Z-scores of their predictions (provided that these are greater than zero). One group, 427, named AlphaFold 2, shows an incredible improvement with respect to the second best group, 473 (BAKER). One of my colleagues kindly gifted a mug with this image and the sentence &#8220;<a href="https://knowyourmeme.com/memes/you-vs-the-guy-she-told-you-not-to-worry-about">You, and the guy she tells you not to worry about</a>&#8220;</figcaption></figure></div><p>The numbers were almost embarrassing. On the hardest targets, the ones with no template and no close relatives, AlphaFold achieved a near-perfect score. Two experimental groups, comparing the predictions against their own data, realised it was <em>their</em> data they had misread. A third, who had spent two years unable to crack a structure from their crystallography, solved it in an afternoon by <a href="https://en.wikipedia.org/wiki/Molecular_replacement">molecular replacement</a> once they had a prediction to work from. John Moult, who has watched this assessment since 1994 and is not a man given to hyperbole, said the problem could, <a href="https://www.nature.com/articles/d41586-020-03348-4">in a meaningful sense, be considered solved</a>.</p><p>I have a reputation, in more than one circle, as a skeptic. I have spent my professional life wincing through grifter after grifter promising that <em>their</em> model would change the world. So let me be entirely upfront: this was the greatest scientific advance I have ever witnessed. Every single reasonable person I have spoken to in the field would say the same. Having been one of the reviewers of the AlphaFold 2 paper is a fact I intend to tell my grandchildren.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gNiN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gNiN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 424w, https://substackcdn.com/image/fetch/$s_!gNiN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 848w, https://substackcdn.com/image/fetch/$s_!gNiN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 1272w, https://substackcdn.com/image/fetch/$s_!gNiN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gNiN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png" width="556" height="86.01104972375691" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:168,&quot;width&quot;:1086,&quot;resizeWidth&quot;:556,&quot;bytes&quot;:59998,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gNiN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 424w, https://substackcdn.com/image/fetch/$s_!gNiN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 848w, https://substackcdn.com/image/fetch/$s_!gNiN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 1272w, https://substackcdn.com/image/fetch/$s_!gNiN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9b65a26-2262-4f5b-ae52-3782c3804b2e_1086x168.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">A reaction from the CASP14 Discord server, by one of the field&#8217;s most prominent academics. The server was inundated with similar comments by researchers of all levels and tenures, across all stages of grief.</figcaption></figure></div><p>I don&#8217;t think this piece is the right place to talk in depth about the architecture, or the science, because I <a href="https://www.blopig.com/blog/2021/07/alphafold-2-is-here-whats-behind-the-structure-prediction-miracle/">already wrote about it</a> when the model was published (and I have some outdated <a href="https://carlos.outeiral.net/ai4proteins">lecture notes</a>!), as did <a href="https://moalquraishi.wordpress.com/2020/12/08/alphafold2-casp14-it-feels-like-ones-child-has-left-home/">other experts of the field</a>. Instead, I would like to focus on the consequences of the release. I will keep to the basics needed to follow the rest of the essay. The model can be conceived as two main parts. The first, the <em>Evoformer</em>, is a pair of communicating transformers that extracts coevolutionary information from the multiple sequence and uses it to build an abstract, latent representation of structure. The second, the <em>structure module</em>, takes the refined representations and turns them into coordinates. The two modules work in unison through recycling steps. The final output is a full 3D structure of the protein, including side chains (which are added post-hoc by predicting torsion angles with an accessory network).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yF76!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yF76!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 424w, https://substackcdn.com/image/fetch/$s_!yF76!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 848w, https://substackcdn.com/image/fetch/$s_!yF76!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 1272w, https://substackcdn.com/image/fetch/$s_!yF76!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yF76!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png" width="1456" height="490" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:490,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:530224,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yF76!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 424w, https://substackcdn.com/image/fetch/$s_!yF76!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 848w, https://substackcdn.com/image/fetch/$s_!yF76!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 1272w, https://substackcdn.com/image/fetch/$s_!yF76!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49fda7ec-7622-4054-b4d1-61cb64404db9_2157x726.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><p>The paper was released together with a 60-page <a href="https://static-content.springer.com/esm/art%3A10.1038%2Fs41586-021-03819-2/MediaObjects/41586_2021_3819_MOESM1_ESM.pdf">supplementary information</a> describing every method in intricate detail, alongside full inference (but not training!) <a href="https://github.com/google-deepmind/alphafold">code</a>. This format started a tradition that machine learning in biology has followed ever since: the main text became a demonstration of capabilities, while the <em>really</em> interesting content, in far more detail than is conceivable for standard scientific manuscripts, is relegated to the Supplementary Information. More importantly, it answered the questions of a field that had spent over six months wondering whether the AlphaFold team would release the model at all. </p><p>And the moment the code was released, the field began to play. Within days, <a href="https://twitter.com">people on Twitter</a> noticed you could trick the model into predicting protein <em>pairs</em>, simply by feeding it two sequences separated by a large gap, or a glycine linker. That trick worked because the model had quietly learned something it was never built to do: read coevolution <em>between</em> chains, not just within them. DeepMind followed shortly with a proper version, <a href="https://www.biorxiv.org/content/10.1101/2021.10.04.463034v2">AlphaFold-Multimer</a>. It was the first hint of how much signal was hiding in those alignments.</p><p>It was, after several years of very slow progress, the opening of the floodgates for computational molecular biology.</p><h5>Protein structure becomes a commodity</h5><p>The most immediate consequence of AlphaFold 2 is the simplest to state: protein structure stopped being scarce. What once would cost years and a small fortune now cost a few hours and a few dollars, and, thanks in large part to Sergey Ovchinnikov and collaborators&#8217; wonderful <a href="https://github.com/sokrypton/ColabFold">ColabFold</a>, you no longer even had to be a computational scientist to run it. Structure became something you could reach for casually, to play with an idea.</p><p>Then DeepMind went further and simply gave it all away. In a <a href="https://www.nature.com/articles/s41586-021-03828-1">companion</a> to their AlphaFold paper, they released the <a href="https://alphafold.ebi.ac.uk/">AlphaFold database</a>, which today contains nearly 215M predicted structures across more than one million organisms. There is a scene in the documentary <a href="https://www.youtube.com/channel/UC0SOuDkpL6qpIF1o4wRhqRQ">The Thinking Game</a> that follows the &#8220;war room&#8221; as the team decides to predict and release a structure for essentially every protein in UniProt (<a href="https://youtu.be/d95J8yzvjbQ?si=xgti84CGcrqSLuBZ&amp;t=4487">1:14:47</a> in case you are interested). Hundreds of millions of structures, handed to the world, for free. </p><div id="youtube2-d95J8yzvjbQ" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;d95J8yzvjbQ&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/d95J8yzvjbQ?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>That sudden abundance set off a small field of its own: new ways to search and mine this ocean of structures (<a href="https://www.nature.com/articles/s41587-023-01773-0">Foldseek</a>), to organise it (<a href="https://www.science.org/doi/abs/10.1126/science.adq4946">the Encyclopedia of Domains</a>), and even to fold it back into the sequence models it came from, as <a href="https://proceedings.mlr.press/v162/hsu22a">ESM-IF</a> (or, more recently, <a href="https://www.science.org/doi/10.1126/science.ads0018">ESM3</a>) did so beautifully. The abundance of structures has also inspired lots of models, from protein modelling and property prediction (see <a href="https://www.biorxiv.org/content/10.64898/2026.05.28.728196v1.abstract">here</a> for my former PhD student Annie&#8217;s work on this!) to, perhaps more importantly, protein design.</p><h5>Protein design starts to <em>really</em> work</h5><p>One of the ironic things about AlphaFold 2 is that, even though it was supposed to teach us about <em>existing</em> proteins (and it has!), one of its deepest impacts has been on <em>inventing new ones</em>.</p><p>The bottleneck in protein design has always been verifying if a designed protein will succeed. An algorithm can propose hundreds of candidate proteins in an afternoon, but finding out whether any of them actually fold as expected (or at all!) means expressing, purifying and assaying them, which can take months in the best cases, and thousands of dollars. As you can imagine, AlphaFold broke the bottleneck. You could now ask, <em>in silico</em> and in minutes, whether a proposed sequence folds to the shape you intended, and throw away many of the failures before ever lifting a pipette. In practice, this meant that the success rate of designed proteins <a href="https://www.nature.com/articles/s41467-023-38328-5">jumped roughly 10-fold</a> in a matter of months.</p><p>That single capability reorganised the field around what you might call the <a href="https://x.com/AllThingsApx/status/1908513080532754464">central dogma of protein design</a>: generate a backbone, thread a sequence onto it, and check it folds back. Lots of new methods were created to generate backbones <em>de novo</em>, the most popular by far being <a href="https://www.nature.com/articles/s41586-023-06415-8">RFdiffusion</a>, but with a long tail of competitors of which <a href="https://www.nature.com/articles/s41586-023-06728-8">Chroma</a> is my favourite. These &#8220;imagined&#8221; backbones are passed to <a href="https://www.science.org/doi/10.1126/science.add2187">ProteinMPNN</a>, which designs a sequence to fit them, and AlphaFold sits at the end as judge, confirming the sequence folds where it should.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JwZy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JwZy!,w_424,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 424w, https://substackcdn.com/image/fetch/$s_!JwZy!,w_848,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 848w, https://substackcdn.com/image/fetch/$s_!JwZy!,w_1272,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 1272w, https://substackcdn.com/image/fetch/$s_!JwZy!,w_1456,c_limit,f_webp,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JwZy!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif" width="800" height="450" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:450,&quot;width&quot;:800,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6201477,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/gif&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JwZy!,w_424,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 424w, https://substackcdn.com/image/fetch/$s_!JwZy!,w_848,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 848w, https://substackcdn.com/image/fetch/$s_!JwZy!,w_1272,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 1272w, https://substackcdn.com/image/fetch/$s_!JwZy!,w_1456,c_limit,f_auto,q_auto:good,fl_lossy/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa493a1a7-9047-42d5-b269-8e2c7a67187b_800x450.gif 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Animation (<a href="https://www.bakerlab.org/2023/03/30/rf-diffusion-now-free-and-open-source/">source</a>) showing the first version of RFdiffusion generating a protein binder to the insulin receptor. The model generates the backbone (no side chains, so no amino acid identities), and then models like ProteinMPNN assign a protein sequence.</figcaption></figure></div><p>Here is where the story gets really interesting. A designed protein has, by definition, no evolutionary history at all, and therefore no coevolutionary signal to mine. You would expect AlphaFold to be helpless here, but it turns out to be one of the most powerful tools in modern protein design! The reality is that structure prediction is a combination of two problems: a search problem over the astronomically large space of possible protein folds, and a scoring problem that selects the most stable structure. There is a good portion of scientific literature that suggests this is exactly how AlphaFold works: the Evoformer builds a model of the interactions, and <a href="https://journals.aps.org/prl/abstract/10.1103/PhysRevLett.129.238101">the structure module defines which is the right one</a>.</p><p>Designed proteins work so well because the first part of the problem can be removed almost completely. Designers build idealised, well-behaved folds with <a href="https://www.nature.com/articles/35011000">smooth energy landscapes</a>, so the search space converges much more quickly. The structure module then acts like a learned energy function, which is far more powerful than the rudimentary atomistic potentials that we used before. And there are catches, even there. The analysis of designed protein binders in Adaptyv competitions has suggested that the strongest predictor of a successful binder is <a href="https://www.adaptyvbio.com/blog/po104/">similarity to structures already in the PDB</a>. Which is to say: we still measure success by resemblance to what we have already seen.</p><p>There is a lot that could be written about protein design; I strongly recommend <a href="https://www.nature.com/articles/s41586-026-10328-7">this review</a> by David Baker&#8217;s group that highlights all the major achievements in the last few years. For the methodological side, I also co-authored a <a href="https://www.sciencedirect.com/science/article/pii/S0959440X24000216">Current Opinion review</a> on exactly this, led by the brilliant Adam Winnifrith, with Brian Hie.</p><p>But the story is not over here.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://couteiral.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">If you got this far, you enjoyed my essay at least a little. Subscribe (free) for more essays on frontier AI in the life sciences.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>The AlphaFold 3 era (2024-2026)</h2><p>The next update to the field was the release of AlphaFold 3 on May 8th, 2024<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a>. I will once again omit architectural details (I wrote about it <a href="https://carlos.outeiral.net/alphafold-3/">back in the day</a>), but the take-home message is simple: <strong>the architecture changed, the paradigm did not</strong>. The Evoformer became the Pairformer, the structure module gave way to a diffusion decoder, and the model swallowed a much larger chemical vocabulary. Underneath it all, though, the engine was the same one that powered AlphaFold 2: exploit coevolutionary signal, and turn it into coordinates.</p><p>The release itself became a minor scandal. The article was <a href="https://www.nature.com/articles/s41586-024-07487-w">published in Nature</a>, alongside a traditionally detailed supplementary information&#8230; but there was no code, nor weights. This led to a <a href="https://x.com/RolandDunbrack/status/1788262978166596053">small Twitter revolution</a> as the gesture was perceived by many (rightfully so) as a challenge to the basic principles of reproducibility in scientific publishing. Under mounting pressure the team settled on an awkward compromise: code under a noncommercial licence, weights under a licence so restrictive it permits little beyond academic work. </p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VlcT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VlcT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 424w, https://substackcdn.com/image/fetch/$s_!VlcT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 848w, https://substackcdn.com/image/fetch/$s_!VlcT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 1272w, https://substackcdn.com/image/fetch/$s_!VlcT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VlcT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png" width="550" height="214.43873179091688" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:455,&quot;width&quot;:1167,&quot;resizeWidth&quot;:550,&quot;bytes&quot;:97076,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VlcT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 424w, https://substackcdn.com/image/fetch/$s_!VlcT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 848w, https://substackcdn.com/image/fetch/$s_!VlcT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 1272w, https://substackcdn.com/image/fetch/$s_!VlcT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe865d6dc-096f-4f53-86f7-98baf7000974_1167x455.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>What happened next was, of course, that multiple teams set out to reproduce the model, and did so with startling speed. Within a year there were several near-replicas, differing mostly in how openly they were shared. The fastest were <a href="https://www.biorxiv.org/content/10.1101/2024.10.10.615955v1">Chai</a> and <a href="https://www.biorxiv.org/content/10.64898/2026.02.05.703733v1">Protenix</a>, though neither released their weights in the first instance. Much smarter was the move by the folks at <a href="https://www.biorxiv.org/content/10.1101/2024.11.19.624167v1">Boltz</a> (back then at Regina Barzilay&#8217;s group at MIT, not yet a company), who put everything in the open and, by doing so, captured the entire mindshare of the AI-for-biology community more or less overnight. Also open, if a little slower, was the team at <a href="https://github.com/aqlaboratory/openfold-3">OpenFold</a>. While AlphaFold 3 still kept a lead in terms of overall accuracy (though OpenFold<a href="https://portal.openfold.omsf.io/reports/of3p2_technical_report.pdf"> is getting closer</a>), within a year there were multiple models that were 90% there.</p><p>Beyond the gossip, which of course is fun, the most interesting point is not just how AlphaFold 3 was built, but the new doors that it unlocked. AlphaFold 2 had focused on protein-only complexes, the task where coevolution is strongest. AlphaFold 3 pointed the same weapon at harder ground: how proteins meet other molecules, from DNA to drugs. And because these are the testing grounds where the coevolutionary signal provides less information, here we see the first cracks in the model.</p><h5>Small molecules everywhere, all the time</h5><p>The headline capability of this generation is the ability to model small molecules. Isomorphic Labs was not the only player here (for example, Charm Therapeutics had been chasing <a href="https://charmtx.com/dragon-a-top-performing-structure-prediction-model-for-small-molecule-discovery/">similar approaches</a> for a while, with less obvious success) but the ambition was shared: put a drug and its target into the same model and predict the pose.</p><p>Here is the thing about small molecules: they are far harder than they look. Yes, they are small, and their conformational space is trivial next to a protein&#8217;s. But what they lack in size they make up in <em>chemical</em> complexity. There are 20 proteinogenic amino acids (22 if you count <a href="https://en.wikipedia.org/wiki/Selenocysteine">selenocysteine</a> and <a href="https://en.wikipedia.org/wiki/Pyrrolysine">pyrrolysine</a>, the latter found only in some archaea and bacteria), perhaps a few tens more if you include common post-translational modifications like phosphorylation and hydroxylation, and a couple of hundred if you throw in every known one. Small molecules, by contrast, number somewhere around <a href="https://pubs.rsc.org/en/content/articlelanding/2010/md/c0md00020e">10^60 drug-like structures</a>. If you count the stranger things the field has taken to building lately (bifunctionals, anyone?), the number is even greater.</p><p>They are also far less forgiving. One of the most frequent words in the medicinal chemist&#8217;s vocabulary is <a href="https://pubs.acs.org/doi/full/10.1021/jm201706b">activity cliff</a>: the phenomenon where two nearly identical compounds show wildly different activity, usually in binding potency. Here is a real example, from a real <a href="https://onlinelibrary.wiley.com/doi/abs/10.1002/anie.201303207">medicinal chemistry paper</a>. The two molecules in the figure below are part of the chemical series for <a href="https://www.medchemexpress.com/gb/en/Doramapimod.html?srsltid=AfmBOopZ51QZ6Sc-hTm4-p6VVjQEGiimHEYCie5hewiLbyRq5Wu0EAtl">doramapimod</a>, an inhibitor of the <a href="https://en.wikipedia.org/wiki/P38_mitogen-activated_protein_kinases">p38 MAP kinase</a> developed by Boehringer Ingelheim for inflammatory disease that unfortunately failed in Phase III. They differ in just a single methyl group. This tiny difference, just one heavy atom, is enough to change the <a href="https://en.wikipedia.org/wiki/IC50">IC&#8325;&#8320;</a> 250-fold. The structural perturbation that it produces is subtle: it twists the two rings out of plane, forcing the molecule into a shape that somehow fits much better into the pocket.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_OMb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_OMb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 424w, https://substackcdn.com/image/fetch/$s_!_OMb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 848w, https://substackcdn.com/image/fetch/$s_!_OMb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 1272w, https://substackcdn.com/image/fetch/$s_!_OMb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_OMb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png" width="468" height="208.55361216730037" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:586,&quot;width&quot;:1315,&quot;resizeWidth&quot;:468,&quot;bytes&quot;:58073,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_OMb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 424w, https://substackcdn.com/image/fetch/$s_!_OMb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 848w, https://substackcdn.com/image/fetch/$s_!_OMb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 1272w, https://substackcdn.com/image/fetch/$s_!_OMb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97201c12-2d9b-4983-8857-5a8463cdbf93_1315x586.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>Here is the worrying part: this difference should be <em>easy</em> to explain if you have structural methods. A good medicinal chemist can look at the structure and rationalise how the methyl group might change the torsion of the molecule and reduce the strain of the molecule in the pocket. Classical <a href="https://pubs.acs.org/doi/abs/10.1021/jm3003697">computational modelling</a> with physics-based methods can reproduce the torsion, and even estimate the affinity difference with good accuracy. And yet, AlphaFold-like models cannot even predict how the pose changes (perhaps unsurprising, since they also can&#8217;t predict what happens to the pose <a href="https://www.nature.com/articles/s41467-025-63947-5">if you remove the pocket</a>).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HO3a!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HO3a!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 424w, https://substackcdn.com/image/fetch/$s_!HO3a!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 848w, https://substackcdn.com/image/fetch/$s_!HO3a!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!HO3a!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HO3a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png" width="408" height="328.97802197802196" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1174,&quot;width&quot;:1456,&quot;resizeWidth&quot;:408,&quot;bytes&quot;:1094934,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HO3a!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 424w, https://substackcdn.com/image/fetch/$s_!HO3a!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 848w, https://substackcdn.com/image/fetch/$s_!HO3a!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 1272w, https://substackcdn.com/image/fetch/$s_!HO3a!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3eb67786-d58c-4932-920a-6a6a66af0f39_1500x1209.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Boltz-2 prediction of the two ligands above, in complex with the p38 MAP kinase. Both ligands are predicted with the same pose, save for a slight displacement (not a torsion) of the phenyl group. The affinity of the methylated ligand is predicted to be lower (unsurprisingly, as the experimental affinity results are reported on ChEMBL, which was used to train Boltz-2).</figcaption></figure></div><p>These results shouldn&#8217;t be surprising if we remember that AlphaFold-like models are interpolating machines that work by <em>extracting coevolutionary information</em> and <em>turning it into structure</em>. In the small-molecule world, <strong>there is no coevolutionary information to extract</strong>. In fact, a large portion of the success of cofolding methods has been attributed to their ability to locate binding sites, which are known to be <a href="https://arxiv.org/abs/2006.15222">well-encoded in coevolutionary representations</a>. A drug has no evolutionary history, no genomic relatives, no alignment of millions of variants with a hidden signal whispering which residues sit close in space. You also can&#8217;t understand a molecule by its analogues, because even single-atom changes can cause dramatic changes in behaviour. In a world where you had access to lots of counterfactual experiments (very similar molecules with small changes, and matching crystal structures and measurements), then it might be possible to get the models to learn some physics. Perhaps this will work once some model developers gain access to the internal databases of private pharma companies, where there is a wealth of protein-ligand structural and assay data. But today is not that day.</p><p>The one result that certainly wasn&#8217;t on my bingo card is that the problem seems much easier to solve in the other direction. It appears easier to <em>design</em> a protein that binds a small molecule strongly than to <em>predict</em> how a given small molecule binds a given protein (see <a href="https://www.nature.com/articles/s41467-026-70953-8">here</a> and <a href="https://www.nature.com/articles/s41592-025-02626-1">here</a>). That might sound backwards, but reminds us of the lesson from designed proteins: when you get to build an idealised, well-behaved binding site from scratch, you are no longer fighting the model&#8217;s blind spots, but are playing to its strengths.</p><h5>Antibodies become designable, all at once</h5><p><em>This section grew into an essay of its own, <a href="https://couteiral.substack.com/p/the-antibody-revolution-is-real-the">The antibody revolution is real</a>, so I will keep it very short here and point you there for the full story.</em></p><p>The short version is that AlphaFold 3 solved the problem of designing new antibodies, which had been a long-standing field problem. In 2023, AbSci reported the <a href="https://www.absci.com/wp-content/uploads/2023/11/Absci___Unlocking_de_novo_antibody_design.pdf">first generative design of antibodies</a>, which turned out to be a flop. Barely two years later, four frontier companies (Nabla Bio, Latent Labs, Chai Discovery, and Xaira, the last under David Baker&#8217;s wing) independently reported models that design antibodies de novo at high-single to low-double-digit success rates, several of them without a single round of optimisation. A frontier problem became, in the span of a few months, something of a commodity.</p><p>The AlphaFold 3 architecture has likely been one of the differentiators. In the original paper by David Baker&#8217;s group reporting <a href="https://www.nature.com/articles/s41586-025-09721-5">atomic design of antibodies</a>, they observed that AF3 itself could achieve an AUROC of 0.86 (using just AF3&#8217;s internal confidence metric) to predict which antibodies were successful and which ones weren&#8217;t. </p><p>The fact that AlphaFold 3 can work with antibodies so successfully is worth some thought. Antibodies are the cleanest example in protein biology where coevolution does not help. The sequence of antibodies arises from a combination of <a href="https://en.wikipedia.org/wiki/V(D)J_recombination">gene recombination</a> and <a href="https://en.wikipedia.org/wiki/Somatic_hypermutation">somatic hypermutation</a> to maximise variability, which means that there is not a source of evolutionary information that can be exploited. And yet (and this is the pattern, again) the vast majority of antibody design has restricted the structure significantly, needing to predict only a handful of loops. My read is that AlphaFold&#8217;s learned energy function (as crystallised in the diffusion module) is simply powerful enough to elucidate the best conformation when the search space is sufficiently small.</p><h2>The next frontier (2026-?)</h2><p>And now, here is today. The methodology has stood roughly still since AlphaFold 3. There are still a few people innovating on the architecture: for example, Minkyung Baek, first author on the RoseTTAFold paper, has published <a href="https://www.biorxiv.org/content/10.1101/2023.05.24.542179v1.full.pdf">work</a> on alternatives to the standard recipe; and the AlQuraishi group put out <a href="https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1">Genie 3</a> (not to be confused with the <a href="https://deepmind.google/models/genie/">world model</a> that shares the name). Others, like Apple&#8217;s team, have demonstrated that you can make <a href="https://arxiv.org/abs/2509.18480">the architecture even </a><em><a href="https://arxiv.org/abs/2509.18480">simpler</a></em><a href="https://arxiv.org/abs/2509.18480">, just a small variation of the transformer</a>. But these are variations over an existing piece, rather than a completely novel melody.</p><p>One might expect that DeepMind (or Isomorphic Labs, which feeds from very similar people) would once again update the field with a new model architecture. The closest thing we have is the recent <a href="https://storage.googleapis.com/isomorphiclabs-website-public-artifacts/isodde_technical_report.pdf">IsoDDE release</a>, which is AlphaFold 3.5 in all but name. The authors offer little information about the architecture (though they hint that it remains close in spirit to AlphaFold 3), but they show that on the hardest, most out-of-distribution cases, it roughly doubles AlphaFold 3's accuracy. One can hypothesize that a lot of the changes may derive from the army of computational chemists that Isomorphic has spent the last two years hiring, as well as their own proprietary data. From the outside, it is impossible to tell how much of this performance increase is due to the model and which one is to the data. All in all, it seems like IsoDDE is an incremental improvement rather than a step change.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G9AV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G9AV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 424w, https://substackcdn.com/image/fetch/$s_!G9AV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 848w, https://substackcdn.com/image/fetch/$s_!G9AV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 1272w, https://substackcdn.com/image/fetch/$s_!G9AV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G9AV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png" width="418" height="310.02659069325733" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cefad496-6e4a-47ae-9285-7567993da036_1053x781.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:781,&quot;width&quot;:1053,&quot;resizeWidth&quot;:418,&quot;bytes&quot;:131674,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/207650372?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!G9AV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 424w, https://substackcdn.com/image/fetch/$s_!G9AV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 848w, https://substackcdn.com/image/fetch/$s_!G9AV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 1272w, https://substackcdn.com/image/fetch/$s_!G9AV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcefad496-6e4a-47ae-9285-7567993da036_1053x781.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">IsoDDE&#8217;s performance on the <a href="https://www.nature.com/articles/s41594-026-01797-5">Runs N&#8217; Poses</a> benchmark, developed to measure the generalisation of protein-ligand cofolding models to examples far from the training set. The IsoDDE model is significantly better than AF3-era models for ligands that are dissimilar to the training set &#8212; yet it doesn&#8217;t quite solve the problem.</figcaption></figure></div><p>The field is also stalled in answering many of the crucial questions about how biomolecules interact, that go beyond simply predicting the crystal structure.</p><p>One, for example, is protein dynamics. Even if most of our understanding of proteins comes from samples at -200 &#176;C, <em>proteins are not rocks</em>. They flex and flip between conformations, and a lot of valuable functions depend on these conformations. Ever since AlphaFold 2, people have been trying to trick the models into generating ensembles of structures, for example by  <a href="https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1">perturbing the MSA</a>, or more recently, <a href="https://europepmc.org/article/med/42078495">by manipulating the latent space</a>. The most mature line of work is the <a href="https://www.science.org/doi/full/10.1126/science.aaw1147">Boltzmann generators</a> pioneered by Frank No&#233; and others (Frank has since gone to Microsoft AI4Science and built <a href="https://www.science.org/doi/full/10.1126/science.adv9817">BioEmu</a>, among other lovely things), which try to learn the equilibrium distribution of a protein directly from lots of molecular dynamics simulation. The catch is of course that they inherit every defect of molecular dynamics itself, including the <em>very </em>approximate nature of energy functions. Today, I do not think we have made real progress on the ensemble problem at all, not solely because of the methods, but perhaps even more because there are no good assessments with reference data to keep us honest.</p><p>That, by the way, also applies to protein <em>folding.</em> No matter how much DeepMind claims it, or how many press releases suggest it, AlphaFold has <em>not</em> solved the protein folding problem, only the protein <em>structure prediction</em> problem. The folding problem is a different beast entirely, asking <em>how</em> the protein chain gets to that structure, with enormous implications for both basic and applied science. The distinction is not pedantic. I became rather obsessed with it during my PhD (the paper is <a href="https://academic.oup.com/bioinformatics/article/38/7/1881/6517779">here</a>), to the point of building a small benchmark out of the pitifully few experimental examples we have of folding pathways. You can mine a billion sequences and never recover the order in which a chain zips itself together, because that order left no signal in the sequence record.</p><p>So, we seem to have a stalemate. Does anyone have any idea how we are going to make progress? Well, here are some.</p><h5>Prediction 1: physics will be more and more important</h5><p>The entire premise of this essay can be summarised in a sentence: nearly every success in the last decade can be attributed to the power of <em>coevolution</em>. The whole field owes an enormous debt to a few pioneers (David Jones, Debbie Marks, several others). However, as we seem to have exhausted coevolution, we need to start looking elsewhere.</p><p>The problem with coevolution is that it only carries information in some situations. The strongest coevolutionary signals come from an individual protein, because this is highly conserved. Protein-protein interactions come close behind, at least when the interaction between the proteins is sufficiently conserved such that coevolution has imprinted the sequence record. However, there is far less information when we talk of protein-DNA, and essentially none when we talk of protein-ligand interactions. When we are lacking that information, we need to endow the models with new capabilities that extend beyond, much like modern LLMs have learned to use tools to expand their imperfect information corpora. In protein science, that toolkit is <em>physics</em>.</p><p>I am not arguing that the current models do not know any physics, by the way. There is a <a href="https://journals.aps.org/prl/pdf/10.1103/PhysRevLett.129.238101">beautiful paper</a> by James Roney and Sergey Ovchinnikov that shows how AlphaFold 2 gets an idea of the energetics. There is also anecdotal evidence that AlphaFold 3 can make some astounding, physically reasonable predictions, such as <a href="https://x.com/TimothyDuignan/status/1788390250097905978">the structure of electrolyte solutions</a>, or <a href="https://x.com/fenguita/status/1789177480667959728">the assembly of membrane systems</a>. And yet there is also <em>lots</em> of evidence that nearly all known protein or ligand structure models have a tendency to violate the most basic laws of molecular physics. Famous already are the decoys generated by models like <a href="https://arxiv.org/abs/2210.01776">DiffDock</a> that defy the very rules of chemistry.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mo0Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mo0Z!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 424w, https://substackcdn.com/image/fetch/$s_!mo0Z!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 848w, https://substackcdn.com/image/fetch/$s_!mo0Z!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 1272w, https://substackcdn.com/image/fetch/$s_!mo0Z!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mo0Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png" width="400" height="197.4937343358396" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72627615-8216-4586-ae92-da0cde9f1517_1197x591.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:591,&quot;width&quot;:1197,&quot;resizeWidth&quot;:400,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;When machine learning docking goes wrong&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="When machine learning docking goes wrong" title="When machine learning docking goes wrong" srcset="https://substackcdn.com/image/fetch/$s_!mo0Z!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 424w, https://substackcdn.com/image/fetch/$s_!mo0Z!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 848w, https://substackcdn.com/image/fetch/$s_!mo0Z!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 1272w, https://substackcdn.com/image/fetch/$s_!mo0Z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72627615-8216-4586-ae92-da0cde9f1517_1197x591.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption"><em>On the left, an example of a</em> molecular pretzel<em>, a common feature of some recent machine learning docking methods where the structure is entirely unphysical, even though quality scores (e.g. RMSD to crystal structure) suggest it is of good quality. On the right, the same molecule in a physically reasonable conformation. Image reproduced from the <a href="https://posebusters.readthedocs.io/en/latest/">PoseBusters documentation</a>.</em></figcaption></figure></div><p>The failures become even worse when you look at the AlphaFold 3 class of models. <a href="https://www.nature.com/articles/s41467-025-63947-5">One of my favourite papers of last year</a> tested the leading cofolding models against extremely simple adversarial examples. For example, they took <a href="https://en.wikipedia.org/wiki/Cyclin-dependent_kinase_2">CDK2</a>, with ATP bound, and mutated every single residue in the pocket to glycine, effectively deleting the side chains that anchor the triphosphate. And yet, all four models tested in the paper (AlphaFold 3, RoseTTAFold-All-Atom, Chai-1 and Boltz-1) place ATP exactly where it was. When the authors instead mutated all residues to phenylalanine, so that a conglomerate of aromatic rings physically <em>fills </em>the pocket, the models still kept the ligand in the pocket <strong>despite the enormous steric clashes</strong>. </p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4D_U!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4D_U!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 424w, https://substackcdn.com/image/fetch/$s_!4D_U!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 848w, https://substackcdn.com/image/fetch/$s_!4D_U!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 1272w, https://substackcdn.com/image/fetch/$s_!4D_U!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4D_U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png" width="513" height="400.9574175824176" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1138,&quot;width&quot;:1456,&quot;resizeWidth&quot;:513,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Fig. 1: Binding site mutagenesis challenges against co-folding models using the CDK2 system (PDB: 1B38).&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Fig. 1: Binding site mutagenesis challenges against co-folding models using the CDK2 system (PDB: 1B38)." title="Fig. 1: Binding site mutagenesis challenges against co-folding models using the CDK2 system (PDB: 1B38)." srcset="https://substackcdn.com/image/fetch/$s_!4D_U!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 424w, https://substackcdn.com/image/fetch/$s_!4D_U!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 848w, https://substackcdn.com/image/fetch/$s_!4D_U!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 1272w, https://substackcdn.com/image/fetch/$s_!4D_U!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F95587eae-7768-45dd-abf4-46349f607b55_1719x1344.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Results of predicting the CDK2-ATP interactions using different models, reproduced from Figure 1 of <a href="https://www.nature.com/articles/s41467-025-63947-5">Masters </a><em><a href="https://www.nature.com/articles/s41467-025-63947-5">et al</a>.</em> Predicted binding-site residues are shown as cyan sticks, predicted ligand poses are shown as green sticks, and the original co-crystallized ligand pose is shown as gray sticks. In many cases the ligand is still predicted within the binding site and can adopt a low RMSD pose, despite clear alterations to the binding pocket.</figcaption></figure></div><p>People are thinking how to fix this, and there are a few ideas out there. For example, the guys at Boltz introduced a physical steering process. Their <a href="https://www.biorxiv.org/content/10.1101/2025.06.14.659707v1">Boltz-2</a> model can use <a href="https://arxiv.org/abs/2501.06848">Feynman-Kac steering</a> (which is a very fancy way of saying, they constrain the denoising steps) to make sure that the structures are physically realistic according to a classical force field. This helps to reduce mistakes such as steric clashes, but it is still not enough to tell the model not to put the ligand inside of a pocket with only glycines.</p><p>I envision physics entering these models the way tool calling helped LLMs. Not long ago, if you asked a language model to multiply two large numbers it would give you a stupid (not just wrong) answer, because it was matching over text rather than computing anything. The solution was to grant it access to a calculator. Cofolding models are in a similar space: they can generate plausible structures, but have no way to use physics to understand which of those structures make sense. Unfortunately, I don&#8217;t think the solution is as simple as plugging in a physical model (these are <em>very</em> limited in molecular biology, at least). But I am confident that novel ideas that get the models to better understand the priors of the laws of physics will result in better predictions across the board.</p><h5>Prediction 2: biological data will play an increasingly important role</h5><p>The open frontier in biological AI is, and has always been, data. The lesson seems to have been learned quite well in neighbouring areas in cell modelling, with enormous amounts of data. There are many companies collecting massive datasets on cell perturbations (e.g. <a href="https://www.tahoebio.ai/">Tahoe</a>, and at least some of the work that <a href="https://www.xaira.com/">Xaira </a>is doing) or tissue samples from patients (e.g. <a href="https://www.noetik.ai/">Noetik</a>, <a href="https://www.relationrx.com/">Relation</a>, etc). The need for data in the cell biology world is far from forgotten. Somehow, in the more molecular-centric world there has been less of a desire to collect &#8220;machine learning-grade&#8221; data. However, they are not alone.</p><p>One of my favourite companies is <a href="https://www.leash.bio/">Leash Bio</a>, a company that is using <a href="https://en.wikipedia.org/wiki/DNA-encoded_chemical_library">DNA-encoded libraries</a> to generate the world&#8217;s largest dataset of protein-small molecule interactions. The company&#8217;s journey has been the subject of an <a href="https://www.owlposting.com/p/an-ml-drug-discovery-startup-trying">incredible blog post</a> by The Owl Posting that I would recommend to anyone in the field. There are a few problems with this technology, however. Because of the way that DNA-encoded libraries work (tag, enrich, deconvolute with DNA sequencing), the information is very binary: you know if the small molecules bind or not, but it is hard to infer a quantitative affinity readout that tells you which molecules bind more strongly. You also don&#8217;t get any structural information. And, because the DNA tags are tethered to the molecule, there is a fair number of false negatives simply because clashes with the tag preclude binding.</p><p>An approach that is likely to be more successful is <a href="https://openbind.uk/">OpenBind</a>, which is based around the high-throughput X-ray crystallography technology at the UK&#8217;s national synchrotron, <a href="https://www.diamond.ac.uk/">Diamond Light Source</a>, at Harwell. Their objective is to collect, over the next five years, the largest open dataset of protein&#8211;ligand structures ever assembled: on the order of a couple of thousand new complexes a week at full scale, using robotics to soak crystals directly with reaction crudes. Most importantly, every structure will also come with affinity data attached.</p><p>Another is <a href="https://www.aalphabio.com/">A-Alpha Bio</a>, which has a <a href="https://aalphabio.substack.com/p/alphaseq-turning-binding-measurements?utm_source=post-email-title&amp;publication_id=6508565&amp;post_id=209659410&amp;utm_campaign=email-post-title&amp;isFreemail=true&amp;r=54i69p&amp;triedRedirect=true&amp;utm_medium=email">high-throughput yeast display system</a> to measure protein-protein interactions. The team has been using their platform intensively for antibody design (well-suited, because yeast display operates in the exterior of the cell, it captures a lot of the environmental factors important for antibodies such as oxidation state, etc). As I was editing this piece, they announced <a href="https://endpoints.news/a-alpha-bio-starts-ai-antibody-data-consortium-with-gsk-boltz/">a partnership with GSK, Boltz, Cradle and Dyno Therapeutics</a>, to generate data to power biomolecular AI models.</p><p>I believe that these approaches are all going in the right direction: generating data that fills in the gaps left by coevolution. However, something that I have learned in many years in the field is that data <em>quality</em> and data <em>abundance</em> are often at opposite ends. I have also learned that consortia can start with incredible momentum and soon fade as interest (or grant money) runs thin. I don&#8217;t think that the question of generating the right data is quite solved.</p><h5>Prediction 3: progress will stagnate because of siloing</h5><p>I have been in the field long enough to remember the time of stagnation. Remember that for nearly a decade between ~2000 and 2010, our protein structure prediction abilities remained roughly unchanged. In the following decade, until the AlphaFold breakthrough, the little progress that was made happened one CASP at a time. Because every group wanted to be the one to come out first, progress updates were withheld until the biennial assessment, and the participant groups were forced to independently rediscover the wheel over and over.</p><p>Now, imagine if the information wasn&#8217;t shared at all. A lot of the field&#8217;s progress in the last few years has stemmed from DeepMind/Isomorphic&#8217;s releases providing a &#8220;gradient update&#8221; to our architectural understanding. Unfortunately, Isomorphic Labs seems to have learned from the AF3 fiasco and has decided, at least for their latest IsoDDE model, not to put out any more information than necessary. </p><p>The problem is likely to get worse. One of the recent headlines was of Isomorphic Labs <a href="https://www.isomorphiclabs.com/articles/isomorphic-labs-announces-series-b-investment-round">raising over 2 billion USD</a>, taking the total amount of funding they secured in just over a year to well over 2.5 billion dollars. Together with some of the other funding rounds recently e.g. <a href="https://www.fiercebiotech.com/biotech/new-ai-drug-discovery-powerhouse-xaira-rises-1b-funding">Xaira&#8217;s 1 billion dollar seed</a>, and Chai&#8217;s <a href="https://www.businesswire.com/news/home/20260713849009/en/Chai-Discovery-Announces-%24400M-Series-C-to-Advance-AI-Driven-Molecular-Design">recent 400M USD round</a>, it is clear that there are many players with extremely deep pockets. This, by the way, makes complete sense<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a>: these huge raises enable them to focus on progress without the distractions of constant fundraising, especially in a field like biotech where it is crucial to have multiple shots on goal. However, this will also lead to another issue: talent concentration.</p><p>Think about what is happening in the natural language world: nearly all the frontier research is being carried out inside private corporations. While it is well-known that research secrets in the Bay Area travel like a cough in flu season, this is a consequence of the <a href="https://www.newyorker.com/magazine/2018/10/22/did-uber-steal-googles-intellectual-property">particular legal structure of California competition law</a>. In biotech the same conditions do not hold. The money is spread across jurisdictions where non-competes are perfectly enforceable, and the culture is built around patents rather than preprints, because the key output (the <em>only</em> output that ultimately matters) is the drug. If enclosure in language models is porous, in biology it will not be.</p><p>So my third prediction is a paradox: more money than the field has ever had will lead to less progress per pound than it has ever managed, because the ideas that would compound stop moving between the people who have them.</p><h2>Conclusions</h2><p>AlphaFold 2 is the greatest scientific advance I have witnessed in my career&#8230; which is not something I expect to say about any model for a long time. However, five years after the Nature paper was published, and having witnessed two new generations of the model, I have been worrying that the field is reaching a stalemate. The problem is that coevolution, the single idea underpinning the technology, has given us nearly everything it has to give.</p><p>Coevolution was a gift: evolution spent four billion years running an optimisation experiment on every protein, and left a faint statistical shadow in the sequence record. The genius in AlphaFold was learning to extract this information from multiple sequence alignments, and transform it into a protein structure. But the same success that made the model excel at protein structure prediction is not providing the information that it needs to excel. Like a tourist with a phrasebook, it often produces useful sentences; but it also fails disastrously at grammar.</p><p>There are, as far as I can see, only two paths to success. One is to teach the models the physics that they need. They would need to understand enough about the physical world to avoid making atoms fatally crash, while retaining enough of its latent space behaviour to excel at the task; kind of like humans, navigating between a system 1 and a system 2 and using the best of both worlds. Unfortunately, this is something that has been preached for years but never achieved in practice. The other is to generate, deliberately and at industrial scale, the data to fill in the gaps left by evolution. Data of multiple sorts, including structural as well as energetic data. Unfortunately, that data does not come cheap.</p><p>Coevolution took us further than anyone in 2017 could have dared to imagine. Here, I have argued that it took us as far as it can. The next breakthroughs in AI for proteins belong to those who break the barrier.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://couteiral.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free if you would like to read more essays on frontier AI and the life sciences.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>For brevity reasons, I have omitted some of the story with regards to deep learning appearing around CASP12 (2016). In that edition, Jinbo Xu&#8217;s <a href="https://pubmed.ncbi.nlm.nih.gov/31298436/">RaptorX</a> method, which used a convolutional neural network to predict protein-protein contacts, achieved the best performance, and was shortly followed by similar methods like David Jones&#8217; <a href="https://onlinelibrary.wiley.com/doi/10.1002/prot.25779">DeepMetaPSICOV</a>. These methods were predicting <em>binary</em> contacts <em>i.e.</em> predicting which residue pairs were within 8 &#197; of one another. Jinbo Xu also published the first (as far as I know) neural network <a href="https://www.pnas.org/doi/10.1073/pnas.1821309116">predicting protein distances</a>, rather than binary contacts, which did very well at CASP13 (2018). Despite the deep impact that all these contributions had on the field at that time, I have chosen to omit these details to simplify the story.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><a href="https://www.ibbr.umd.edu/profiles/john-moult">John Moult</a>, who founded the assessment, would frown at calling it a &#8220;competition&#8221;, so I will refrain from saying AlphaFold &#8220;won&#8221;.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>I remember the date with some accuracy, because it is my birthday.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>For the companies. For the investors&#8230; we&#8217;ll see.</p></div></div>]]></content:encoded></item><item><title><![CDATA[The antibody revolution is real. The business, less so.]]></title><description><![CDATA[7.1k words, 28-35 min reading time]]></description><link>https://couteiral.substack.com/p/the-antibody-revolution-is-real-the</link><guid isPermaLink="false">https://couteiral.substack.com/p/the-antibody-revolution-is-real-the</guid><dc:creator><![CDATA[Carlos Outeiral]]></dc:creator><pubDate>Thu, 02 Jul 2026 18:42:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a5f2b5bd-dcf1-448b-adde-451cd1efa012_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I have been meaning to write about the recent antibody revolution. The problem is, you see, that I wanted to make sure it was real. AI for drug discovery has tended to experience revolutions the way other industries have quarterly results: four times a year, always the one that will finally change everything. After a decade in the field, you build a reflex: you read the paper, then look for the catch.</p><p>Take the 2023 antibody &#8220;revolution&#8221;. Reports of the first <a href="https://www.biorxiv.org/content/10.1101/2023.01.08.523187v1">&#8220;</a><em><a href="https://www.biorxiv.org/content/10.1101/2023.01.08.523187v1">de novo</a></em><a href="https://www.biorxiv.org/content/10.1101/2023.01.08.523187v1"> designed&#8221; AI antibody</a> by <a href="https://www.absci.com/">AbSci</a> made it to <a href="https://www.forbes.com/sites/johncumbers/2023/01/10/this-company-is-using-generative-ai-to-design-new-antibodies/">international newspapers</a> and created a frenzy on social media. The catch? They made <strong>a variant of <a href="https://en.wikipedia.org/wiki/Trastuzumab">trastuzumab</a></strong>, an antibody targeting <a href="https://en.wikipedia.org/wiki/HER2">HER2</a> in breast cancer, <strong>approved in 1998</strong>, with just one loop replacement. Also, it was not really <em>de novo</em>: they screened<strong> half a million variants</strong>, to find only three winners&#8230; the bar being, again, improving on a twenty-five year old molecule! The experimental design also had <a href="https://www.linkedin.com/posts/surge-biswas-a8b61270_surge-biswas-on-twitter-activity-7019046023041839104-oDk4/">clear caveats</a>, including lack of appropriate controls. But despite all of this, the <a href="https://investors.absci.com/news-releases/news-release-details/absci-first-create-and-validate-de-novo-antibodies-zero-shot">press release</a> read like the biology equivalent of the moon landing.</p><p>Then, last year, not one but <em>four</em> companies again announced the ability to design <em>de novo</em> antibodies using AI/ML: <a href="https://www.biorxiv.org/content/10.1101/2025.07.05.663018v1.full.pdf">Chai Discovery</a>, <a href="https://nabla-public.s3.us-east-1.amazonaws.com/2025_Nabla_JAM2.pdf">Nabla Bio</a>, <a href="https://arxiv.org/abs/2512.20263">Latent Labs</a> and Xaira Therapeutics (though licensing work from <a href="https://www.nature.com/articles/s41586-025-09721-5">David Baker&#8217;s lab</a>). If you remember the trastuzumab fiasco, you surely can&#8217;t blame me for facing the news with skepticism. One of the founders of Chai, Joshua Meier (a solid researcher originally from the Meta team that developed the popular <a href="https://www.pnas.org/doi/abs/10.1073/pnas.2016239118">protein language model ESM</a>) was also the corresponding author on the infamous AbSci technical report. There is something new, though. This time, the different labs have all reported <strong>novel</strong> antibodies against <strong>multiple, novel</strong> targets, with little to no optimization. The hit rates are now in the double digits, and the overall pharmacological properties of these antibodies start to look like proper drugs. And an increasing number of offshoot competitors, and even some academic groups, are reporting similar results.</p><p>This time, I think, the science is real. But the excitement it deserves has been buried under hype from an AI-heavy commentariat that has mostly not bothered to ask what it means for drug discovery. Some of the things they are missing are:</p><ul><li><p>why the <em>de novo</em> antibody discovery technology is a genuine scientific achievement</p></li><li><p>why it might not necessarily result in the industry making better drugs</p></li><li><p>why it probably is not a good business, unless these companies think very carefully about how to create value</p></li></ul><p>I would like to address this today.</p><h2>Why we care about antibodies</h2><p>Most of the excitement rests on an assumption that seems obvious if you haven&#8217;t been in the field before: that designing an antibody and designing a drug are the same problem. That thesis falls right at the outset. While it is true that antibodies <em>can</em> be drugs, or perhaps more accurately, that <em>some drugs</em> are antibodies, most antibodies are not, and could not, be drugs. The argument is not trivial, and my best chance to get it through is by starting with what an antibody actually is.</p><p><a href="https://en.wikipedia.org/wiki/Antibody">Antibodies</a> are natural proteins expressed by your immune system to fight off invaders. They are frankly incredible, the byproduct of millennia of evolution to produce a lean, mean, defensive machine that can recognise and inactivate a variety of threats. For this essay, what matters most is that <strong>an antibody can recognise a </strong><em><strong>specific</strong></em><strong> target protein with incredibly strong affinity while leaving almost anything else untouched</strong>. If you have ever taken a <a href="https://en.wikipedia.org/wiki/Lateral_flow_test">pregnancy test</a> (or a CoVID lateral flow test), you have been using a strip impregnated with an antibody that immediately latches on to the target protein while ignoring everything else. This <a href="https://en.wikipedia.org/wiki/Magic_bullet_(medicine)">magic bullet</a> property is also what has made them incredibly promising as drugs.</p><p>The implication of an antibody drug is that a specific (or, as we say, <a href="https://en.wikipedia.org/wiki/Monoclonal_antibody">monoclonal</a>) antibody can be designed and injected to attack some proteins in the human body. Any proteins? Not really. Because they are large and bulky, antibodies cannot cross the membrane that separates a cell from its surroundings, so they are limited to act on proteins displayed on the cell&#8217;s exterior (there are <a href="https://insight.jci.org/articles/view/127474">some variants</a> that can get <em>inside</em> the cell, but they are still prototypes, and none have been approved by the FDA yet). Limiting to extracellular proteins is fine because they are roughly one fifth of all human proteins, and they are disproportionately represented in disease.</p><p>How do these antibody drugs work? Take Humira (<a href="https://en.wikipedia.org/wiki/Adalimumab">adalimumab</a>), one of the largest-grossing drugs in history. Humira acts as a <em>protein-protein interaction inhibitor</em> for a protein called <a href="https://en.wikipedia.org/wiki/Tumor_necrosis_factor">tumor necrosis factor</a> (TNF). TNF is an alarm signal for the cell: whenever there is an infection or injury, cells release it to switch on inflammation, recruiting the immune system and opening the body&#8217;s resources to combat damage. This is usually a good thing. But if you have an autoimmune disease like rheumatoid arthritis, or Crohn&#8217;s disease, or psoriasis, or many others, your body is stuck generating TNF and turning on itself, and switching it off has therapeutic value.</p><p>And that&#8217;s exactly what happens when you get a Humira injection. The antibody latches onto the TNF protein and coats it, preventing it from reaching the immune system cells that it would otherwise activate and switching off the autoimmune disorder. Because of the generality of the mechanism, the enormous patient population and (let&#8217;s not beat around the bush) the fact that it needs to be taken regularly because autoimmune diseases are chronic, Humira became one of the best selling drugs in history, peaking at $20B/year before it went off-patent in 2023.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4i5S!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4i5S!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 424w, https://substackcdn.com/image/fetch/$s_!4i5S!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 848w, https://substackcdn.com/image/fetch/$s_!4i5S!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 1272w, https://substackcdn.com/image/fetch/$s_!4i5S!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4i5S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png" width="1326" height="812" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:812,&quot;width&quot;:1326,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1261649,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4i5S!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 424w, https://substackcdn.com/image/fetch/$s_!4i5S!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 848w, https://substackcdn.com/image/fetch/$s_!4i5S!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 1272w, https://substackcdn.com/image/fetch/$s_!4i5S!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3bac7a96-0165-435f-a22d-83f9d2613e83_1326x812.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Infographic describing the mechanism of Humira (adalimumab). The antibody binds free tumor necrosis factor (TNF) in the bloodstream, blocking it from binding the TNF receptor on immune cells and inducing inflammation. Infographic lightly edited from a GPT Image-2 generation, June 2026.</figcaption></figure></div><p>Since we are talking about Humira, it may be worth talking about the <em>king</em> of antibodies, the one that will figure in the pitch deck of any AI for antibody discovery company. Keytruda (<a href="https://en.wikipedia.org/wiki/Pembrolizumab">pembrolizumab</a>, ~$32B in revenue in 2025) targets a protein called <a href="https://en.wikipedia.org/wiki/Programmed_cell_death_protein_1">PD-1</a>, which is expressed by T-cells, specialised cells of the immune system that hunt down other abnormal cells, like tumour cells. PD-1 works like a password recogniser: it inspects healthy cells for another protein called <a href="https://en.wikipedia.org/wiki/PD-L1">PD-L1</a> that flags them as healthy. Unfortunately, tumour cells learn this trick quite quickly and express PD-L1 on their surface to pass the check. Keytruda blocks PD-1 so the false password is never read, and the T-cell attacks the tumour anyway.</p><p>If you are paying attention, you may now be wondering why these T-cells do not attack healthy cells. Well, they do. It is one of the major side effects. But, as it often happens in drug discovery, smart engineering can adjust a molecule to produce a <a href="https://en.wikipedia.org/wiki/Therapeutic_index">very large effect on diseased cells while only barely inconveniencing healthy ones</a>. In effect, the Phase III clinical trials for Keytruda (first for <a href="https://www.nejm.org/doi/full/10.1056/NEJMoa1503093">melanoma</a>, then for many others) showed a <em>doubling</em> in the overall survival, with respect to the standard of care, and with relatively minimal side effects.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YroV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YroV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 424w, https://substackcdn.com/image/fetch/$s_!YroV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 848w, https://substackcdn.com/image/fetch/$s_!YroV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 1272w, https://substackcdn.com/image/fetch/$s_!YroV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YroV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png" width="1456" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/937beeef-00e9-4360-a886-217528c9341f_1693x929.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1847828,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YroV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 424w, https://substackcdn.com/image/fetch/$s_!YroV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 848w, https://substackcdn.com/image/fetch/$s_!YroV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 1272w, https://substackcdn.com/image/fetch/$s_!YroV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F937beeef-00e9-4360-a886-217528c9341f_1693x929.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Infographic describing the mechanism of Keytruda (pembrolizumab). The antibody binds the PD-1 receptor on the surface of T cells, forcing them to destroy tumour cells. Infographic generated using GPT Image-2, June 2026.</figcaption></figure></div><p>More broadly, there is a reason to care about antibodies that goes beyond any single drug: they help us push back against <a href="https://en.wikipedia.org/wiki/Eroom%27s_law">Eroom&#8217;s law</a>, the observation that drug discovery is becoming slower and more expensive over time. Antibodies have significantly higher success rates than small molecules. One of the most extensive studies found that <strong>they are around 50% more likely to be approved</strong> (<a href="https://go.bio.org/rs/490-EHZ-999/images/ClinicalDevelopmentSuccessRates2011_2020.pdf">12% vs 8% Phase I-to-approval</a>), which is driven mostly by success at Phase II/III and regulatory review. In other words, an antibody program is far likelier to reach patients than a small-molecule one, and development times have also <a href="https://www.nature.com/articles/s41587-020-0512-5">shortened significantly</a>.</p><p>Antibodies also come in increasingly sophisticated variants. <a href="https://en.wikipedia.org/wiki/Bispecific_monoclonal_antibody">Bispecific antibodies</a>, for example, are designed to have different arms with different specificity for different proteins. Perhaps the best example is <a href="https://en.wikipedia.org/wiki/Emicizumab">emicizumab</a>, which brings together two factors in the coagulation cascade and which has been transformative in the treatment of hemophilia. And the last few years there have seen a frenzy around <a href="https://www.nature.com/articles/s41392-022-00947-7">antibody-drug conjugates</a>, which are antibodies with a cytotoxic attachment, where the antibody identifies a cancerous cell and delivers their payload. But the near totality of work in the AI/ML field has focused on two types of antibodies, <a href="https://en.wikipedia.org/wiki/Single-domain_antibody">nanobodies</a> and <a href="https://en.wikipedia.org/wiki/Single-chain_variable_fragment">single-chain variable fragments</a> (scFvs), that follow the principles we discussed above.</p><p>Let&#8217;s leave it at this: antibodies are protein drugs that target proteins on the cell&#8217;s exterior, with high specificity, the ability to reach undruggable targets, and better clinical success rates than small molecules. They are, frankly, remarkable. But there are some caveats.</p><h2>What it entails to make an antibody</h2><p>I am saying a lot of great things about antibodies here, which is not really my style. I am leaving aside several problems here for the sake of brevity, although they are extraordinarily important: the cost of manufacturing (roughly ~100x more expensive than small molecules), the supply chain issues (need to be refrigerated) and the pharmacology problems (can&#8217;t be orally dosed, need to be injected/infused or in the best case administered by sub-cutaneous injection, which is painful and inconvenient). None of these are trivial for the ultimate mission of medicine, which is to give universal, affordable care to everyone.</p><p>But there is one really great thing about antibodies: <em>immune systems know how to make them</em>. Let me explain. In the realm of <a href="https://en.wikipedia.org/wiki/Small_molecule">small molecule drugs</a>, where most of pharma still spends most of their time, there are rarely guarantees of success. Many targets are <a href="https://en.wikipedia.org/wiki/Druggability">undruggable</a>, which means they are proteins without a <a href="https://en.wikipedia.org/wiki/Binding_site">druggable pocket</a> where a small molecule can bind. Some of these undruggable targets (<a href="https://en.wikipedia.org/wiki/Myc">Myc</a>, <a href="https://en.wikipedia.org/wiki/Catenin_beta-1">&#946;-catenin</a>, <a href="https://en.wikipedia.org/wiki/Alpha-synuclein">&#945;-synuclein</a>, amongst many others) have been known for decades and we are no closer to find an effective drug for them than we were fifty years ago. In contrast, it is almost guaranteed that given some extracellular target, you can make <em>some</em> antibody that binds to it.</p><p>The most common technology for antibody design can be summarised in: find a way to <strong>expose an organism to some foreign protein, and wait until their immune system engineers an antibody against it</strong>. The classical approach is the <a href="https://en.wikipedia.org/wiki/Hybridoma_technology">hybridoma technology</a>, which relies on a simple idea: take an animal (generally a mouse, or a rabbit), inject it with the protein you want to raise antibodies against, then look at the antibodies that the animal generated and isolate one. In recent years there have been a few improvements, such as using <a href="https://en.wikipedia.org/wiki/Humanized_mouse">transgenic humanised mice</a> approaches where the immunoglobulin genes from the mice have been replaced with human ones. This has the advantage that the developed antibodies are more human-like, and therefore less likely to have side effects such as <a href="https://www.frontiersin.org/journals/immunology/articles/10.3389/fimmu.2024.1401178/full">anti-drug antibodies</a>. While it has some limitations (for example, you generally can&#8217;t make an antibody against an endogenous protein), it is an overwhelmingly successful technology for antibody discovery.</p><p>Of course, rather than injecting a foreign entity, you can also just look at humans that have naturally evolved ones. The whole field of <a href="https://www.nature.com/articles/s41592-024-02243-4">natural immune repertoire analysis</a> relies on this, having led to multiple <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC9631452/">therapies for infectious disease</a>. Alternatively, you can use something called <a href="https://en.wikipedia.org/wiki/Phage_display">phage display</a>, or its close cousins <a href="https://en.wikipedia.org/wiki/Yeast_display">yeast</a> and <a href="https://en.wikipedia.org/wiki/Ribosome_display">ribosome display</a>, to test hundreds of millions of randomly generated antibodies to see if one of them binds the target. In these display technologies, antibody fragments are expressed fused to a coat protein on the surface of an organism, with the encoding gene packaged inside the particle (cell or phage). Modern setups can screen around one trillion antibodies in around two weeks.</p><p>There are lots of caveats, the main one being: the antibodies generated by these methods are <em>binders</em>, and a lot of work may be required to turn them into drugs. We will come back to this. But for now, let&#8217;s think about how AI/ML methods can change the way we come up with novel antibodies.</p><h2>The new era of antibody design</h2><p>Let&#8217;s go back to the breakthroughs of 2025. Four frontier BioAI companies (Nabla.Bio, Latent Labs, Chai Discovery, and Xaira Therapeutics, the latter under the wing of <a href="https://en.wikipedia.org/wiki/David_Baker_(biochemist)">David Baker</a>&#8217;s lab) reported models capable of designing antibodies <em>de novo</em>, with high-single to low-double digit success rates. </p><p>The sequence started with a <a href="https://www.nature.com/articles/s41586-025-09721-5">paper from David Baker&#8217;s lab</a> (preprinted in March 2025, published in Nature in October 2025), presenting the technology that was licensed to Xaira Therapeutics and likely was behind their <a href="https://www.fiercebiotech.com/biotech/new-ai-drug-discovery-powerhouse-xaira-rises-1b-funding">billion dollar seed round</a>. Their paper was the first demonstration that you could find a binder from scratch at all: zero-shot and against real, but well-trodden targets (hemagglutinin, toxin B from <em>C. difficile</em>, and others), although with relatively weak potencies (single-digit micromolar, or triple-digit nanomolar) and a hit rate around 1-2%. Just three months later, Chai Discovery published <a href="https://www.biorxiv.org/content/10.1101/2025.07.05.663018v1">a technical report</a> showing much higher-quality zero-shot antibodies, this time against unprecedented targets, with much higher potencies (generally single-digit to double-digit nanomolar). The hit rate was also much higher, reported as double digit, although of course that depends on the definition of &#8220;hit&#8221; (which here, without in any way diminishing their accomplishment, was generous).</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!408R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!408R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 424w, https://substackcdn.com/image/fetch/$s_!408R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 848w, https://substackcdn.com/image/fetch/$s_!408R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 1272w, https://substackcdn.com/image/fetch/$s_!408R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!408R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png" width="639" height="491.0740157480315" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:976,&quot;width&quot;:1270,&quot;resizeWidth&quot;:639,&quot;bytes&quot;:178105,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!408R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 424w, https://substackcdn.com/image/fetch/$s_!408R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 848w, https://substackcdn.com/image/fetch/$s_!408R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 1272w, https://substackcdn.com/image/fetch/$s_!408R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85f30df9-9bd9-4ba8-9c66-9a0e99a634b2_1270x976.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Overview of targets and success rates reported by Chai in Figure 3 of their <a href="https://www.biorxiv.org/content/10.1101/2025.07.05.663018v1.full.pdf">technical report</a>. The team reports double-digit success rates across nearly half of the targets they tried, many of them novel and without any data in the literature. That said, the definition of &#8220;hit&#8220; is relatively loose so, while remarkable, the real hit rates may be lower.</figcaption></figure></div><p>The Chai technical report was impressive, but the story doesn&#8217;t stop there. Towards the end of the year, Nabla.Bio released <a href="https://nabla-public.s3.us-east-1.amazonaws.com/2025_Nabla_JAM2.pdf">another technical report</a> where they showed an improved model, with similar hit rates to Chai-2, but with more rigorous benchmarking, and succeeding at some of the targets where their competitor had failed. They also reported not just antibody <em>binders</em>, but they looked at many of the properties that make an antibody a good drug (we&#8217;ll talk about this in a second). For example, they report wet lab assays for thermal stability, monomericity, expression titer, and many others, where the antibodies look reasonably attractive. And if this sounds amazing, then it was topped by Latent Labs, right before Christmas, with another <a href="https://arxiv.org/abs/2507.19375">technical report</a>, showing similar results to other labs but with a pretty extensive amount of developability data, including even a demonstration that their antibodies had low immunogenicity when tested on T cell samples from healthy volunteers (although of course this is poorly predictive of what might happen in a real patient).</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WyP6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WyP6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 424w, https://substackcdn.com/image/fetch/$s_!WyP6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 848w, https://substackcdn.com/image/fetch/$s_!WyP6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 1272w, https://substackcdn.com/image/fetch/$s_!WyP6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WyP6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png" width="1456" height="282" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:282,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:114409,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WyP6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 424w, https://substackcdn.com/image/fetch/$s_!WyP6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 848w, https://substackcdn.com/image/fetch/$s_!WyP6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 1272w, https://substackcdn.com/image/fetch/$s_!WyP6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37d2b3c9-5e07-4f6a-9188-657f1dbad019_1787x346.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Developability patterns reported by Latent Labs in Figure 2b of their <a href="https://arxiv.org/pdf/2512.20263">technical report</a>. The properties look generally good across designs; yield and monomericity are remarkable, and thermal stability is really good. Hydrophobicity and, specially, polyreactivity, still need some work.</figcaption></figure></div><p>While I was editing this essay, Nabla published <a href="https://nabla-public.s3.us-east-1.amazonaws.com/2026_Nabla_JAM2_multispecific_pMHC.pdf">yet another technical report</a> where they used their JAM-2 model to design antibodies against <a href="https://en.wikipedia.org/wiki/Major_histocompatibility_complex">peptide-MHC (pMHC) complexes</a>. Some background will be useful here: vertebrate cells are continuously breaking apart some of their proteins, and taking small fragments (peptides) that are presented in the surface of the cell via the major histocompatibility complex (MHC). In turn, these complexes are recognised by cells of the immune system, which identify potentially foreign cells and destroy them. This system, however, is not perfect. For example, mutations in <a href="https://en.wikipedia.org/wiki/KRAS">KRAS</a>, a protein that is mutated in 20% of human cancers and has been one of the most coveted drug targets for decades, do not elicit an immune response because a single change is often insufficient to make the peptide-MHC complex reliably distinguishable as &#8220;foreign&#8221; by T cells.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WF-r!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WF-r!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 424w, https://substackcdn.com/image/fetch/$s_!WF-r!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 848w, https://substackcdn.com/image/fetch/$s_!WF-r!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 1272w, https://substackcdn.com/image/fetch/$s_!WF-r!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WF-r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png" width="506" height="328.26366001734607" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:748,&quot;width&quot;:1153,&quot;resizeWidth&quot;:506,&quot;bytes&quot;:1089356,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9956651-b87a-4b0a-88eb-f3f66022d693_1153x863.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WF-r!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 424w, https://substackcdn.com/image/fetch/$s_!WF-r!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 848w, https://substackcdn.com/image/fetch/$s_!WF-r!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 1272w, https://substackcdn.com/image/fetch/$s_!WF-r!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7005047-7a45-43f0-9091-839b3b44f0b0_1153x748.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Infographic describing the mechanism of the major histocompatibility complex (MHC), bound to peptides. Peptides from naturally degraded proteins are presented in the surface of the cell for recognition by the immune system. Infographic lightly edited from a GPT Image-2 generation, June 2026.</figcaption></figure></div><p>And here is the power of Nabla&#8217;s model: they <strong>managed to design one antibody that can selectively recognise pMHC complexes of two mutants of KRAS</strong> (G12V and G12C, the second and third most common) <strong>while leaving wild-type KRAS essentially untouched</strong>. This is an incredible achievement for two reasons. One, they are the first report, as far as I know, of a computationally designed pMHC-antibody. This is really hard, because binding a peptide is much harder than binding a globular protein, because of its flexibility and because one needs to engage specific atoms. Making it selective is even harder. But, second and perhaps more importantly, this opens the possibility of designing antibodies <em>against intracellular proteins </em>(with lots of caveats like HLA-restriction, the low density of the pMHC on the surface, and the fact that the broader TCR-mimic field has a graveyard all its own), which opens many doors in diseases with highly unmet medical need.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0z4N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0z4N!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 424w, https://substackcdn.com/image/fetch/$s_!0z4N!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 848w, https://substackcdn.com/image/fetch/$s_!0z4N!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 1272w, https://substackcdn.com/image/fetch/$s_!0z4N!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0z4N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png" width="310" height="179.35714285714286" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:324,&quot;width&quot;:560,&quot;resizeWidth&quot;:310,&quot;bytes&quot;:204942,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0z4N!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 424w, https://substackcdn.com/image/fetch/$s_!0z4N!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 848w, https://substackcdn.com/image/fetch/$s_!0z4N!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 1272w, https://substackcdn.com/image/fetch/$s_!0z4N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F65881f69-4d00-442d-9292-e6c836c6dbb6_560x324.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Crystal structure of the Nabla-designed pMHC antibody. The circle highlights the side chain of the amino acid Val12. The only difference from the wild-type peptide (Gly12) is that the valine side chain inserts into a complementary VHH pocket, the basis for variant-selective recognition. Reproduced from the Figure 4 of <a href="https://nabla-public.s3.us-east-1.amazonaws.com/2026_Nabla_JAM2_multispecific_pMHC.pdf">Nabla&#8217;s technical report</a>.</figcaption></figure></div><p>The experimental data released by all of these labs is, frankly, one of the most thorough technological demonstrations I have seen in some time. In AI for science, the proof is in the pudding, and <strong>generating novel chemical matter for a novel target that other people have been trying to solve for a long while is as high a bar as you can possibly reach for</strong>. If you do it for a dozen, previously unsuccessful targets, well, that is even more impressive! While ignoring the caveats (sure, their hit rates will be lower once you discard the weak signals, or the sensograms that look weird), I think that this is an <em>extremely solid</em> demonstration that the technology is working. It is also quite impressive to see the field advance in real time throughout two years, as the technology improved and as the bar to make a splashy announcement kept moving higher and higher.</p><p>We do not know all the details of the private models because they are, well, private. We can however infer some information from the paper by David Baker&#8217;s lab, which I had the pleasure to peer-review for Nature. At first glance, pure architectural innovations in David Baker&#8217;s paper are relatively modest. They use their celebrated diffusion model <a href="https://www.nature.com/articles/s41586-023-06415-8">RFdiffusion</a>, together with some fine-tuning on published antibody structures, to design a large number of structural decoys. These decoys have a structure, but not a specific set of amino acids (<em>i.e.</em> they design only the <a href="https://en.wikipedia.org/wiki/Amino_acid">alpha-carbons</a> and omit side chains), so the <a href="https://www.science.org/doi/10.1126/science.add2187">ProteinMPNN</a> neural network is used to generate a protein sequence that may fold into that structure. Finally, they used a fine-tuned version of their <a href="https://www.biorxiv.org/content/10.1101/2023.05.24.542179v1">RoseTTAFold 2</a> model (which you can understand as an alternative to AlphaFold 2) to predict the structure of the complex, as well as the probability of binding using a specially designed and trained head, to validate that the generation is correct. There are lots of small technical details to how this worked in practice, but the general pipeline follows the <a href="https://x.com/AllThingsApx/status/1908513080532754464?s=20">central dogma of protein design</a> or RFdiffusion -&gt; ProteinMPNN -&gt; AlphaFold 2.</p><p>Buried in the supplementary methods, however, are some proprietary datasets, like the Baker Lab&#8217;s entire database of yeast display experiments (1.6M measurements against 41 proteins). Data collection has become an explicit engine for some of these companies, particularly Nabla.Bio. Others appear far more computational: as far as I know, neither Chai Discovery nor Latent Labs has opened a wet lab, though perhaps they are generating large amounts of data at cloud labs like <a href="https://adaptyvbio.com/">Adaptyv</a>. But the range of approaches is itself the interesting part. Given how rapidly and simultaneously these labs converged on similar technology (some data-heavy, some not) I am tempted to <a href="https://endpoints.news/isomorphic-labs-ceo-demis-hassabis-bets-on-biotechs-ai-future/">side with Demis Hassabis that data was not the problem</a>, at least to reach the current capabilities of these models.</p><p>There is another point that is very interesting in the paper by the Baker Lab. After running the whole design experiments and wet lab validation, they reran the predictions with the, back then, recently open-sourced AlphaFold 3, and observed that <strong>AF3 was an extremely powerful predictor of what designs were true binder versus which ones were not</strong> (AUROC 0.86). This is perhaps the most interesting point in this whole discussion. <a href="https://www.blopig.com/blog/2020/12/casp14-what-google-deepminds-alphafold-2-really-achieved-and-what-it-means-for-protein-folding-biology-and-bioinformatics/">When AlphaFold 2 was announced</a>, I claimed that immune proteins appeared to be one of the hardest tasks because <a href="https://www.blopig.com/blog/2021/07/alphafold-2-is-here-whats-behind-the-structure-prediction-miracle/#:~:text=Consider%20the%20following,get%20the%20idea.">coevolutionary information</a> was intrinsically not helpful due to the <a href="https://en.wikipedia.org/wiki/V(D)J_recombination">recombination</a> and <a href="https://en.wikipedia.org/wiki/Somatic_hypermutation">somatic hypermutation</a> that these proteins experience. Indeed, for a long time all the structure prediction models have significantly struggled with antibody prediction and (even more) with antibody-antigen prediction. However, it seems that <strong>something in the new AlphaFold 3 model has endowed the model with the ability to better understand antibody-antigen interactions</strong>. I wish I could pinpoint the reason, but this would require a large number of ablations that I would never have the compute to achieve. My intuition is that it has something to do with the model&#8217;s exposure to non-protein information, such as small molecules, forcing it to learn some of the physics that govern the interaction as well as the standard pattern matching.</p><p>If I had to bet, I would expect most of these companies to be exploiting AlphaFold 3-class models that are jointly modelling structure and sequence <em>i.e.</em> they can predict a structure from a sequence, but they can also generate a sequence from a structure, like ProteinMPNN, as well as generating both at the same time. These models will be based on transport-based generative approaches (like diffusion, or flow matching) and use similar loss functions to AlphaFold 3. I am sure there are a million small technical tricks, and certainly clever training recipes, but on the prior of having four independent groups arriving to this roughly at the same time, I would be happy to bet that this is a large portion of their success.  I would also be willing to bet that a small but elite team with 5-10M USD (a pittance with today&#8217;s AI seed rounds!) could likely get something on a similar level to what these guys have put out. After all, big portions of the architecture are public, and several open source models (see <a href="https://www.biorxiv.org/content/10.1101/2025.11.20.689494v1">BoltzGen</a> and <a href="https://www.biorxiv.org/content/10.1101/2025.09.19.677421v3.full.pdf">Germinal</a>) have achieved performance not far behind their well-funded competitors.</p><p>This will be important later, when we talk about the business models. But, for now, let&#8217;s return to the original question: are these models going to help us design new drugs?</p><h2>From antibody binder to antibody drug</h2><p>Antibody <em>binders</em> are very different from <em>antibody drugs</em>.</p><p>The non-life-sciences world tends to assume the hard problem in drug discovery is designing the drug. It is not. We have been taking antibodies to the clinic with enormous success: <strong>there are some 200 FDA-approved antibodies, with over 1,500 at some stage of the clinic</strong>, and with the recent success of antibody-drug conjugates (ADCs) and other novel modalities of antibodies, there will only be more and more in the next few years. <strong>Finding antibodies has never been the problem</strong>. You can search online for a few dozen companies that run hybridoma or phage display analysis at very competitive prices, with very reasonable success rates. The problem is converting these binders into drugs. </p><p>First of all, to be a successful drug, <strong>the binder has to become developable</strong>, which is a polite word for a long list of ways a molecule can betray you. It has to be <strong>stable enough</strong> to survive months in a vial without unfolding or losing potency, and to tolerate the freeze-thaw cycles and the heat excursions of a real supply chain. It has to stay a monomer: antibodies that quietly aggregate are both less effective and more dangerous, because aggregates are one of the things the immune system likes to notice. It has to <strong>express well</strong> in CHO cells at a titre that makes manufacturing economic, fold correctly, and purify cleanly; a beautiful binder that yields a hundred milligrams per litre instead of several grams is a binder that costs too much to make. It <strong>cannot be too hydrophobic</strong>, or it sticks to columns and to itself; it cannot be polyreactive, latching nonspecifically onto unrelated proteins, or you inherit toxicity and a half-life problem. Then there is <strong>immunogenicity</strong>: a fully human-looking sequence that nonetheless carries a T-cell epitope can provoke anti-drug antibodies that neutralise the therapy outright.</p><p>Beyond developability there are a million other problems related to how the antibody drugs perform inside a living body. For example, one is half-life and pharmacokinetics <em>i.e.</em> whether the molecule actually lingers in circulation long enough to dose monthly rather than daily, which is mostly a question of how it engages the <a href="https://en.wikipedia.org/wiki/Neonatal_fragment_crystallizable_receptor">FcRn recycling pathway</a> and how fast the kidneys clear it. Tissue and cross-reactivity are also problems, such as when the antibody binds to some other things that you could not even have imagined. What is more, <strong>the further you advance into the preclinical process, the properties that matter become slower and more expensive to validate</strong>.</p><p>The argument that has been floated around is that, if antibodies are <em>better from the get go</em>, these companies will accelerate the preclinical phase, leading to more shots on goal to reach the clinic. There are some early reasons to be interested in this. If you read the <a href="https://arxiv.org/pdf/2512.20263">Latent Labs whitepaper</a>, the majority of the antibodies they designed seem to have some excellent developability properties: the stability, monomericity and yield are comparable to therapeutic antibodies and frankly much better than most first-round designed antibodies; although the hydrophobicity and polyreactivity still need some work. Similar metrics are reported across other companies&#8217; papers.</p><p>I agree that these methods will sometimes speed up some parts of the preclinical process. However, <strong>the real problem is whether the antibody produces a real benefit in a real patient, which is a property not of the molecule but of the biology</strong>, and which no amount of binding affinity or easily assayable properties can buy you. Of the 88% of antibodies (roughly 7 in 8) that fail in the clinic, <strong>every single one of them passed the filters that these companies claim to be able to overcome, and even demonstrated efficacy in advanced animal models, yet failed for a completely different reason. </strong></p><p>One of the best examples here are the anti-amyloid antibodies developed in the never-ending quest for cures of Alzheimer&#8217;s disease. Take <a href="https://en.wikipedia.org/wiki/Solanezumab">solanezumab</a>: it was an extremely good antibody by all quantitative metrics, binding monomeric &#946;-amyloid at picomolar affinity, and sailing through any developability panel you could possibly imagine, yet <a href="https://www.nejm.org/doi/full/10.1056/NEJMoa1312889?__cf_chl_f_tk=0GQXYDD3XdM8I4sSGeHNATkvIhGDdZoUlAaNHoPxXMY-1783015089-1.0.1.1-7beo5hEZ7.9ckfHbYc7I3E3UzErN4Az8j6SSzY8i_T4">it failed its Phase III trials outright</a>.  In contrast, <a href="https://en.wikipedia.org/wiki/Aducanumab">aducanumab</a> bound the same target some ten thousand times more weakly (micromolar, which these models, and frankly most of the people I know, would discard), and yet it was the first in the class to <a href="https://pubmed.ncbi.nlm.nih.gov/35542991/">show a human signal</a>, because it was able to bind the aggregate rather than the monomer. But, and here is the catch, <strong>both antibodies <a href="https://www.alz.org/alzheimers-dementia/treatments/aducanumab">failed to produce a meaningful clinical effect</a></strong> because the biological hypothesis (that reducing amyloid deposits would halt disease progression) <a href="https://www.bmj.com/content/372/bmj.n156">doesn&#8217;t seem to hold</a>. And this isn't limited to diseases like Alzheimer's, where the translational difficulties are notorious; there are others, in antibodies (<a href="https://en.wikipedia.org/wiki/Afelimomab">afelimomab </a>in sepsis) and in drug discovery generally.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_QqS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_QqS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 424w, https://substackcdn.com/image/fetch/$s_!_QqS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 848w, https://substackcdn.com/image/fetch/$s_!_QqS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 1272w, https://substackcdn.com/image/fetch/$s_!_QqS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_QqS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png" width="606" height="281.35714285714283" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:676,&quot;width&quot;:1456,&quot;resizeWidth&quot;:606,&quot;bytes&quot;:1338188,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://couteiral.substack.com/i/202630550?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_QqS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 424w, https://substackcdn.com/image/fetch/$s_!_QqS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 848w, https://substackcdn.com/image/fetch/$s_!_QqS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 1272w, https://substackcdn.com/image/fetch/$s_!_QqS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F776a9d03-70c0-4260-97fc-db4ce23b2e9d_1491x692.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Infographic describing how different epitopes of the same protein can lead to significantly different biological effects. Infographic lightly edited from a GPT Image-2 generation, June 2026.</figcaption></figure></div><p>This does, however, point to a very attractive advantage of <em>de novo</em> design: <strong>the ability to control the <a href="https://en.wikipedia.org/wiki/Epitope">epitope</a> of the antibody.</strong> In short, since proteins are three-dimensional machines, <em>where</em> an antibody binds matters as much as <em>whether</em> it binds. Two antibodies with identical affinity for the same protein can have opposite biological effects: one lands next to the active site and shuts the protein down, whereas the other binds a harmless patch elsewhere and does nothing. For some targets this is extremely important. Take G protein-coupled receptors (GPCRs), the target of nearly a third of all approved drugs and a notoriously hard one for antibodies. Indeed, Nabla.Bio <a href="https://www.biorxiv.org/content/10.1101/2025.05.28.656709v1.full.pdf">showed their ability to design many antibodies against this target class</a>, including validated agonists and antagonists <em>i.e.</em> they didn&#8217;t just bind, they elicited a biological effect.</p><p>In summary, the hard truth in the drug discovery industry is that <strong>most drugs fail even in the best conditions</strong>. Every single drug that went into clinical trials worked on the animal models, was approved by the key opinion leaders, and ticked every box to convince the drug developer to invest tens of millions of dollars into clinical trials. As they say, if cancer in humans was like cancer in mice (or even in primates), we would have cured most of it already. The ultimate hard truth about drug discovery is that the bottleneck is human studies. An enormous graveyard of extremely promising drugs sits behind us, none of which went anywhere.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://couteiral.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">If you got this far, you enjoyed my essay at least a little. Subscribe (free) for more essays on frontier AI in the life sciences.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h2>Is this a business?</h2><p>We have concluded that there is a lot of good science in this two-year revolution, but that there is (or rather, I have) some skepticism about their ability to truly transform drug discovery. The next potential question is: do these companies have a business? </p><p>The four companies alluded to in the introduction are the near definition of <em><a href="https://centuryofbio.com/p/on-biotech-platform-strategy">platform</a></em><a href="https://centuryofbio.com/p/on-biotech-platform-strategy"> companies</a>: they are working on a technology engine that can generate <em>multiple</em> drugs. Because of the nature of the pharmaceutical industry, where all of the revenue comes from a few extremely profitable drug launches, these companies end up with two broad strategies: they either <em>make a drug </em>(what I will call the &#8220;biotech&#8221; model), or <em>help someone else make a drug</em> (what I will call the &#8220;outsourcing&#8221; model).</p><p>The outsourcing model is very attractive: <strong>you find a few Big Pharmas, and get them to sign you a </strong><em><strong>biobucks</strong></em><strong> deal to develop drugs for them</strong>. What this means is that the pharma company will pay an <em>upfront</em> amount for usage of the company; if you are a hot company, this upfront can range from the mid-double-digit millions to the low-triple-digit millions. But that is not all, the pharma will also give you <em>milestone</em> payments, which means that everytime that the molecule goes through some value inflection point (say, IND approval, or first injection in a Phase I trial), the Pharma will wire an even larger amount of money. You will also have access to <em>royalties</em>, which means a share of the revenue produced by the sales of the molecule if/when it reaches approval. These humongous amounts of money will often be reported, in a headline, as an even more humongous number that contains every possible payment. <a href="https://www.statnews.com/2016/11/28/biotech-biobuck-deals/">Whether these humongous numbers will actually materialise</a> is an entirely different story, but even the cash upfront can be an outrageous amount of money for an otherwise cash-starved biotech company.</p><p>Here is the deal: <strong>if any single company had created this technology, they would be the sweetheart of the pharmaceutical industry</strong>. When four companies reach the same milestones within months, then their ability to capture value is reduced because the technology nears <a href="https://en.wikipedia.org/wiki/Commoditization">commoditisation</a>. A pharma company shopping for antibody frontier models now has four doors to knock on instead of one, and that competition compresses margins. I should caveat, however,  that Chai Discovery has performed particularly well here, because their announced partnerships with <a href="https://www.businesswire.com/news/home/20260108131261/en/Chai-Discovery-Announces-Collaboration-with-Eli-Lilly-and-Company-to-Accelerate-Biologics-Discovery">Eli Lilly</a> and <a href="https://www.biospace.com/press-releases/chai-discovery-announces-license-agreement-with-pfizer-to-accelerate-drug-discovery-with-ai">Pfizer</a> report <em>recurrent revenue</em> from access to this. While they don&#8217;t get to report the <em>very humongous</em> number of the total deal value, the disclosed Eli Lilly fee is reportedly in the eight-figure range. It remains to see, however, whether they will be able to maintain these deals in the light of increased competition.</p><p>Commoditisation is a challenge to any industry, but it is even worse in the software world because <strong>open source is racing behind</strong>. Within months of the commercial releases, the open source <a href="https://www.biorxiv.org/content/10.1101/2025.11.20.689494v1">BoltzGen</a> pipeline was published with the ability to generate nanobodies (<em>note: while editing this article, the <a href="https://boltz.bio/">Boltz team</a> announced <a href="https://boltz.bio/api">BoltzProteo</a>, which is an upgraded version of this pipeline with a better model; their whole thesis is to be the default model provider in the commoditisation era, and they are delivering like they mean it</em>). Okay, you may argue that they are using <a href="https://www.biorxiv.org/content/10.1101/2025.06.14.659707v1">Boltz-2</a>, which was trained on Recursion&#8217;s <a href="https://top500.org/system/179939/">BioHive supercomputer</a>, so they are not really your garden variety academic group. But you can also get <a href="https://www.biorxiv.org/content/10.1101/2025.09.19.677421v3.full.pdf">Germinal</a>, a model trained by academics at Stanford and the Arc Institute. While these guys have access to quite a lot of compute for an academic group, their method is really simple: they backpropagate through AlphaFold Multimer with a mix of gradients from AF-M itself, and an antibody language model (IgLM); then they map the structure to a sequence using AbMPNN and verify with AlphaFold 3. This, plus a few tweaks from the team (and these are all really simple things, like constraining the generation to avoid certain secondary structures) is enough to deliver double-digit success rates in the nanomolar range.</p><p>If the outsourcing, &#8220;B2B SaaS&#8221; play is not an attractive one, the alternative strategy is for these companies to play the biotech model. Here, they would <strong>pick some valuable targets, develop some antibodies against them, push them through the drug discovery pipeline, and enter the clinic as soon as possible</strong>. This very much seems like what Xaira (and, I would bet, at least some teams within Isomorphic Labs) are doing in the background; of course, we do not hear that much about them because, <a href="https://www.fiercebiotech.com/biotech/new-ai-drug-discovery-powerhouse-xaira-rises-1b-funding">having raised billions of dollars</a>, they don&#8217;t really have to worry about making announcements for investors and can focus on the long-term delivery.</p><p>The problem is that <strong>asset development is </strong><em><strong>fundamentally</strong></em><strong> a terrible business model</strong>. After all, a drug that hits the clinic has a 88% probability of <em>failure</em> (that&#8217;s because it&#8217;s an antibody, a small molecule would have a whopping 92%). That is, I will not get tired of saying it, a byproduct of years of work that has shown success in multiple animal models and any possible experimental test before taking it into humans. Despite the absurdly large amounts of money that some drugs generate, <em><a href="https://www.deloitte.com/ch/en/Industries/life-sciences-health-care/analysis/measuring-the-return-from-pharmaceutical-innovation.html">the pipelines of most pharmaceutical companies return less than the cost of capital</a></em>. And yet they have dozens of drugs entering the clinic every year, and can diversify the risk, so what hope has a small biotech? When the attrition-adjusted cost of taking a drug to approval is about $2 billion dollars, even extremely well-funded companies like Xaira ($1B seed) and Chai ($225M raised, rumoured to be raising more) may be short-changed.</p><p>There is only one way to win, really. They need to <strong>pick one differentiated clinical candidate, prove it, and then let it fund the platform</strong>. But for the reasons that I have mentioned before, they can&#8217;t afford to go against a random target. They need to <strong>pick some target that is sufficiently biologically well-validated that they minimise the risk of clinical trials going awry</strong>, or otherwise they face the same inefficiencies as most pharmaceutical companies. At the same time, they cannot afford competition, so they need to focus on something that no other competitor can do: something that their platform unlocks. The only one of the companies that seems to be doing this is Nabla, with their published work <a href="https://www.biorxiv.org/content/10.1101/2025.05.28.656709v1">developing antibodies for G-protein coupled receptors</a> or the previously discussed <a href="https://nabla-public.s3.us-east-1.amazonaws.com/2026_Nabla_JAM2_multispecific_pMHC.pdf">peptide-MHC antibody mimics for KRAS</a>.</p><p>There can be a lot of discussion of what constitutes a great target for one of these companies, but notice that the two requirements pull in opposite directions. The biology has to be well-trodden enough that the target is not the gamble, but the particular way the molecule is looking at it. Likewise, the molecule has to be one that only <em>de novo</em> design could have produced, which by definition pushes you toward the targets classical discovery never cracked. The ideal candidate is therefore a strange chimera: a target safe enough that a small biotech can survive its clinical trial, and hard enough that nobody could have made the antibody any other way. The success is clear: a success like Revolution Medicines&#8217; <a href="https://en.wikipedia.org/wiki/Daraxonrasib">daraxonrasib</a> <a href="https://www.nejm.org/doi/full/10.1056/NEJMoa2505783">Phase III results</a>, <strong>finding a novel way to succeed with KRAS, a drug target known for fifty years that however remained uncracked</strong>. Unfortunately, the intersection is narrow, and everyone capable of fishing in it is fishing in the same pond.</p><p>So here is the open question for these antibody drug discovery companies: can they deliver a differentiated clinical candidate? I will be very interested to see what they come up with.</p><h2>The data play</h2><p>There is a counterargument I owe an honest hearing, because it is the best one the believers have: the value of data in delivering an exponential improvement in capabilities.</p><p><a href="https://centuryofbio.com/p/on-biotech-platform-strategy">Platform biotechs</a> are not new; in fact, some of the great pharmaceutical successes of the last fifty years were built this way (<a href="https://en.wikipedia.org/wiki/Amgen">Amgen</a> on recombinant proteins, <a href="https://en.wikipedia.org/wiki/Gilead_Sciences">Gilead</a> on antivirals, <a href="https://en.wikipedia.org/wiki/Regeneron_Pharmaceuticals">Regeneron</a> on humanised antibodies, etc). What is supposed to make this wave different is the <strong>data flywheel</strong>. In a traditional biotech, a failed project is killed and you learn almost nothing, but <strong>in an AI biotech, a failed project becomes valuable training data</strong>. Every campaign, win or lose, makes the model better. The moat compounds, and the frontier pulls away from the open-source competition trailing it by only a few months. If you believe that, then the commoditisation worry I just raised should just be temporary, and the company that generates the most data wins.</p><p>The main argument against the data flywheel in the antibody field is that the data that is most valuable is also the most expensive to get. At every step of the preclinical journey, the problem increases in value and cost: finding binders is relatively easy, optimising them for developability is a bit harder, getting the pharmacokinetics is harder still, animal experiments are more expensive the closer you get to human, and testing biological hypotheses in a human trial is a multi-million dollar experiment that so far has no real alternative. If the data you need to power the flywheel is at the top of the value chain, as seems the case, then you are in trouble.</p><p>The question to always keep in mind is: <strong>does the marginal data get cheaper than the marginal value it unlocks?</strong> If the data required to keep improving the models grows faster than the improvement it buys, then the flywheel does not run away from the competition. The equilibrium point is somewhere around where the cost of the next accuracy increment exceeds the industry&#8217;s willingness to pay. </p><p>In this model, the only way to make the flywheel work is to build an advantage at the very early parts of the discovery. Here is where I think Nabla.Bio&#8217;s strategy is really clever: their focus is on modalities that have so far been inaccessible but the biological interest is clear, such as the GPCR modulators and the pMHC mimetics. Their <a href="https://www.biorxiv.org/content/10.1101/2025.01.21.633066v1.full.pdf">JAM platform</a> combines <em>de novo</em> design with what amounts to neighbourhood expansion: generating large amounts of data around a promising binder with a robotised pipeline, essentially. Because their data is closest to physics binding, it is unsurprising that their platform gets intricate control on the binding mode of the antibodies and enables some of their breakthroughs.</p><p>The LLM world offers a clear parallel. In 2026, it is no longer controversial to say that <a href="https://arxiv.org/abs/2510.13786">pure scaling by training on data is largely considered stalled at most frontier labs</a>. The scaling these days has taken three forms: reinforcement learning at scale, generally with <a href="https://arxiv.org/abs/2501.12948">verifiable rewards</a>; post-training on specific tasks (such as the many white-collar datasets collected by companies like Scale AI and Mercor); and test-time scaling, which is just letting the models think for longer. These do not have a clear parallel in biology. While there are some reports of test-time scaling in biology, most of them can be summarised in running an inference process for longer (e.g. more diffusion samples or more recycling steps). Reinforcement learning is far harder in biology, because we can't cheaply (or at all) simulate outcomes, so the reward signal has to come from the wet lab, where it's slow and expensive. Therefore, if you agree as we have posed that more pretraining will give you diminishing returns (as the <a href="https://en.wikipedia.org/wiki/Neural_scaling_law">scaling law</a> foretells), the only play left is post-training: focusing on specific problems and moving forward. But post-training is a lot about becoming very good on specific assets, and less about getting an overly powerful model that takes forward the data flywheel.</p><p>I believe in the value of the data flywheel in AI for life sciences with all my heart. So much so that I have dedicated my career to it. However, unless the strategy focuses at the preclinical end, where the data is cheap and the models can accrue value, I do not think it makes sense.</p><h2>Conclusions</h2><p>The developments in <em>de novo</em> antibody discovery that we saw in 2025 and 2026 are one of the most compelling advances I have seen in the field. They remind me a lot of AlphaFold 2 in 2021; not on the magnitude, but on the amount of proof delivered. Having spent a lot of time over the last months thinking about it, I have concluded that there is no catch&#8230; in the science.</p><p>The catch is in the business model. There are many questions about how well these model companies will perform as a business. In general, I am skeptical that the platform-as-a-service model will work here. I think we are likely to see a scenario mirroring LLM commoditisation, where the open source model tracks closely behind the frontier. This time, unlike natural language models, there will be no clear reinforcement learning loops or additional training techniques; and, if the data required to increase the accuracy of the models needs to grow exponentially while the cost of data acquisition stays constant, models will halt as soon as the cost of developing them surpasses what the industry is willing to pay for them.</p><p>The open question for the field is whether these companies will be able to play the platform biotech game. In almost all cases, this will require them to commit to at least one drug candidate where their technology truly differentiates them. If this candidate(s) fare well, the revenue they generate may power the rest of the technology. Otherwise, we may see another wave of AI for drug discovery companies struggle.</p><p>The revolution happened; will the business sustain it?</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://couteiral.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free if you would like to read more essays on frontier AI and the life sciences.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>