NEWS
Flow Matching Turns Biology From Snapshots Into Paths
A Missouri-led review maps flow matching for biology, but the gain is trajectory data and a virtual cell that still sits about a decade out.
A nine-author team led from Lawrence Berkeley National Laboratory and the University of Missouri put flow matching on the biology map in Nature Machine Intelligence. The review published on 23 April 2026 runs in volume 8, pages 517 to 534, and had drawn 7,963 accesses, 5 citations and a 20 Altmetric score by early September 2026.
Jianlin “Jack” Cheng, a Curators’ Distinguished Professor in bioinformatics at Missouri, sold the work as a way to watch life change rather than freeze it. The paper underneath that pitch is a review of other labs’ models, and it leaves a harder bill: if you want paths, you need measurements at more than one state.
The April Review Maps a Crowded Field
Nature Machine Intelligence listed the article as a Review, not a new model. First author Alex Morehead is at Lawrence Berkeley National Laboratory. Cheng is the Missouri corresponding voice, with co-authors at the Broad Institute’s Eric and Wendy Schmidt Center, Mila and Université de Montréal, and the University of California, Berkeley.
THE PAPER AT A GLANCE
- The venue: Nature Machine Intelligence, volume 8, pages 517 to 534, dated 23 April 2026.
- The authors: Nine names, with Morehead first and Cheng the Missouri senior author.
- The reach: 7,963 accesses, 5 citations and a 20 Altmetric score on the journal page.
- The scope: Small molecules, proteins, DNA and RNA, their contacts, plus single-cell and imaging models aimed at a virtual cell.
A University of Illinois survey posted on 23 July 2025, and revised on 8 March 2026, had already claimed a first wide look at flow matching in the life sciences. The Missouri-led paper is the version that landed in a Nature journal, with a public reading list attached.
Flow Matching Learns Paths Between Biological States
Cheng put the method in plain language for the university announcement. Flow matching, he said, helps computers learn how biology changes from one state to another, from protein folding to cell development and cancer progression.
Flow matching is becoming a unifying framework for generative AI in biology. It has the potential to fundamentally change how we model, study and understand living systems.
Jianlin Cheng, Curators’ Distinguished Professor in Bioinformatics, University of Missouri
The review’s own abstract is more precise. Many biology problems are mappings from one state of a system to another, or searches for new points inside a constrained space. Hand-built maps, such as taking a diseased cell back toward a healthy one, eat expertise. Flow matching learns a vector field that carries samples between two high-dimensional distributions, including pairs that are not Gaussian noise and data.
That is why Cheng can talk about folds, cells and tissues in one breath. The same training idea can move a molecule toward a binder, a cell population toward a drug response, or an image of a cell toward a later morphology, if the datasets exist.
How Flow Matching Differs From Diffusion
Yaron Lipman and colleagues posted simulation-free training for generative flows on 6 October 2022 and presented it at ICLR in May 2023. The trick is to regress a vector field along a chosen probability path, then sample with an ordinary differential equation solver instead of a long denoising chain.
Meta later said the same family had replaced classical diffusion in several of its own generators and published a flow matching guide and code on 10 December 2024. Biologists inherit that shortcut: the source distribution can be another biological state, not only noise, and linear or optimal-transport paths can be shorter than diffusion curves.
SNAPSHOT MODELS VERSUS FLOW MATCHING
| Approach | Typical starting point | What it learns | Biology example |
|---|---|---|---|
| Snapshot predictor | One sequence or structure | A single likely state | AlphaFold 3 |
| Diffusion generator | Gaussian noise | A path that ends at data | RFdiffusion |
| Flow matching | Any source distribution | A vector field between two states | FoldFlow, CellFlux |
Cheng also said computers can see connections across enormous amounts of data that humans cannot, and that this helps researchers move faster and ask better questions. Speed is real when the path is short. It does not invent the two ends of the path.
Paired Time Points Are the Hidden Cost
The review’s cellular section is blunt about homework. To model a cell’s dynamics in silico, the first step is to collect a sufficient quantity of paired measurements of that cell’s state at, at least, two different time points. Then a network can encode those states and learn the flow between them.
Most popular stores of biology are still snapshots. The Protein Data Bank holds structures, not movies. Standard single-cell RNA sequencing destroys the cell, so a lab gets a cloud of cells at time A and another cloud at time B, with no line joining one cell to its later self. Flow matching can still push one cloud toward the other, which is why unpaired optimal-transport variants exist. A true trajectory, the kind Cheng invokes for development and cancer, wants time and perturbation, not one freeze-frame.
WHAT THE METHOD ASKS OF A LAB
- Two or more states: Healthy and diseased, unperturbed and treated, unfolded and folded, measured on the same assay.
- Population depth: Enough cells or molecules in each state for a distribution, not a handful of heroes.
- Perturbation labels: Drugs, gene knockouts or morphogens, so the flow can be conditioned instead of guessed.
- A scale match: Atomic coordinates for folds, embeddings for phenotypes, pixels for morphology, kept consistent from source to target.
Fabian Theis, a computational biologist at Helmholtz Munich, has made the same point from the single-cell side: the live problem is dynamics beyond snapshots, and simulation-free flow matching is how groups now try to model whole populations without integrating a neural ODE at every training step. The algorithm got cheaper. The wet-lab calendar did not.
Missouri Speaks for a Multi-Lab Cast
Missouri’s announcement led with Cheng and the College of Engineering. The author list is wider, and it explains why the paper can talk like a methods tutorial rather than a campus exclusive.
Morehead and Aditi Krishnapriyan sit on the Berkeley side of the collaboration, Krishnapriyan in computer science and chemical engineering. Alexander Tong, at Mila and Université de Montréal, is a builder of conditional flow matching, the training objective the review flags as a practical way to fit these models. Lazar Atanackovic is listed with the Broad Institute’s Eric and Wendy Schmidt Center. Akshata Hegde, Yanli Wang, Frimpong Boadu and Joel Selvaraj are the Missouri group around Cheng and NextGen Precision Health.
HOW FLOW MATCHING REACHED BIOLOGY
- 6 October 2022: Lipman and colleagues post flow matching as simulation-free training for continuous normalizing flows.
- May 2023: The work appears at ICLR, and the first biology-scale flow models follow the same year.
- 2024: The review counts 8 notable molecular flow-matching methods in that year alone, with protein backbones, ligands and sequences all in play.
- 10 December 2024: Meta releases its guide and PyTorch package; a Cell perspective sets out an AI virtual cell as a shared goal.
- Early 2025: The review notes 3 foundational single-cell flow methods among a wider wave of morphology and spatial models.
- 23 April 2026: Nature Machine Intelligence publishes the Morehead-to-Cheng review and points readers to code.
Cheng’s quotes do the public work. The engineering work of pairing distributions sits with Tong’s line of code and with the Berkeley and Broad co-authors who already treat flow matching as infrastructure.
The Models That Left the Review Behind
By the time the journal issue landed, flow matching was already a production choice in protein and molecule papers, not a proposal waiting on a roadmap. The authors keep an open catalog of flow-matching methods on GitHub, and the application table there is the honest scoreboard.
OPEN FLOW MODELS THE REVIEW POINTS TO
| Model | Job | Benchmark figure on the authors’ list |
|---|---|---|
| EquiFM (Song et al., 2023) | Small-molecule generation | 98.9% validity on GEOM-Drugs, +7% |
| SemlaFlow (Irwin et al., 2024) | Small-molecule generation | 93.9% validity on GEOM-Drugs |
| FrameFlow (Yim et al., 2023) | Protein backbone generation | 0.81 designability, +93% |
| FoldFlow (Bose and Huguet, 2024) | Protein backbone generation | 0.82 designability, +34% |
CellFlux, cited in the review, treats control and perturbed cell images as two distributions and learns the flow between them. AlphaFlow fine-tunes structure predictors under a flow objective so a protein is an ensemble, not a single pose. In early September 2026, new preprints were still arriving on Dirichlet flows for inverse folding, flow models for RNA-protein ensembles, and water placement around crystal structures. The pattern is the same: pick two states, learn the path, sample.
Kevin K. Yang at Microsoft Research described a related next step as expanding flow maps, models that insert new coordinates or tokens while they denoise, so the output can be larger than the input. That is a builder’s concern, not a press-release horizon, and it is where the field was already arguing while the review was being read.
What a Virtual Cell Still Cannot Do
The long bet in both the paper and Cheng’s comments is an AI-powered virtual cell, a digital stand-in that lets a scientist try an idea on a computer before a bench run. Cheng said that, over time, this could reduce reliance on animal and human studies and accelerate progress toward more personalized medicine. The review’s preprint language is slower: with data-driven mappings, an AI-based virtual cell may be within reach in the coming decade.
Charlotte Bunne and colleagues set priorities for an AI virtual cell in Cell on 12 December 2024, volume 187, pages 7045 to 7063. They described a multi-scale, multi-modal network that could represent molecules, cells and tissues across states. On 6 November 2025, Biohub named a unified AI model of the cell as a grand challenge and said it would expand compute toward 10,000 GPUs by 2028. The Missouri review is academic scaffolding for that race, not the model itself.
WHAT WE KNOW
- The method: Flow matching can, in principle, carry a molecular distribution, a cell-state embedding or a microscopy image from one condition to another.
- The code: Meta’s package and Tong’s conditional-flow-matching library are public, and the review points to both.
- The ambition: Authors from Missouri to Montreal treat a virtual cell as the stack of those mappings, not a single network trained next month.
WHAT IS UNCONFIRMED
- Animal and human studies: Cheng’s hope that virtual cells cut in-vivo work is a forecast, not a result in this paper.
- Personalized treatment: No trial, dose or patient cohort is attached to the review.
- One unifying model: The GitHub list is still a pile of specialist generators, not a single cell that folds, signals and divides.
A virtual cell that can be trusted will be a data product before it is a software product. The review is useful because it says that out loud in a methods journal, even when the campus announcement sounds like a breakthrough in hand.
Frequently Asked Questions
Who Introduced Flow Matching?
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel and Matt Le presented the method at ICLR in Kigali, Rwanda, held 1 to 5 May 2023, after posting the paper on 6 October 2022. Lipman later co-wrote Meta’s public guide while based at Meta FAIR in Tel Aviv.
Is the Missouri Paper a New Biomedical Model?
No. Nature Machine Intelligence classified it as a Review Article that surveys theory and applications and points to open-source implementations. It does not report a newly trained protein, drug or virtual-cell model from the Missouri group.
What Does an AI Virtual Cell Mean in This Literature?
Bunne and colleagues defined an AI virtual cell as a multi-scale, multi-modal large neural network that can represent and simulate molecules, cells and tissues across diverse states, and they framed it as a community build rather than one lab’s checkpoint.
Where Is the Conditional Flow Matching Code the Review Cites?
Alexander Tong’s group maintains a PyTorch library at github.com/atong01/conditional-flow-matching, which the journal’s code-availability note lists beside Meta’s package as a starting point for training CFM models.
The reading list is public and the generators are already running on folds, ligands and cell images. The missing piece is still the second measurement, the one that turns a snapshot archive into a path a virtual cell would have to get right.
-
ENTERTAINMENT3 weeks agoBravo Cuts Nathan Gallagher but Still Airs Below Deck
-
NEWS3 weeks agoGoogle AI Mode Adds Paginated Follow-Ups With Skip
-
NEWS4 weeks agoApple Uses a Returned MacBook to Press OpenAI Hardware
-
GAMING3 weeks agoDawnwalker Hits 1 Million as Players Stretch Its Clock
-
NEWS3 weeks agoAustralia’s Teen Social Media Ban Still Lets Most Kids In
-
NEWS3 weeks agoCompliance AI Adds Work for More Teams Than It Saves
-
NEWS3 weeks agoHarvard Study Cuts Wasted Cache Promotions 20 to 60 Percent
-
NEWS3 days agoGoldman’s $1.2 Trillion AI Capex Needs $300 Billion in Revenue
