15 Challenges for Generative AI in Cell Biology: What the Cell Paper Says
On August 17, 2026, the Perspective “Fifteen challenges for generative AI applications to cell biology” was published in Cell. Instead of declaring another particularly large model a breakthrough, the author team poses a more fundamental question: What tasks would generative AI actually need to solve before we can speak of reliable prediction and targeted control of cells?
The answer is 15 challenges, ranging from molecular interactions, cell states, and mechanisms of action to biomarkers, organism responses, and clinical studies. The paper is therefore less a ranking of current models than a proposal for better scientific targets: success should be measured by whether a model provides useful, experimentally verifiable predictions in unknown biological situations.
In Brief
- The paper formulates 15 concrete challenges for generative AI in cell biology and assigns them to four levels: molecular interactions, molecular function, cell/system function, and translation.
- The core point is not whether a model convincingly reconstructs biological data, but whether it makes correct and practically relevant predictions outside of known training situations.
- Tasks involving causality are particularly demanding: Which intervention changes a cell state, why does a drug work, how do cells communicate with each other, or what result will a study not yet conducted have?
- Current foundation models are therefore not automatically “virtual cells”. For example, a benchmark published in 2025 showed that several deep learning models could not consistently outperform simple baselines in predicting genetic perturbations.
- The 15 challenges connect AI development with experimental biology: good benchmarks should be designed so that progress can be translated into measurable biological insights.
What the Cell Paper Actually Aims to Achieve
The work by Léo Dupire and 14 other authors – including Theofanis Karaletsos, Shana O. Kelley, Emma Lundberg, Jian Ma, Stephen R. Quake, and Andrea Califano – is a Perspective, meaning a scientific orientation and discussion contribution. It does not present a single new cell model with a new best performance. Instead, it tries to align the field towards tasks whose solutions are biologically meaningful.
This is important because “good performance” in computational biology strongly depends on what exactly is being measured. A model may classify known cell types very well or plausibly fill in missing values, and still fail to predict the response to a new genetic perturbation, an unknown tissue, or a new drug. For truly predictive cell biology, this generalization is crucial.

Source: pexels.com
The benchmark for biologically useful AI is not just a good benchmark score. What matters is whether predictions hold up in experiments with new conditions and help researchers with their next decision.
The authors therefore structure the 15 tasks along an increasing biological scale. Below are the original English titles of the challenge headings. The German explanations are an editorial classification of what type of prediction or design task is to be tested in each case; they are not a literal translation of the full paper text.
The 15 Challenges at a Glance
Level 1: Molecular Interactions
The first level concerns the elementary relationships from which cell behavior arises. A model must not only recognize which molecules occur together but also be able to reliably handle regulatory context, chromatin state, and communication between cells.
| No. | Original Title | What it practically entails |
|---|---|---|
| 1 | Regulatory and signaling interactions | To predict regulatory and signaling relationships in such a way that interpretable responses to an intervention or stimulus can be derived from them. |
| 2 | Epigenetic interactions | To capture how chromatin and other epigenetic states contextually alter gene activity and cell response. |
| 3 | Cell-cell interactions | To predict the communication and mutual influence of different cells, rather than modeling each cell in isolation. |

Source: commons.wikimedia.org / National Institutes of Health
Cellular behavior arises from structures and interactions at multiple levels. The NIH image shows HeLa cells with actin in red, microtubules in cyan, and nuclei in blue – a vivid example of the simultaneously present molecular and cellular organization.
Level 2: Molecular Function
The second level shifts the question from observed relationships to function. Here, it becomes interesting whether AI can infer biological effects from sequences and molecular components, or even design new functional mechanisms.
| No. | Original Title | What it practically entails |
|---|---|---|
| 4 | Synthetic mechanisms | To design new, engineered molecular mechanisms with a desired behavior and reliably predict their function. |
| 5 | Genome to biochemical function | To move from genomic information to specific biochemical capabilities and processes – not just to statistical sequence similarities. |
| 6 | Drug mechanism of action | To predict the target structures, signaling pathways, and contextual conditions through which a drug actually acts. |
Level 3: Cell and System Function
At this level, molecular prediction becomes a systems task. A cell has many possible states, responds to its environment, and can be reprogrammed through interventions. This is precisely where it becomes apparent why a “language model for genes” is not automatically a complete virtual cell.
| No. | Original Title | What it practically entails |
|---|---|---|
| 7 | Genome to phenotype | To reliably infer phenotypic properties and behavior from a genome, even when combinations of variants or conditions are novel. |
| 8 | Cell state reprogramming | To find interventions that specifically move a cell from an initial state to a desired target state. |
| 9 | Synthetic circuit design | To design robust artificial genetic circuits whose behavior in real cells matches specifications. |
| 10 | Systems-level mechanisms | To capture system-level mechanisms where cell state, environment, and multiple coupled processes jointly determine the observed function. |

Source: pexels.com
Many of the proposed tasks require a closed loop of model and experiment: formulate predictions, test in controlled cell culture, measure deviations, and derive the next model generation from them.
Level 4: Translation
The fourth level is the toughest real-world test. Now, molecular and cellular predictions are to lead to questions relevant to drug development and medicine. This increases not only biological complexity but also the importance of data quality, patient heterogeneity, and prospective validation.
| No. | Original Title | What it practically entails |
|---|---|---|
| 11 | Complex biomarker identification | To find multidimensional biomarkers that robustly indicate a clinically relevant state or response and prove effective outside the discovery dataset. |
| 12 | Drug toxicity | To predict toxic effects of a drug as early as possible and in relevant biological contexts. |
| 13 | Drug efficacy | To predict efficacy not only on average, but for different biological states and potentially different patient groups. |
| 14 | Organismal responses | To generalize from individual cells and tissues to the responses of an entire organism where immune, metabolic, and organ systems are coupled. |
| 15 | Clinical trial outcomes | To predict how a clinical trial will turn out under real-world conditions – a task that goes far beyond reconstructing existing omics data. |

Source: pexels.com
The final challenges extend into translation. The closer a prediction gets to drug safety, efficacy, or clinical trials, the higher the requirements for independent data, reproducibility, and prospective confirmation.
Why Current Benchmarks for “Virtual Cells” Are Insufficient
The 15 challenges hit a sore spot in current AI biology: a benchmark can be technically sound and still create a false impression of progress. If training and test data are very similar, a model can interpolate patterns without adequately capturing the underlying mechanism. Therefore, for real predictive power, test cases are important that deliberately deviate from the known distribution – for example, new perturbations, different cell types, different tissues, or combinations that the model did not see during training.
The relevance of this objection was shown in 2025 by an open-access paper in Nature Methods. Constantin Ahlmann-Eltze, Wolfgang Huber, and Simon Anders compared five single-cell foundation models and two other deep learning approaches with deliberately simple baselines in predicting transcriptomic changes after genetic perturbations. In the tasks examined, none of the complex models could consistently outperform the simple baselines. The authors themselves emphasize that this does not mean that deep learning is fundamentally unsuitable for single-cell biology. Rather, the finding shows how important strong baselines and demanding tests are for statements about generalization.
The Cell paper addresses exactly this point: a future model should not be convincing by realistically replicating known data, but by its predictions improving the next experimental decision. This shifts the guiding question from “How similar does the output look?” to “Which new biological statement can we test with a reasonable error rate?”
Three Structural Problems Behind the 15 Challenges
1. Biological data is abundant, but not automatically the right training data
Single-cell atlases today contain enormous amounts of measurements. Nevertheless, for many predictive tasks, precisely the data that allows for causal statements are missing: controlled perturbations, time series, dose-response relationships, spatial context, multi-omics measurements, and a sufficient number of independent biological systems. One billion observed cells are less informative for some questions than a smaller, cleanly designed dataset with targeted interventions.
The Chan Zuckerberg Initiative and Biohub are therefore pursuing, parallel to the creation of large cell datasets, the explicit goal of an AI-based virtual cell that is intended to predict and explain cell behavior. The perspective "How to build the virtual cell with artificial intelligence" published in Cell in 2024 also describes the virtual cell as a multi-scale, multi-modal model for molecules, cells, and tissues. The 15 Challenges from 2026 can be read as more concrete benchmarks against which this ambitious goal can be measured.
2. Cell biology is combinatorial and context-dependent
Language has rules and recurring patterns; cells do too – but their state space is simultaneously shaped by gene regulation, protein interactions, metabolism, epigenetics, spatial neighborhood, developmental history, and environment. Even small changes can alter their effect depending on the cell type or context. A model that correctly knows a single component is therefore far from being able to correctly predict the consequence of a combination of multiple interventions.
This argues for models that do not treat biological prior knowledge as an afterthought, but consider it in architecture, training, and evaluation: for example, known interaction networks, physical constraints, reaction pathways, or explicit perturbation data. Larger parameter counts can be helpful but do not automatically replace this structure.
3. Correlation is not enough for many challenges
Cell state reprogramming, mechanisms of action, or drug effects involve interventions: What happens if we change X? This is precisely the question that separates correlation from causation. A model can perfectly associate a marker with a disease state without knowing whether the marker is the cause, the consequence, or just an accompanying symptom. This distinction is immediately relevant for Challenges 8, 11, 12, and 13.
A vivid example of the desired validation cycle is theZerlo report on the DeepMind/Yale cancer hypothesis confirmed in the lab: There, a model led to a concrete, context-dependent hypothesis that was subsequently tested in human cell models. Such a hit does not yet prove a general "virtual cell," but it shows why prospective experiments are significantly more informative than purely retrospective benchmarks.
What a truly useful "AI Virtual Cell" would need to be able to do
The term AI Virtual Cell is often used for foundation models that learn representations from large single-cell datasets. This is an important building block, but the 15 Challenges suggest a higher bar. A useful virtual cell would not only need to encode the current state but also be able to simulate interventions and their consequences.
- Generalize: Predictions must remain robust for unseen cell types, perturbations, or environments.
- Connect multiple scales: Molecular interactions must be linked to cell state, tissue, and ultimately organismal responses.
- Indicate uncertainty: A model must indicate when data is lacking for a robust prediction, rather than making every output appear equally certain.
- Be experimentally falsifiable: Predictions should be specific enough to be refuted or confirmed in the lab.
- Be mechanistically useful: For research and drug development, it is not only important that an effect is predicted, but also, as far as possible, through which mechanisms and under which conditions it arises.
This also makes it clear why Challenge 15 – clinical trial outcomes – is at the end of the scale. Between a gene expression profile and the outcome of a trial lie numerous levels: dosage, pharmacokinetics, immune responses, comorbidities, patient selection, study protocol, and statistical endpoints. A model that reliably connects all these factors would be qualitatively different from a cell type classification system.
What AI and biology teams can learn from the paper
A practical development strategy can be derived from the 15 Challenges. The following is an editorial synthesis from the paper and the current benchmark debate, not an additional list from the authors:
- First, define the biological decision. Before choosing a model metric, it should be clear which experimental or translational decision is intended to be improved by the prediction.
- Take simple baselines seriously. A complex foundation model should compete against naive, linear, and domain-specific baselines. Otherwise, it remains unclear whether the additional complexity provides genuine information gain.
- Plan test data for generalization. Meaningful splits hold back precisely the situations that are intended to be predicted later: new perturbations, new cell contexts, or entire independent datasets.
- Separate retrospective and prospective evaluation. Retrospective benchmarks help compare methods. Prospective experiments, on the other hand, test whether the model can actually predict new, unknown biology.
- Integrate biological prior knowledge and uncertainty. Networks, mechanisms, and physical constraints can structure the search space; uncertainty estimates help avoid risky overinterpretations.
- Publish negative results. If a large model does not beat simple baselines, it is scientifically valuable. Such results prevent the field from confusing computational effort with biological progress.
Why the paper is relevant beyond cell biology
The perspective is an example of a general problem in generative AI: the more impressive a model's output, the easier it is to confuse plausibility with correctness. In cell biology, this difference is particularly striking because a plausible prediction can fail in an experiment – and because false assumptions later cost time, animals, samples, or clinical resources.
The 15 Challenges therefore provide a useful counterpoint to pure scaling thinking. They demand tasks whose solution becomes measurably more difficult and scientifically more valuable the further one moves from molecular reconstruction towards causal cell control and clinical translation. The paper deliberately leaves open whether current transformers, graph-based models, hybrid mechanistic systems, or new architectures will best solve these tasks.
FAQ
What does "Fifteen challenges for generative AI applications to cell biology" mean?
This is the title of a perspective published in Cell in 2026. The author team formulates 15 scientific challenges by which generative AI for cell biology can be meaningfully tested – from regulatory interactions to the prediction of clinical trial outcomes.
What four levels do the 15 Challenges cover?
The challenges range from molecular interactions to molecular function and cell/system function to translation. This gradually raises the bar from local biological relationships to predictions about drugs, organisms, and clinical trials.
Does the Cell paper say that a virtual cell already exists?
No. The paper is not a success story about a virtual cell that has been fully solved. It describes tasks that future generative models would need to accomplish for their predictions to be more biologically and translationally convincing.
Why are good scores on existing single-cell benchmarks not enough?
A good score can arise from tasks that are very similar to training or are solvable by simple patterns. For predictive biology, it is crucial whether a model generalizes to new perturbations, cell types, tissues, or other previously unseen conditions and provides better predictions there than strong simple baselines.
Can larger foundation models simply solve the 15 problems through scaling?
This is not proven. Larger models can ingest more data and complex representations, but the challenges also involve data quality, causality, biological constraints, generalization, and experimental validation. Furthermore, the Nature Methods benchmark from 2025 shows that model complexity alone does not guarantee a reliable advantage in perturbation predictions.
Which challenge is most far-reaching for medicine?
Challenge 15, "Clinical trial outcomes," is at the end of the translational level. It requires a model to sufficiently integrate many preceding levels: molecular mechanisms, cell states, efficacy, toxicity, organismal responses, and clinical heterogeneity.
Conclusion
"Fifteen challenges for generative AI applications to cell biology" is primarily a plea for a more demanding definition of progress. The 15 tasks shift the focus from impressive data reconstruction to generalization, mechanism, intervention, and prospective confirmation. The further the list leads towards translation, the less sufficient is a model that merely plausibly reproduces biological patterns.
For the debate about AI Virtual Cells, this is a helpful benchmark: a real virtual cell would not simply be a very large model for gene expression, but a system that predicts new biological situations, indicates its uncertainty, and repeatedly delivers useful hits in real experiments. Until then, the 15 challenges are less a proof of what has been achieved than a concrete research agenda for what is still missing.