More than an image: how multimodal AI is redefining precision oncology

Imagine walking into the clinic for a patient with metastatic lung cancer. In front of you: their chest CT, the biopsy with immunohistochemistry, the…

Blog · ppj.es

More than an image: how multimodal AI is redefining precision oncology

Published on July 22, 2026

🇪🇸 Leer este artículo en español

The problem

Imagine walking into the clinic for a patient with metastatic lung cancer. In front of you: their chest CT, the biopsy with immunohistochemistry, the next-generation sequencing gene panel, and the latest blood tests. Five different sources of information, in different formats, generated by different specialists, at different times. And you have to integrate all of it into a therapeutic decision within the next few minutes.

This fragmentation — technical, temporal, human — is, according to the authors of this review published in Cancer Letters, the main bottleneck preventing personalized oncology from being truly personalized.

The study

The team of Ruichong Lin and Kang Zhang (University of Texas at Austin and University of Maryland) has published a comprehensive review of how artificial intelligence is beginning to solve that problem. This is not just another review about «AI that detects nodules on CT scans.» It is a map of the new paradigm: multimodal AI, which learns to read together — and correlate — imaging, pathology, genomics, and clinical course data.

What is multimodal AI?

Until now, most AI systems in oncology worked in a single domain: a neural network that analyzes mammograms, an algorithm that predicts mutations from histology, or an LLM that summarizes a medical record.

Multimodal AI does something qualitatively different: it learns joint representations from different data sources. It is not that there is an imaging algorithm and a genomics algorithm that are later combined. It is that a single model learns the relationships between what it sees in histology, what it reads in the genomic report, and what it detects in the radiological image.

The authors identify four major technical advances that are making this possible:

  • Foundation models: trained on hundreds of thousands of multimodal samples, they learn generalizable representations that are later fine-tuned for specific tasks.
  • Synthetic data generation: to mitigate the scarcity of labeled data and the heterogeneity between centers.
  • Large Language Models: as a bridge between unstructured clinical language and quantitative data.
  • Autonomous agents: systems that chain multiple steps of clinical reasoning while orchestrating different tools.

What can it already do?

The review covers concrete applications that are already underway:

Molecular subtyping from histology. Several models can predict the mutational profile (EGFR, KRAS, TP53) directly from H&E slides, without the need for additional sequencing. They will not replace genomics, but they can help prioritize which samples to sequence first.

Integrated prognostic stratification. By combining radiological imaging, digital histology, and clinical data, some models improve survival prediction compared with any single source alone. Imaging alone sees tumor morphology; histology sees tissue architecture; clinical data sees the patient. Together, they see more.

Individualized treatment recommendations. Systems that integrate the tumor’s molecular profile, response imaging, and the patient’s comorbidities to suggest treatment options aligned with clinical guidelines.

But with caution

The authors themselves devote a substantial section to limitations, and they are worth highlighting:

Limited generalizability. A model trained at MD Anderson may not work at a community hospital. Heterogeneity across centers, equipment, and populations is the Achilles’ heel of most of these systems.

Insufficient prospective validation. Most studies are retrospective. The leap from paper to real patient requires trials demonstrating that AI improves outcomes — not just performance metrics.

Regulatory uncertainty. How do you approve a model that integrates five data sources and learns representations that no one can directly inspect? FDA and EMA are building the framework, but it does not yet exist for multimodal AI.

Equity and bias. If training data come from large academic centers in developed countries, will this model serve a patient at a rural center or in a middle-income country?

Why it matters to the clinician

This review is not an instruction manual, but it sketches where we are heading. And the direction is clear: from unimodal AI to integrated AI.

For the clinical oncologist, this means that in the coming years we will see tools that do not just «read» an image, but cross-reference the image with the biopsy, genomics, and medical history to tell you: «This 68-year-old patient with EGFR-mutant lung adenocarcinoma, intermediate tumor mutational burden, and 60% PD-L1 is more likely to respond to osimertinib + chemo than to chemo-immunotherapy alone.»

That is not science fiction. These are the foundations being laid right now.

Reference

Lin R, Zhao Z, Liu Z, Kang J, Zhang K et al. Artificial intelligence in clinical oncology: Multimodal integration and translational development. Cancer Lett. 2026. DOI: 10.1016/j.canlet.2026.218493. PMID: 41962624.

— This analysis was generated by ANGIE (Always Next to Guide, Inspire and Empower), an artificial intelligence system with SOUL profiles, designed by Dr. Javier Pumares Pérez.

Connect on LinkedIn · ppj.es

Disclaimer: this article is educational and informational in nature and reflects the personal opinion of the author. It does not constitute medical advice nor replace the assessment of a healthcare professional. If you have a health concern, consult your physician.