

Topic
Precision Health
Anticipation Committee Chair:

Emmanouil Dermitzakis
Precision Health
The enabling technology is genomic sequencing, which gives the complete DNA dataset of an individual. The first sequencing of a complete human genome, completed in 2003 after 13 years and an investment of approximately $3 billion, established the reference map against which individual genomes are read. 1 Advances in sequencing chemistry, instrumentation, and data-processing have since reduced both cost and turnaround time by many orders of magnitude, making it feasible to sequence the genome of a critically ill patient within hours and to contemplate population-scale programmes that would generate genomic profiles for millions of individuals.2
Alongside DNA sequencing, the parallel development of high-throughput measurement platforms for proteins (proteomics), metabolites (metabolomics), and gene expression (transcriptomics) is providing complementary molecular layers that together describe an individual’s biology with unprecedented resolution.
Realising the clinical potential of this molecular information requires solving a set of interconnected problems that span technology, data science, medicine, and governance. The genetic variant, or mutation, responsible for a patient’s condition must be identified among millions of variants present in any human genome, then correctly interpreted and linked to a therapy. Population-scale datasets from diverse human populations are needed to train the interpretive models on which this depends. The resulting data must be shared across institutions and borders, while respecting individual privacy, community interests, and increasingly assertive national frameworks for data sovereignty. These challenges define the current frontier of precision health.
Key takeaways
Precision health is undergoing a rapid transition from specialist research tool to clinical standard, with an individual’s genomic information beginning to inform diagnosis, prevention and treatment across medicine. It is data-driven health and in some cases can become personalised. Rapid genomic diagnosis and clinical translation is now evidence-based and technically ready for broad implementation, though challenges to infrastructure remain. Gene discovery and mechanistic understanding of disease is accelerating, with new sequencing technologies revealing a far richer landscape of pathogenic variation than previously accessible. Multi-omics integration and population-scale precision medicine is progressing rapidly through large biobank programmes that link genomic, proteomic, metabolomic, and clinical data, the main remaining challenge involving moving from association to understanding causation. Genomic data governance and equitable implementation represents an increasingly urgent programme of work: growing data-sovereignty legislation risks fragmenting the global knowledge-sharing infrastructure on which scientific progress depends, and genomic databases remain heavily biased toward populations of European ancestry.
Rapid genomic diagnosis and clinical translation
Future Horizons:
5-yearhorizon
Diagnostic genomics enters mainstream clinical care
10-yearhorizon
Population-scale genomic screening becomes feasible
25-yearhorizon
Universal genomic profiling is a clinical standard
The reliability of variant interpretation remains a challenge. A third of all genetic tests performed for symptomatic patients return a variant of uncertain significance (VUS) — a genetic change that cannot be confidently classified as pathogenic or benign with currently available evidence.3 Resolving a VUS requires access to evidence from other individuals carrying the same variant, “functional assay” — or experimental — data characterising the molecular consequences of the change, and computational tools that can integrate these inputs at scale.
A further challenge is the multitude of non-coding variants that are associated with disease risk and the applicability of polygenic risk scores (PRS) for risk prediction. The analogue and combinatorial nature of PRS means that they can be interpreted only by population genetics: as technologies and associations improve, use of PRSs will probably play an increasing role in general health risk assessment, but for this to happen, population specific PRS frameworks will need to be shared globally.
Progress therefore depends critically on two developments: global data-sharing infrastructure for linked genomic and phenotypic data, and programmes that systematically generate clinically relevant evidence for variants in disease-relevant genes before they are encountered in patients.4 , 5 Even if these two developments are fulfilled, all the technologies mentioned depend on statistical associations, which creates a further challenge. While this can be achieved with very large numbers, it requires scaled resources for genotyping (identifying the genetic variant) and phenotyping (identifying the physical outcome of the variant, which is the most expensive process), and lacks mechanistic validation.
Thus, a third development is necessary: the integration of experimental physiological, cellular, and genetic information through targeted experimentation. This integration of genetic, genomic, clinical, and in vitro models with molecular sciences is being pioneered in the development of cardiac digital twins.6 Such holistic data integration moves the utility of genetic databases from single pathogenic assignment to a simulated model of each individual patient taking into account both confounding and contributory factors. The technical challenge is the computational integration of digital genomic data with analogue physiology models. Here, we anticipate that advanced AI approaches will be essential.
The pipeline of genetic therapies is expanding in parallel with diagnostic capabilities. Several hundred genetic therapies have received regulatory approval or are in advanced clinical development. Current trends, if sustained, suggest that the number of available therapies could reach several thousand within a decade, since identifying a pathogenic variant increasingly means identifying a treatment pathway. As this progress occurs, there is intensifying pressure to extend genomic medicine beyond specialist centres into mainstream clinical systems — and beyond high-income countries to the full global disease burden.
Gene discovery and mechanistic understanding of disease
Future Horizons:
5-yearhorizon
Ground-truth catalogues are established for key disease genes
10-yearhorizon
Systematic functional interpretation of variants advances
25-yearhorizon
Variant interpretation scales to the full disease genome
Long-read sequencing technologies are revealing a rich and complex landscape of pathogenic DNA variation, such as large chromosomal inversions and translocations, splice-site mutations, methylation defects, structural variants, and non-protein-coding regulatory changes that alter gene expression (the translation of a genetic sequence into a protein, which then carries out a specific biological function) without modifying the protein-coding sequence itself.14 Challenges remain, for example the development of frameworks for interpreting variants in the 98 per cent of the human genome that lies outside protein-coding regions, much of which controls when and where genes are switched on and off in the body (also of functional importance).15
The field requires a systematic programme to generate ground-truth functional data at scale. One pathway is to use saturation mutagenesis — the systematic creation and functional measurement of every possible variant in a defined set of key disease genes. This will produce reference catalogues that allow any subsequently observed variant to be interpreted against experimentally validated benchmarks.16 Such catalogues would enable large language models trained on genomic and clinical data to reinterpret existing patient datasets, dramatically reducing the proportion of unresolved cases. A further challenge is to characterise non-protein coding variants and understand why the same pathogenic variant produces different severities of disease in different individuals. This requires adding data concerning genetic background effects (such as the rest of an individual’s genetic sequence), environmental exposures, and immune and metabolic states that modify the consequences of a primary genetic variant.
Unique population structures are proving scientifically valuable. Consanguineous (closely related) populations, for example, in which a higher proportion of individuals express rare recessive variants, provide natural means of monitoring human gene function. Individuals who carry copies of a loss-of-function variant in a gene and are nevertheless healthy provide data that increases the precision of disease models and accelerates the interpretation of variants observed clinically.17 Research programmes anchored in these populations, particularly across the Arabian Peninsula, are yielding gene discoveries and therapeutic insights of global relevance.18
At the mechanistic level, investigation of rare developmental conditions is exposing the complex relationship between genetic mutation and disease. In oncogenic (cancer-causing) pathways such as the RAS-MAPK cascade, for example, somatic mutations — those arising in a subset of cells after fertilisation rather than being inherited — are being identified in non-disease-causing proliferative conditions including complex vascular malformations. This demonstrates that the phenotypic outcome of a mutation depends critically on where and when in the body’s development it occurs.19 Similarly, the emerging biology of clonal haematopoiesis, in which somatic mutations accumulate in blood stem cells during ageing, is expanding understanding of how genetic change in adult tissues contributes to cardiovascular disease and neurodegeneration, besides the better-understood link to cancers.
Multi-omics integration and population-scale precision medicine
Future Horizons:
5-yearhorizon
Proteomic profiling transitions from research to clinical platforms
10-yearhorizon
Dynamic multi-omic risk models enter clinical practice
25-yearhorizon
Precision health shifts to lifelong disease prevention
This includes describing how gene activity varies across tissues and cell types (transcriptomics), how proteins are expressed (proteomics), how metabolites reflect the chemical state of the body (metabolomics), and how these molecular layers interact with each other and with environmental exposures over a lifetime. Large biobank programmes that link genomic data to deep molecular profiling in hundreds of thousands of individuals are already yielding insights that would have been unattainable a decade ago. Examples include the identification of genetic variants that influence the abundance of thousands of circulating proteins, the discovery of shared biological mechanisms underlying apparently distinct diseases, and the demarcation of disease subtypes by molecular rather than clinical criteria.20
A central scientific challenge is the transition from association to causation. Genomic variants are useful here: they are uniquely well-suited to causal inference because germline genetic variation (that inherited from parents) is fixed in the DNA at conception and cannot be confounded by disease status.21 The other molecular layers, on the other hand — protein levels, metabolite concentrations, and gene expression — are dynamic and can be altered by disease as much as they cause it. This means that extracting true causal relationships from such data requires approaches that anchor multi-omic associations in genetic evidence. It also requires new models of interventional study design: small, precisely instrumented trials that perturb specific exposures and measure downstream molecular consequences in a standardised way.
The emerging field of exposomics is increasingly recognised as a necessary complement to genomics.22 Exposomics is the systematic measurement of the totality of environmental exposures experienced by an individual, from chemical pollutants and pharmaceutical agents to diet, microbiome composition and psychosocial stressors. Wearable sensors are generating streams of physiological data at population scale for this purpose. Together, these environmental and physiological data streams constitute what might be termed the “wearable omic”: a layer of longitudinal biological information distinct from, but potentially as informative as, static molecular measurements. Integrating such data with genomic and proteomic profiles, and doing so in a way that can establish causal relationships, will be a defining methodological challenge for the coming decade.
Researchers hope to be able to convert multi-omics knowledge into actionable clinical guidance for individuals who are not yet ill. Risk-stratification tools based on polygenic scores alone are insufficient to motivate preventive intervention: when communicated as a 20 or 30 per cent lifetime risk of a common disease, they rarely change behaviour.23 Predictive tools that can identify individuals who face a greater than 50 percent probability of developing a specific disease will be required to shift the centre of gravity of medicine toward prevention, because this is the threshold at which clinical action becomes clearly justified and preventive interventions economically defensible. This will mean developing ways to integrate genomic, proteomic, and longitudinal physiological data into reliable, dynamic, and continuously updated risk estimates.
Genomic data governance and equitable implementation
Future Horizons:
5-yearhorizon
Federated research becomes the norm; governance matures
10-yearhorizon
Policy interoperability replaces technical interoperability as the central challenge
25-yearhorizon
Genomic data operates as a global scientific commons
The interpretation of a patient’s genetic variant, for instance, requires evidence from thousands of other individuals carrying similar variants. Similarly, the training of AI models that can generalise across populations requires diverse, globally representative datasets, and the study of rare diseases requires connecting the small number of patients worldwide who share a condition. Yet the infrastructure for doing this — legal, technical, institutional, and cultural — is fragile and under increasing strain.
Unfortunately, the tools historically used to protect individual privacy and institutional interests — consent requirements, institutional review processes, data-use agreements — were designed for a world of small-scale, single-institution studies and are poorly adapted to the scale and connectivity of modern precision-health research.24,25 Governance frameworks that treat privacy tools as a bottleneck to be managed around, rather than as instruments for responsible science, are a significant systemic problem. It needs to be appreciated that the risks of not sharing data — failed diagnoses, missed discoveries, AI systems that perform poorly in under-represented populations — are as real and as serious as the risks of sharing it.
Legislation concerning data sovereignty represents a rapidly escalating structural challenge. Frameworks emerging across Asia, the Middle East, and elsewhere are treating genomic data as a strategic national asset and restricting its flow across borders.26 This has led to a shift toward federated research models, in which computational analyses are brought to the data rather than the data being centralised. While federated approaches preserve sovereignty and are technically sound, they do not resolve the underlying governance challenge. Even where data physically stays within borders, the rules governing its analysis must be interoperable for collaborative international science to function. The next phase of governance development requires not merely technical interoperability but policy interoperability: harmonised ethical and legal frameworks that allow federated systems to operate as a coherent global scientific resource.
The representational composition of genomic databases is not just an ethical problem: it also skews the science. Databases and AI models trained primarily on individuals of European ancestry produce diagnostic and predictive tools that perform less reliably in other populations, compounding rather than resolving health disparities.27,28 It will be vital to anchor population genomics programmes in the Middle East, Africa, and South and East Asia. These can leverage population structures, disease patterns, and genetic architectures distinct from those found in European cohorts, providing both a scientific opportunity and a means by which non-European countries can develop precision-health capabilities on their own terms. Furthermore, the concentration of analytical and AI capacity in a small number of countries and institutions risks reproducing, at a new technological level, the inequities that have characterised earlier phases of biomedical research.
Finally, equitable implementation requires not only the willingness to share data, but also solutions to the cost of informational access. As sequencing becomes a commodity, the data integration, processing, and access will be the critical bottleneck for broad implementation — and the cost of data access, integration, and maintenance is significant and growing. The inclusion of AI in the interrogation of such data further augments the cost. While common databases like OMIM29 are freely available, more advanced AI-based medical tools such as OpenEvidence30 are proprietary. The health solutions that can be guided by genomic approaches (such as gene therapy) are also often costly and will not be available to most people in most countries. A path forward here will require focussed discussions between economic and financial experts that explore ways to rationally finance these powerful innovations.31
Citations
Topic brief
- International Human Genome Sequencing Consortium (2001). Initial sequencing and analysis of the human genome https://doi.org/10.1038/35057062