Notes
When prediction is not explanation
From protein language models to the person as data

Artificial intelligence is learning regularities in proteins, genomes and health trajectories with extraordinary effectiveness. But a protein does not interpret a prediction about itself; a person can. Between those two scales lies a scientific, anthropological and political boundary that should not be erased.
Long before the current expansion of artificial intelligence, Alfonso Valencia’s scientific career was already shaped by a question that has now become central. How far can we infer what a biological system does from the data we have about it?
In 2000, Damien Devos and Alfonso Valencia published Practical limits of function prediction. At the time, the problem was protein function. The number of known sequences was growing much faster than the experimental capacity to establish what each protein did, so bioinformatics sought to transfer functions from known proteins to similar ones. Their analysis showed that finding patterns can support useful predictions, but that this transfer has limits arising from database inaccuracies, evolutionary divergence and the functional plasticity of proteins themselves (Devos & Valencia, 2000).
Twenty-six years later, the scale has changed. Deep-learning models predict molecular structures, explore genome regulation, process medical records, simulate tumours and generate synthetic biological data. Yet the underlying question remains.
What exactly do we know when a prediction is correct?
The issue becomes especially interesting when we follow the path from an amino-acid sequence to a person. At each change of scale, not only the number of variables increases; so do the kinds of relationships we must take into account. These include development, environment, history, language, institutions, inequality and the capacity to act upon one’s own future.
A digitised medical record does not contain a life. It contains those parts of a life that a particular health system managed — and chose — to turn into data.
When evolution becomes a corpus
The similarity between language models and models that work with proteins deserves careful explanation.
A language model learns regularities within sequences. In natural language, those sequences are made up of discrete units — tokens — whose probability depends on context. Some models are trained by predicting the next element; others by reconstructing hidden elements. In both cases, learning emerges from detecting dependencies across enormous quantities of sequences.
Proteins offer a formal structure that is strikingly well suited to a similar problem. They, too, consist of sequences of discrete units. In this case, amino acids. The alphabet is much smaller — twenty standard amino acids — and the natural sequences available to us are the product of an immense evolutionary history.
The analogy is not that a protein ‘speaks’. It is that evolution has produced a vast corpus of sequences in which some combinations appear, disappear or remain conserved under physical, chemical and functional constraints. Those regularities can be learnt statistically.
Rives and colleagues (2021) trained language models on 250 million protein sequences containing around 86 billion amino acids. Without being given explicit annotations about structure, the models’ internal representations began to organise information related to biochemical properties, homology, and secondary and tertiary structure. The apparently simple objective of reconstructing sequences forced the system to capture some of the constraints that evolution had left embedded within them.
This reveals a first important connection with large language models. Context can be used to infer relationships that are not explicitly labelled in the data.
AlphaFold2, however, is not strictly an LLM.
Its architecture is specifically designed for the protein-structure problem. Its Evoformer core processes multiple-sequence alignments and pairwise residue representations using attention-based components alongside other specialised mechanisms. The evolutionary information contained in those alignments helps infer which amino acids have varied together over evolutionary time and, therefore, what spatial relationships may exist between them (Jumper et al., 2021).
The connection with language models becomes even clearer with ESMFold. Lin and colleagues (2023) showed that a large language model trained directly on protein sequences could infer atomic structures from a single sequence. When scaled to 15 billion parameters, its internal representations contained enough structural information to produce three-dimensional predictions at high speed.
This allows the resemblance to be stated precisely.
Protein language models work because sequences contain contextual and evolutionary regularities from which properties that are not explicitly written at each position can be inferred.
In a sense, evolution generated the training corpus.
But this is also where part of the analogy ends.
Amino acids are not words. A protein does not hold a conversation, negotiate the meaning of a sequence or alter its behaviour because it has learnt that someone has made a prediction about it. To confuse the computational usefulness of the linguistic metaphor with an identity between proteins and human language would be to turn a methodological tool into an ontological claim.
A protein is not a person
That jump in scale recalls a warning developed by Alfredo Francesch Díaz in a different context.
In Sabores y sinsabores de un programa darwinista para las ciencias sociales [Pleasures and pitfalls of a Darwinian programme for the social sciences], Francesch Díaz (2010) critically examines what happens when an explanatory programme built in biology is extended into social phenomena. His work does not discuss artificial intelligence or anticipate protein language models. The connection proposed here is an analogy. A framework can be extraordinarily productive within one domain without thereby becoming a universal explanation.
AI now presents a particularly tempting case.
If a transformer can discover relevant relationships among amino acids; if another model can predict a protein’s structure; if AlphaGenome can learn relationships between sequence and regulation, why not continue upwards to model tissues, organs, behaviour and even persons?
The answer is not to deny that such modelling is possible. It is to recognise that each scale introduces new properties and relationships.
A protein is more than a sequence, but its structure is strongly constrained by physical properties and by constraints accumulated through evolution. A person, by contrast, develops within environments that they change and that change them, learns categories, participates in institutions, interprets meanings and can react to classifications produced about them by other people or by systems.
Adding more variables does not automatically resolve this change in the nature of the problem.
The genome is not a closed programme
The history of genomics offers a second example.
Sequencing the human genome encouraged a powerful image for years. If we could read the entire code, we would understand the organism. Yet between knowing a sequence and explaining a phenotype lies a network of gene regulation, development, cellular interaction and environmental conditions that cannot be reduced to a linear reading of DNA.
Bioinformatics itself has progressively moved away from isolated sequences towards increasingly heterogeneous systems. Cirillo and Valencia (2019) describe personalised medicine as a problem of integrating genomic and multi-omic data, imaging, medical records and data produced by personal devices. They also warn that large biomedical datasets are heterogeneous, incomplete and imprecise, and that our ability to generate data is advancing faster than our ability to interpret it.
Eugenia Ramírez Goicoechea’s biosocial anthropology offers a particularly productive framework for thinking through this difficulty. In Antropología biosocial: Biología, cultura y sociedad [Biosocial anthropology: Biology, culture and society], human evolution does not appear as the linear unfolding of a genetic programme, but as a process in which phylogeny, ontogeny, epigenesis, plasticity and environment are intertwined (Ramírez Goicoechea, 2013).
Her formulation is especially clear. ‘Lo genético nos define como posibilidad a realizarse, pero sólo en la ontogenia devenimos humanos’ [The genetic defines us as a possibility to be realised, but it is only through ontogeny that we become human] (Ramírez Goicoechea, 2013, p. 180).
The gene therefore ceases to occupy the position of a sufficient cause. Development is not merely the period during which a completed programme is executed; it is the process through which particular potentials are realised under specific spatial, temporal, biological and social conditions.
AlphaGenome shows how far prediction can now advance within this territory. The model analyses sequences of up to one million base pairs and predicts multiple modalities of regulatory activity. In the published evaluations, it matched or outperformed specialised models in 25 of 26 variant-effect prediction tasks (Avsec et al., 2026).
This is an extraordinary result, but it is not equivalent to having simulated the complete development of an organism. Between a genetic variant and a life remain cells, tissues, microbiome, diet, environmental exposure, disease, learning and social relationships.
The problem therefore ceases to be only how much we know about the genome and becomes what that genome does within a system that is continuously developing.
From clinical trajectory to embodied biography
Another transformation appears when we move from the genome to medical records.
Delphi-2M, published in 2025, was trained on data from approximately 400,000 UK Biobank participants and externally validated using 1.9 million people in Denmark. The system probabilistically models the temporal evolution of more than one thousand diseases and can generate possible long-term health trajectories (Shmatko et al., 2025).
For the model, a clinical trajectory is a sequence of encoded events. It includes diagnoses, medicines, ages and temporal relationships.
From a biocultural perspective, we can also read it as the partial trace of an embodied biography.
Behind hypertension, respiratory disease or metabolic alteration was a body that developed in particular environments. Diet, pollution, material conditions, work, care, stress, access to healthcare and other experiences can have bodily consequences without necessarily being recorded with the same level of detail in a database.
Ramírez Goicoechea (2013) insists precisely that the sociocultural is not a later addition to a previously completed human nature. Human biology develops within concrete historical and social environments.
This does not mean that every disease can be reduced to culture. It means that a clinical trajectory represents some dimensions of a life and leaves others out.
And what was not measured, categorised or incorporated into healthcare infrastructure can disappear from the model even if it materially contributed to the process the model is trying to predict.
A prediction about a person can change the person
Here, Carissa Véliz’s philosophy introduces a decisive difference.
In Prophecy: Prediction, Power, and the Fight for the Future, from Ancient Oracles to AI, Véliz (2026) examines the long social history of prediction and draws attention to a particularly important property when people are the object of prediction. Predictions can alter expectations, decisions and behaviour and, in doing so, help transform what they purported merely to describe.
The difference from a protein is fundamental.
If a model predicts the structure of a protein, the protein does not learn about the prediction and decide to fold differently in response.
A clinical prediction, by contrast, enters a social circuit. It can alter the behaviour of the person who receives it, decisions made by healthcare professionals, the intensity of medical monitoring, insurance arrangements or the institutional allocation of resources. The prediction becomes part of the future conditions of what it seeks to anticipate.
Not all predictions therefore have the same status.
A molecular prediction can be assessed primarily by its correspondence with a physical phenomenon. A prediction about a person can additionally acquire performative effects.
This makes the conversion of probabilities into identities particularly dangerous. A risk score does not say who a person is, nor does it necessarily describe what will happen to them. It expresses an inference produced from particular data, categories and reference populations.
Véliz (2026) therefore turns an epistemological problem into a question of power. Prediction is not only about describing possible futures; it can also distribute the capacity to make decisions about them.
Better prediction also means knowing where not to generalise
Explainability and bias are not objections imposed from outside by the social sciences upon a biomedicine indifferent to its own limits.
They are part of Valencia’s own research.
Cirillo and colleagues (2020), with Alfonso Valencia among the authors, analysed sex and gender differences and biases in artificial intelligence for biomedicine and healthcare. The study warns that ignoring these dimensions can produce suboptimal outcomes, errors and discriminatory effects. It also emphasises the importance of appropriate datasets and explainable AI tools in moving towards personalised medicine that does not amplify existing inequalities.
It is therefore not enough to increase the average accuracy of a prediction. We also need to know the conditions under which it can fail, which populations are represented, what categories the system uses and what degree of confidence can be placed in its results.
This concern connects very different moments in Valencia’s work. In 2000, it appeared as a limit on predicting protein function (Devos & Valencia, 2000). In the community assessment of AlphaFold, it reappears in the need to interpret confidence metrics critically (Akdel et al., 2022). In biomedical AI, it becomes a question of representation, explainability and equity as well (Cirillo et al., 2020).
Better prediction also means learning where we should not generalise.
From tumour to digital twin
Multiscale models reveal another boundary.
Ponce-de-Leon and colleagues (2022), with Valencia among the authors, developed a multiscale tumour-growth model to explore treatment schedules and the emergence of resistance. The aim was not to produce a complete replica of a person, but to represent particular tumour processes and experiment computationally with their possible responses.
That distinction is essential when discussing digital twins.
A model can incorporate case-specific information and simulate particular biological behaviours without thereby becoming a complete copy of the person.
The difficulty grows with every change of scale. A protein structure can be delimited with relative clarity. A cell involves molecular networks and relationships with its microenvironment. A tissue introduces interactions between cell populations. An organism requires the integration of distinct systems, temporalities and environmental conditions.
When we reach a person, the problem is no longer simply one of adding enough variables.
A person interprets, learns, remembers, establishes relationships, modifies their environment and is modified by it. They may even change their behaviour because they know about a prediction concerning their own future.
That relational character is also central to Ramírez Goicoechea (2013). The human mind and the capacities associated with humanisation do not emerge outside the processes of attachment, communication, learning and intersubjectivity through which a person develops.
Speaking and thinking are not merely sequences
The analogy between language and proteins becomes even more interesting when we return to human language.
In Hablar y pensar, tareas culturales: Temas de antropología lingüística y antropología cognitiva [Speaking and thinking, cultural tasks: Topics in linguistic and cognitive anthropology], Honorio Velasco Maillo (2013) approaches the origins of language through its biological dimensions — the brain and vocal tract — but also through the social group, interaction, sociability, tools, symbols, speech communities and linguistic diversity.
The brain is necessary for speaking.
But an isolated brain does not explain language.
This observation allows us to return to protein language models from the opposite direction. Their success shows that treating a biological sequence as if it were a language for computational purposes can be extraordinarily productive. It does not demonstrate that human language can be reduced to an equivalent sequence.
In a protein, dependencies between positions arise from biophysical, functional and evolutionary constraints. Human language also involves intention, conventions, social relationships, history, pragmatics and shared contexts. Words do not mean only because of where they appear in a string.
We can record words, tokenise them and learn their statistical regularities. LLMs demonstrate how much structure can be extracted from that procedure.
But the ability to predict linguistic sequences does not exhaust what it means to speak.
Velasco therefore enables us to invert the metaphor. The conclusion is not that proteins are language, but to ask why the idea of language works so well as a computational tool in biology while remaining insufficient to explain human language in full.
Synthetic data and the question of fidelity
The same problem reappears with synthetic data.
A review developed within ELIXIR and co-authored by Alfonso Valencia examines the growing use of synthetic data in the life sciences. Synthetic datasets can help address scarcity, privacy constraints or limited access to real-world data, but their usefulness depends on evaluating whether they preserve the relevant properties of the phenomenon they are intended to represent. The review notes that generation techniques are advancing rapidly while standardised procedures for measuring fidelity, utility and reliability remain insufficient in many domains (Fragkouli et al., 2026).
The question is technical, but it is also epistemological.
A synthetic dataset does not need to reproduce every detail of the original. It needs to preserve what is deemed relevant for a particular purpose.
The question, therefore, is not simply whether a simulation appears realistic, but who defines which properties it must preserve in order to count as valid.
The same dilemma appears at different scales in a predicted protein structure, a simulated tumour, a synthetic clinical cohort or a possible human digital twin.
Making something computationally legible requires deciding what can be omitted.
Legibility, power and governance
Every database is also a selection.
For a machine to analyse a reality, that reality has to be translated into variables, sequences, images, categories, labels and quantifiable relationships. This translation can be extremely useful, but it is neither transparent nor politically neutral.
Ignacio Iturralde Blanco’s political and legal anthropology offers another point of support here, although the connection with AI is again ours.
In his analysis of the Mixe normative system and conflicts between Indigenous autonomy and state law, Iturralde Blanco (2011) shows how a dominant normative order can name, classify and only partially recognise local systems whose logics do not fully coincide with state categories. The problem is not computational, but it shares a fundamental question with data governance. Who has the authority to define the categories through which a reality becomes administratively legible?
The issue is particularly important in biomedicine.
So far, many of the most visible cases come from major infrastructures in the Global North, including DeepMind, the Barcelona Supercomputing Center, UK Biobank and European health records. The geography of data matters.
The AGenDA project provides a significant counterpoint. As part of H3Africa, it begins from an explicit observation. African populations remain substantially underrepresented in global genomic research and databases despite the continent’s extraordinary genetic diversity. AGenDA identified underrepresented groups in nine African countries for whole-genome sequencing and incorporated community engagement, ethical approval, compliance with national legislation and a shared governance framework. The project is led from Africa by African teams that participate in decisions about data sharing (Ramsay et al., 2026).
The case allows us to reformulate the question of algorithmic legibility.
A population that is barely present in the data is not equally legible to the model.
But correcting that absence is not simply a matter of extracting more data.
It also requires asking who collects the data, where it is stored, who determines the conditions of reuse, which institutions subsequently develop the models and how potential benefits return to the communities that made the knowledge possible.
Couldry and Mejias’s (2019) concept of data colonialism can serve here as a framework for caution, not as an automatic label. Not every collection of biomedical information constitutes a colonial relationship. But when different dimensions of life are systematically transformed into appropriable resources for classification, prediction and value generation, the power relations organising that transformation also become part of the object of study.
Personalised medicine does not begin or end with the algorithm.
What we think we are explaining when the model gets it right
The path from the limits of functional prediction through protein language models, AlphaFold, AlphaGenome, clinical trajectories, simulated tumours and synthetic data reveals something more interesting than a succession of technological advances.
Each change of scale also changes the kind of problem.
In proteins, millions of years of evolution have produced an immense corpus of sequences subjected to constraints that models can learn. Transformers and protein language models are particularly effective precisely because those sequences contain contextual, evolutionary and structural regularities.
But a person is not simply a longer sequence.
Ramírez Goicoechea helps explain why the genome is not enough. Development is constitutive. Velasco Maillo shows why language and thought cannot be exhausted by formal relations among signs. Francesch Díaz offers a warning about extending paradigms beyond the domains in which they demonstrated their effectiveness. Véliz adds a decisive difference. Predictions about people can intervene in the conditions of what they predict. Iturralde Blanco allows us to introduce the question of who defines categories and regimes of recognition. And Valencia’s work shows, from within bioinformatics, both the power of prediction and the need to interpret, validate and delimit its results.
The most productive question may therefore not be whether artificial intelligence ‘understands’ life.
There is an earlier question.
What version of life have we had to construct in order for a machine to predict it?
AI can make certain dimensions of our organisms extraordinarily legible. That capacity is already transforming scientific research and medicine.
But a life is not merely what we have managed to record about it.
And a prediction about a person does not merely try to anticipate the future. It can become one of the forces that helps produce it.
Perhaps the real challenge of future digital twins will not be to place an entire person inside a computer, but to recognise precisely what we had to leave out in order to say that the model works, and what effects the model may produce when it returns to the world.
Questions to keep thinking with
- What is the epistemological difference between predicting the structure of a protein and predicting the trajectory of a person?
- Which dimensions of a biography disappear when a life trajectory becomes a data trajectory?
- When does a prediction stop describing a possible future and begin to help produce it?
- Who should decide what a digital model needs to contain in order to represent an organism or a person validly?
References
Akdel, M., Pires, D. E. V., Porta Pardo, E., Jänes, J., Zalevsky, A. O., Mészáros, B., Bryant, P., Good, L. L., Laskowski, R. A., Pozzati, G., Shenoy, A., Zhu, W., Kundrotas, P., Ruiz Serra, V., Rodrigues, C. H. M., Dunham, A. S., Burke, D., Borkakoti, N., Velankar, S., … Beltrao, P. (2022). A structural biology community assessment of AlphaFold2 applications. Nature Structural & Molecular Biology, 29, 1056–1067. https://doi.org/10.1038/s41594-022-00849-w
Avsec, Ž., Latysheva, N., Cheng, J., Novati, G., Taylor, K. R., Ward, T., Bycroft, C., Nicolaisen, L., Arvaniti, E., Pan, J., Thomas, R., Dutordoir, V., Perino, M., De, S., Karollus, A., Gayoso, A., Sargeant, T., Mottram, A., Wong, L. H., … Kohli, P. (2026). Advancing regulatory variant effect prediction with AlphaGenome. Nature, 649, 1206–1218. https://doi.org/10.1038/s41586-025-10014-0
Cirillo, D., Catuara-Solarz, S., Morey, C., Guney, E., Subirats, L., Mellino, S., Gigante, A., Valencia, A., Rementeria, M. J., Santuccione Chadha, A., & Mavridis, N. (2020). Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare. npj Digital Medicine, 3, Article 81. https://doi.org/10.1038/s41746-020-0288-5
Cirillo, D., & Valencia, A. (2019). Big data analytics for personalized medicine. Current Opinion in Biotechnology, 58, 161–167. https://doi.org/10.1016/j.copbio.2019.03.004
Couldry, N., & Mejias, U. A. (2019). The costs of connection: How data is colonizing human life and appropriating it for capitalism. Stanford University Press.
Devos, D., & Valencia, A. (2000). Practical limits of function prediction. Proteins: Structure, Function, and Genetics, 41(1), 98–107. https://doi.org/10.1002/1097-0134(20001001)41:1%3C98::AID-PROT120%3E3.0.CO;2-S
Fragkouli, S.-C., Iqbal, S., Crossman, L., Gravel, B., Masued, N., Onders, M., Haseja, D., Stikkelman, A., Valencia, A., Lenaerts, T., Psomopoulos, F., Ó Broin, P., Queralt-Rosinach, N., & Cirillo, D. (2026). An ELIXIR scoping review on domain-specific evaluation metrics for synthetic data in life sciences. NAR Genomics and Bioinformatics, 8(1), lqag012. https://doi.org/10.1093/nargab/lqag012
Francesch Díaz, A. (2010). Sabores y sinsabores de un programa darwinista para las ciencias sociales [Pleasures and pitfalls of a Darwinian programme for the social sciences]. Antropologia Portuguesa, 26–27, 9–27.
Iturralde Blanco, I. (2011). Sistema jurídico dominante y autonomía indígena: El sistema normativo mixe y los conflictos jurisdiccionales [Dominant legal system and Indigenous autonomy: The Mixe normative system and jurisdictional conflicts]. En M. Aparicio Wilhelmi (Ed.), Contracorrientes: Apuntes sobre igualdad, diferencia y derechos [Countercurrents: Notes on equality, difference and rights] (pp. 87–116). Documenta Universitaria.
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Žídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A. J., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., … Hassabis, D. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589. https://doi.org/10.1038/s41586-021-03819-2
Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., Dos Santos Costa, A., Fazel-Zarandi, M., Sercu, T., Candido, S., & Rives, A. (2023). Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637), 1123–1130. https://doi.org/10.1126/science.ade2574
Ponce-de-Leon, M., Montagud, A., Akasiadis, C., Schreiber, J., Ntiniakou, T., & Valencia, A. (2022). Optimizing dosage-specific treatments in a multi-scale model of a tumor growth. Frontiers in Molecular Biosciences, 9, 836794. https://doi.org/10.3389/fmolb.2022.836794
Ramírez Goicoechea, E. (2013). Antropología biosocial: Biología, cultura y sociedad [Biosocial anthropology: Biology, culture and society]. Editorial Universitaria Ramón Areces.
Ramsay, M., Etheredge, H., Tluway, F., D’Amato, M. E., Chikwambi, Z., Hamdi, Y., Alhudiri, I., Fakim, Y., Ahmad, K. M., Belguith, N., Bentley, D., Boujemaa, M., Calumbuana, N., Chaouch, M., Charfeddine, C., Chinien, G., Dukuze, N., Eljilani, M., Elzagheid, A., … Choudhury, A. (2026). Enriching African genome representation through the AGenDA project. Nature, 649, 565–573. https://doi.org/10.1038/s41586-025-09935-7
Rives, A., Meier, J., Sercu, T., Goyal, S., Lin, Z., Liu, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., & Fergus, R. (2021). Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the National Academy of Sciences of the United States of America, 118(15), e2016239118. https://doi.org/10.1073/pnas.2016239118
Shmatko, A., Jung, A. W., Gaurav, K., Brunak, S., Mortensen, L. H., Birney, E., Fitzgerald, T., & Gerstung, M. (2025). Learning the natural history of human disease with generative transformers. Nature, 647, 248–256. https://doi.org/10.1038/s41586-025-09529-3
Véliz, C. (2026). Prophecy: Prediction, power and the fight for the future, from ancient oracles to AI. Swift Press.
Velasco Maillo, H. M. (2013). Hablar y pensar, tareas culturales: Temas de antropología lingüística y antropología cognitiva [Speaking and thinking, cultural tasks: Topics in linguistic and cognitive anthropology]. Universidad Nacional de Educación a Distancia.