You are listening to a Podhoc podcast — a platform where anything can be turned into a Podcast to Learn in Motion.
The assertion made by Demis Hassabis that in the next ten years we will cure all diseases. How do you see it? There are already publications indicating that by training with patient trajectories, you can predict five years in advance which disease you will have. It's not so inconceivable. Alfonso Valencia, a pioneer in bioinformatics and Director of Life Sciences at the Barcelona Supercomputing Center, and former president of the International Society for Computational Biology, leads the application of supercomputing in personalized medicine.
With over 400 published articles, he is a global reference in the use of artificial intelligence to decipher genomes and diseases like cancer. The impact in biomedicine is enormous because we can do things that are unimagined, things that just three or four years ago we wouldn't have even thought possible. We can approve a drug without knowing how it works. This is achievable, but it comes with its complexities.
In this context, the story will be: do you want us to cure this disease without any idea how it was done or managed, but it works? Or do you want us to explain it to you, and it takes twice as long to get there? I have a feeling we'll reach that decision point, won't we? It's very difficult to think you'll have the cure for all diseases in such a short time. What things do you see artificial intelligence making possible that you didn't think you'd see?
We've moved from artificial intelligence being a mere aid to doing things a bit faster, to AI helping us think and do things in a completely different way. However, I believe the majority of biological research is not automatable in the next three years. Just today, I received news from a hospital in the United States that published that using Palantir, that artificial intelligence platform, they saved over 800 lives in the past few years. What is the real impact AI is having on science?
Alfonso, it's a pleasure to have you here today. We'll discuss many things, but first, I wanted to temper the hype, or at least see if it's truly warranted. Because when we talk about medicine and science in artificial intelligence, everything is sky-high, right? We're being promised an autonomous scientific developer with AI by 2028. And just today, I received news from a US hospital that published that using Palantir, that artificial intelligence platform, they saved over 800 lives in recent years. So, what is the real impact AI is having on science?
Well, I think we need to differentiate between the impact on science and the impact on medicine. The impact on science, seeing details and specific cases, is I believe obvious and transformative. Things are being done in a way that is different from how they were done before. The impact on medicine is more complex to measure because it's a process, integrated within medical systems, within hospital systems. And even though there are very powerful developments when they are published, it's much harder to actually implement them in a system, in a hospital, and quantify, as in this article you mentioned, the number of lives saved.
In healthcare, as in other industries, the transformation of processes is very slow, and therefore, the implementation of new technologies will take longer. It's not simply about installing new software; it's about changing the way things are done, and that's difficult in a system as complex as a medical system. Therefore, I think we see less than what could be possible if medical systems were easier to manipulate or implement things in. But in science, at least in the part I know in biology, the impact in biomedicine is enormous because we can do things that are unimagined, things we wouldn't have thought possible three or four years ago.
We have questions from the 1990s that remain unanswered due to a lack of data, computational resources, and, of course, methods. Now, suddenly, we find ourselves with the data, the computational resources, and the methods. So, it's a wonderful time to do research in biology because you suddenly have instruments you didn't have before. Of course, this is one of the things that seems incredibly interesting to me when we see grounded data on the impact it's having in science. Many of these studies are derived from having used artificial intelligence for a period, observing the results, and then publishing a paper.
This means we are always a step behind what is happening today. When we see news about the impact in science or medicine, we are seeing a bit of what's in our rearview mirror. But this is advancing so rapidly that what is happening right now won't be in a paper for another year or until it's published, meaning the gains are even greater than the headlines suggest. Well, this has always been the case in science. Publications naturally lag behind discoveries, and as time accelerates, as you say, it becomes harder to keep up with that fast-moving hare.
So, yes, there's more right now. Our capacity for surprise is already at its limit, isn't it? You open another journal, and you find, "Oh, they've published X amount and so on." Furthermore, many of these new inventions and discoveries come from companies whose operations are not entirely clear, so the element of surprise is perhaps greater, isn't it? Because it's even harder to know what they were doing. It was much easier to know what was going to happen because you'd go to a conference, talk to colleagues, and know what was happening.
Nowadays, it's very difficult to know what's happening in a company doing something new. But, on the other hand, I imagine this is a golden age for you, isn't it? A time when you can achieve things that were previously impossible. If you had to explain the impact of AI in science, would you define it as AI helping scientists go faster, or have we reached a point where AI is contributing new science, making discoveries? How would you define it? Obviously, it's helping to do things faster, isn't it? You can make faster summaries, read more literature, collect it, search for things, and it's also suggesting ideas.
You start writing a project, ask an AI, and it ends up suggesting things. That's good; it's like an assistant. That's the minimum level, but it's becoming transformative. You're starting to be able to do things in fields where you couldn't before, or do things in a different way than you could. Or to combine technologies that previously couldn't be combined. So, I think we've gone from artificial intelligence being a mere aid to doing things a bit faster, to AI helping us think and do things in a completely different way. And I still believe we are far from the end of this. We'll see a greater degree of integration, a degree where you'll continuously use AI systems as another part of your lab, as another assistant that interacts with your students, your postdocs, your engineers, your guidance, and everyone interacts with each other, like another component of the system that will make you think differently.
But it always depends on the scientist. That is, in the coming years, who knows what will happen? I believe it will continue to depend on scientists. There will be areas that can be automated, not entirely, but to a good extent. These are areas with large datasets, with automatable, robotizable systems that will generate results from these data and can continue research. You can imagine this in chemistry or screening some compounds; those types of things might be completely automatable. But I believe most biological research is not automatable in the next three years. In five years, in ten years, how much will be automatable? That's a good question, one I'd also like to know the answer to. We don't know.
No, we don't know because, unlike other areas of physics where there's better-understood data on physical or mechanical systems, biological systems remain an area of discovery. We are still not sure how many types of cells we have, or we lack information about all organs. It's difficult to imagine automating everything to generate hypotheses about everything and have integrated systems for everything. Therefore, because there's a lack of knowledge and basic concepts, and we are still discovering things, it's much harder to imagine everything being automatically replaceable because it requires a much higher level of hypothesis generation, knowledge, and abstraction.
Of course, this is where I was going to ask you: what problem are you trying to solve? I mean, as a biologist, for people to understand your work, what is the objective of your work? This is a big question with a not-so-trivial answer. The most global objective of our work is to try to understand the molecular basis of biological processes. And in most cases, we refer to diseases. That is, why does a disease happen? Why does a person have a disease? Explained not from a phenomenological point of view, because they have a virus or whatever, but what is happening internally, what are the molecular processes, what is triggered to cause a disease?
And we translate this in large part to the fact that after having one disease, you have a greater or lesser probability of having a second disease. That is, what is the trajectory of a disease? But not from a statistical point of view, like "statistically, this happens, then that." What interests us is understanding what is happening internally at the molecular level, which proteins, which genes have been activated and deactivated to produce it. Why? Because the ultimate goal of biology is mechanistic knowledge: how things happen. Why? Because if we don't understand how things happen, we won't be able to intervene rationally.
What instruments will you need? Sufficient genomic information, imaging, processes, sufficient computation. But this will be complex in any case, and technologies that allow you to move from those data to how things happen, because one protein activates another protein that, when it malfunctions, causes a mutation and cannot activate the next, leading to a specific blockage. That mechanistic process is what interests us scientifically, and its application is in diseases. And where are we in understanding biology? Are we very early, or do we already know a lot about humans?
This is a difficult question because, as we don't know what "everything" is, it's hard to know where we stand. But certainly, progress has been significant since a series of techniques, mainly genomics in its many varieties and diversification, began to be automated. Now, for the first time, we have data on how an organ is composed in multiple people, how different processes interact. Let's say we are at the end of the paleolithic era of biology. So, from when, more or less, for me to understand? From the genomics explosion. Starting from the late 1990s, the process of producing genomes began to be industrialized, which was the first thing that was going to be produced.
We naively thought that genomes would solve everything. That's not true. And now we can do genomes of individual cells. This has been a very fundamental technology. It's one thing to sequence a person's genome, which is essentially a genome for life, so to speak. And another is to see the expression level of each gene in each cell. That is, we have a single genome for our entire body. All our cells have, generally speaking, the same genome. But then we have different tissues and organs because the expression of those genes in those genomes has been modulated differently throughout development, leading to a specific tissue or organ.
So, this is much more complex than it seemed, right? I remember in the 90s, they talked about how discovering the genome would solve everything. And here we've just begun a new world, haven't we? Well, we scientists truly thought that by understanding the genome, we would understand diseases, but now we realize that's not the case. We need to do much more. And with these technologies, we can look at every cell in a tumor, try to see, or be able to see each cell of the tumor, which are different cell types, including the immune system, the tumor cells themselves, and normal cells. Each individual cell expressing at a given moment, and we can track it.
That level of cellular resolution wasn't available a few years ago. Now we're starting to have it, but there are many more techniques. Genomics has expanded into many different types of techniques that give us complete information. Now we talk about atlases; we have the lung atlas, the kidney atlas, the X atlas, and we're beginning to have this information that's starting to be interesting for modeling systems and gaining the mechanistic knowledge we need to know what process to target to treat a specific disease in a specific person. And the ultimate goal is to understand in order to treat, or to prevent, or to cure. I understand that understanding humans or biology itself will give us the necessary information to know what this disease will do and thus prevent it from happening or eradicate it entirely. That's the objective, isn't it?
This has always been the objective of biology, of molecular biology: to understand processes in order to control them. We've always said, and I still believe, that if we don't understand them and aren't able to model them, we won't be able to control them. We can't control them if we can't model them, and we can't model them if we don't understand them. These things are related. Now, all of this is suddenly a bit shaky because we can have a machine learning system, an AI, that learns from data and tells us how to treat it, and we have no idea about the intermediate process. And here begins an interesting discussion about, [exhales] to what extent is this appropriate, to what extent is it valid?
First, it's valid from a regulatory perspective. Can we approve a drug without knowing how it works? Yes, it has been done; we didn't fully understand how aspirin worked, and it was approved. It can be done, but it has its complications. Do we want to do it? Do we want to rely on these types of systems, or will we be able to engineer them? And here opens up a whole topic where we combine that analytical, deterministic, mechanistic knowledge with artificial intelligence. We make AI explain itself in mechanistic terms. We constrain AI within the limits of the physical and the real, the biological.
How do we do all this? That is, how do we combine the power of learning without knowing the real rules with the real rules we need to know? This is one of the great debates in all aspects I discuss with artificial intelligence, with people from all sectors. It doesn't matter if it's someone very corporate in a bank, a scientist, or someone more focused on the emotional aspect of AI or how we'll relate to it. But in the end, what always comes up is: will we have to demand explainability from AI, and for how long? Because, you know, if this continues to progress at the rate it's progressing, and we're seeing it go from something you couldn't trust to something that can help you, to something that even starts making certain discoveries on its own, like in mathematics, etc., then who's to say tomorrow we won't be a hindrance by trying to make it explain itself?
So, instead, the story here will be: what do you want to do? Do you want us to cure this disease without having any idea how it was done or managed, but it works because pragmatically we see the result? Or do you want us to explain it to you, and it takes twice as long to get there? I have a feeling we're going to face that decision moment, right? Of wanting to give it free rein and take advantage of the benefits, which is a bit like a gift from the gods, as they sometimes say, or do we want to demand to be in charge and in control, even if that means we can't do as much? And I think it's an interesting debate. It is indeed an interesting debate, and it's very difficult to say no to a cheap solution to a real problem.
In virtue of simply wanting to know how it works. Even morally speaking, it's difficult to say no to those who suffer from the disease. And in practice, it's impossible to say no. No, it's not going to be that easy because the number of cases you'll be able to resolve this way, and also have the capacity to intervene, it's very difficult that it won't require knowledge of intermediate aspects related to explainability. If it's just about providing the result of a situation, for example, in medicine, given this patient's trajectory, tell me what disease they will have next. There are already publications that, by training with patient trajectories, allow you to predict what disease you will have five years in advance.
It's not that inconceivable because you know a trajectory leads to a disease, but that can be done, the system is validated, and it appears to be quite stable. Of course, this gives you very few options for intervention. You can monitor it, if it's cancer or whatever, but it doesn't tell you how to tackle it. Tackling it requires more intermediate knowledge. So, it's quite possible that this explainability issue, right? It's quite possible that the explainability we get from current systems will be insufficient for interventions. And here comes the justification for why you need another type of technology to complement what generative systems, which we are using, won't be able to provide.
We also have to consider that we are using the technology we have because it's what we have. Exactly. But it doesn't have to be the appropriate technology to solve these problems. And it's quite possible that around the corner, we'll find another technology and say, "My goodness, how are we doing this?" Correct. It's probable that we'll see an idea that is perfectly explainable, and today we're with this, and it seems like the future, but it won't be. In the world of science and biology, you've been using artificial intelligence for a very long time, haven't you? Machine learning, deep learning, etcetera. I imagine you've been in this for decades, and this isn't the first time, not by a long shot with this current boom, but it does seem like there's been a before and an after. When is that moment when you notice in your daily life, you say?
We always, as you correctly say, computational biology and informatics are largely a translation of computational methods developed in computer science to biological problems. The translation time, how long it took for a new technology to be adopted, could be decades. Suddenly, frameworks like PyTorch and TensorFlow existed for years. Then, someone had the idea to use them in biology, and it produced a real transformation in biology, a scientific transformation. What's happening is that these times have been continuously shortening with the arrival of more people from computational backgrounds into biology. These times have shortened, starting with DeepMind's AlphaFold. Perhaps that's the moment it became evident.
It became evident that for the first time, a computational method would largely solve a well-defined problem in biology. This happened around 2021-2022. Could you explain, if you will, AlphaFold for those who don't know it? This is a technology developed by DeepMind, a startup that was then part of Google, right? Founded by Elon Musk, by the way, and led by Demis Hassabis, who is today the CEO of Google's AI division, right? But let's explain what AlphaFold is and what problem it tackles. For decades, we've been asking ourselves about a problem that is, in principle, a very theoretical, very abstract, very... how to put it? A very scientific problem: genes produce proteins.
That is, our 20,000 genes produce 20,000 proteins, which are, in principle, the active factors in a cell. The protein does X. These proteins are a chain of polypeptides formed by 20 different components, 20 types of different components. When we sequence a gene, we get the sequence of that protein. These proteins are like small machines that do things, and therefore, to do things, they acquire a specific three-dimensional structure. And we know this works because we can do an experiment. We take a protein that is unfolded, put it in suitable conditions, in water and a bit of something else, and it acquires its three-dimensional structure.
That is, it folds into a physically extremely attractive form, very beautiful shapes that you can represent in colors, in three dimensions, and manipulate. They are like small sculptures that move a little and also do things, transforming one compound into another. They perform all these functions; they are very attractive from an aesthetic point of view. The problem is that resolving that three-dimensional structure has experimental methods for doing it, and it's costly. Even though technologies are improving, resolving the structure of a single protein can still take years because it's complicated.
And what's the use of knowing that structure? For example, it's useful for designing drugs, for designing compounds, for designing new proteins, for knowing how they work, for explaining a mutation in the BRCA gene for breast cancer, how that mutation causes the protein to malfunction as Ras21 does when mutated in cancer. Understanding mutations in proteins requires understanding their structure. Once we understand the structure, we gain visibility and the capacity to do things around it, from developing a medication to the entire industry of drug development, biotechnology, using proteins to do things, or explaining mutations and those kinds of things. All of that relies on knowing protein sequences and structures.
That is, the first step to doing anything is understanding that three-dimensional structure. And this is something that used to take us years. Years. In fact, the first part, going from genes to proteins, is the genetic code, a Nobel Prize-winning discovery and all that, and that's a relatively easy problem. Nowadays, it's a trivial problem. You have the genome sequence, and you get the protein sequence. Broadly speaking, there are always exceptions in biology, but generally, that's how it is. However, once you have the structure, that second part, what is the structure like? Well, it was done experimentally, and we developed computational methods. We've been developing computational methods for decades, making progress, small progress, but we were, if I put it in numbers, at about 70% resolution of the problem, to give you a number. That's a way of speaking.
Starting around 2020, DeepMind began to develop a new vision based on things that had been done before, but with a new way of encoding deep neural networks, which were becoming fashionable at that time. And it produced a very important advance, but it didn't fully solve the problem. By 2022, they reconsidered the approach and changed their architecture, and they did things that, while not inconceivably complex, were very difficult to put together. This isn't a conceptual revolution of doing something very different; it's excellent engineering. And basically, they resolved the problem to 95%. That is, they achieved an AI capable of predicting the three-dimensional structure of a protein with 95% success.
And they are also able to indicate, for each piece of the protein, with what reliability it is predicted. It doesn't just tell you, "This protein is predicted with 90% reliability," but rather, "This part is at 70%, and this part is at 90%." Okay, this is a very interesting property. It's as if ChatGPT could predict for each sentence, "You can trust this one 70%, this one 90%." We can't do that with ChatGPT, but we can do it with proteins, and it's one of the few areas where it's possible.
So, suddenly, this is transformative. We can begin, and it's already in databases. For all the sequences we have, all the genomes we have, we have the proteins, and we have a predicted model with a level of reliability that is, in fact, so immense. So, for people to understand, how many proteins are there? Well, humans, in principle, have one protein for each distinct gene, but each organism has a part of its genome, which can be quite large. For instance, a plant genome can have many more proteins. So, the collection of possible proteins, of real proteins, is very large; it's millions of proteins.
But then, just two years later, in 2022, other systems are published that, instead of doing it the way DeepMind did with a neural network and iterating the process, use language models like those of ChatGPT. Instead of training with text to produce text, they train with those 20 amino acids and produce proteins. And now we can not only have proteins produced in the natural world but also any protein. Of course. Of course. So, the number is infinite. We can generate an infinite number. Of course.
I remember watching the DeepMind documentary, which I highly recommend to everyone. It was a documentary publicly available on YouTube, though I'm not sure if it's still public on YouTube. It received an astonishing 500 million views. It's a very interesting documentary, barely 45 minutes long. And I remember a scene because they filmed this entire process. I recall a scene where they are in the office, in the meeting room, where they have managed to solve the problem. The team says, "Yes, we've managed to solve the problem. Until now, when someone needed a protein prediction, they requested it, and someone would start working on it. We've enabled a sort of system where they can ask us to predict it."
And then Demis says, "But, but what are you saying? If this is so easy and so cheap, make them all." And we arrive at the point where today, DeepMind has managed to predict all proteins, and they are available. All the ones that I understand are known and are available. And Demis Hassabis said at Davos not long ago that if human scientists had done this, it would have taken them a billion years, and they did this in... It is true that they wanted to win a Nobel Prize, and they wouldn't have if they hadn't made the protein structures public. It's a noble goal, isn't it? Again, it's the real proteins, the known ones, plus all those we can invent.
The ones we can invent are very important because they allow us to ask why we only see some, why others weren't produced during evolution, what new functions might have existed or not. That is, the complement of the unreal and the real is what allows us to do things, to ask questions about why what exists exists, why what doesn't exist doesn't, how we can complement functions. So, there's a lot of science behind complementing the known with the unknown. That said, and having resolved it, first, it's a fundamental advance, and now everyone uses the public structures plus the predicted ones. Obviously, you appreciate not only having natural proteins but also wanting to make changes and create new predictions. So, AlphaFold and its derivatives and later versions are continuously used.
That said, this resolves it for a specific class of proteins, but the scientific problems surrounding protein structure go beyond that. We don't just want the structure of an isolated protein; we want the structure of a protein when it functions as a machine, along with other machines. The structure of large complexes. There are new methods for this too. AlphaFold has led to methods for this, and other groups have developed methods for this, but it's a much more difficult problem to solve. How does a protein function? I can predict its structure, I can predict multiple, how they function together. Therefore, it's good to think that it's a major change and allows us to do things we couldn't do before, but it has by no means exhausted the problems in that area.
Clearly, there are many more to resolve, aren't there? But are we in agreement if we say that AlphaFold is perhaps the most important thing artificial intelligence has done for the world? I usually say this. I say it's the only thing that is verifiable. We can verify that it has been resolved. We can check how well it has been resolved, and it has numbers that tell us how well it has been resolved. Other things, producing images, producing diagnoses, we evaluate them, but we are very unsure if they are extrapolable to another environment, if they will work with another dataset. There are always many doubts, and in this case, it's a problem that, with its limits defined, the problem is defined in a bounded way, and we can say it has been almost resolved.
Yes, I use it as an example of the biggest, and perhaps only, case where there's an undeniable advance due to artificial intelligence, which also has few negative repercussions. Or, well, it has resolved the problem, it allows us to advance, it continues to pose further questions, it can pose questions about how things work, which are also interesting to investigate, but it's a definitive advance for a student who would start their thesis before AlphaFold and after. For a student starting their thesis when AlphaFold comes out, everything they will do will be completely different. In fact, the framing of their thesis will be different. So, yes, it's transformative, it doesn't have many negative aspects we can consider about its impact, and it's a very powerful tool.
I hope you are enjoying this podcast. As you know, Infojobs is our sponsor and is offering something that will surely interest you. Can you imagine getting paid €1000 to come to one of my AI podcasts? Well, that's exactly what we've prepared with Infojobs. If you're passionate about AI, spend your days trying new tools, watching videos about artificial intelligence, or are simply curious to understand where the future is heading, this opportunity is for you. We are looking for someone who wants to come to our studio, experience a recording from the inside, meet the team, and spend some time with me talking about artificial intelligence, technology, and everything that's changing in the world. And yes, [music] Infojobs will pay you for it. What does this Call Jobs include? Well, €1000 net for a day of experience, VIP access to an AI podcast recording. Plus, you can bring anyone you want. A meet-and-greet with me to talk, chat, take photos, and ask me anything you want.
Additionally, round-trip travel is included from anywhere in Spain, as is accommodation. And importantly, although the offer appears in Madrid, you can apply from anywhere in Spain because we will cover the travel and accommodation. If you love artificial intelligence and want to have a great experience, apply now through the link in the description, the pinned comment, or the QR code on screen. You can also access Infojobs from the website or the app, type "John Hernández" in the search bar, and the job offer will appear. Don't miss this opportunity, and we'll see you here on the podcast. What do you think about the fact that they received a Nobel Prize? Because I saw you comment earlier that, well, they wanted the Nobel Prize, and therefore they did it in a certain way, because I understand it gives DeepMind validation.
Something happened here. Yes, something happened here: the sociology of these things, Nobel Prizes, and such, they mask a lot, right? And the varied anecdotes and so on. The first Nobel, as you know, is divided into two parts: one for DeepMind and another for David Baker. Demis for the AlphaFold part, and David Baker, who is still in this field and a prominent scientist, for engineering in these types of areas. From the community's perspective, it was somewhat frustrating. Communities are always a bit frustrated with Nobel Prizes because they didn't acknowledge enough that it was possible thanks to the collections of protein structures, meaning crystallographers and databases that had accumulated and annotated and maintained all the protein structures, and without that, they couldn't have done anything. Of course, the training, right? In the end, they were trained on what had already been done by humans over decades, right? Exactly. A unique dataset for training. That's why companies...
Not very large, by the way, right? The training dataset, unlike what we tend to think about generative AI, where they train on a massive amount of data and derive emergent capabilities from it. In the case of AlphaFold, from what I saw in the documentary, it was precisely a small dataset. One of the characteristics of having done something that has done a very complex, very abstract job very well, despite having very limited training because we didn't have that many, but it was high-quality data that covered a large space. The structures were surrounded by many sequences that yielded the same structure.
Therefore, you didn't just learn from that protein with that structure but from all other similar proteins that you know will have that structure. So, there were thousands of structures and millions of sequences. So, it wasn't a dataset as large as others, but much better and much more distributed in space, covering space much better. So, proteins, for entirely unexpected reasons, have become a wonderful dataset for training artificial intelligence mechanisms. One could never imagine, when working in proteins, which was a minority field, that major companies would work on proteins. It's absolutely unthinkable. No one could have imagined something like this, right?
So, they developed the method. They didn't give much credit. Now, whenever DeepMind talks about structures and so on, I don't know. But initially, very little recognition for the community, and for the community itself. They demonstrated their method at CASP, which is a protein structure evaluation competition held every two years for many years, where the entire community contributed as evaluators and predictors, and developed the ideas on which AlphaFold was later based. Of course, initially, there was no mention of this, which was very offensive to the community. At the point where they thought they had to make the method accessible, because if they kept it as an application and didn't make the method and data accessible, they would be very poorly regarded. The method being accessible ultimately wasn't that important because within a few months, it could be reproduced. Other groups reproducing it wasn't a big problem. The capacity to do all the structures isn't just theirs. The database that maintains the structures at the European Molecular Biology Laboratory is the database where the structures are. So, it's a joint effort between a European public laboratory and DeepMind to maintain that database of structures.
In any case, they published the second edition of AlphaFold and didn't make it public. They didn't make the software public until the paper was published in Nature without revealing the code. The reviewers protested because they couldn't review a paper if they couldn't access the code for approval. This created a great scandal, and in the end, they released the code, published the paper despite everything, and, seeing that they wouldn't get the Nobel Prize, they ended up making the software of the second version of AlphaFold public as well. So, this is our interpretation from the community, knowing what happened with the paper and all that. We don't know what the committee decided. I mean, it's great that they made it public, and it's appreciated, but they were always reluctant until it became as beautiful as it seems. It's not as beautiful as it seems.
I'm sorry not to be more positive about this, but it has served them very well. But I think it reflects something that is very true in the world of AI: that ultimately, these companies are large corporations pursuing their interests, and sometimes they look very good in the media for things they're doing, but we shouldn't forget that these are absolutely orchestrated image campaigns. So, in the end, DeepMind has benefited greatly from having that image because now they are seen as the most scientific laboratory, the most... The Nobel Prize has done more for DeepMind than DeepMind for the Nobel Prize. This means that, in the end, this is nothing more than a public authority campaign that allows people like me, who have no clue about the subject, to see it and think, "Wow, this must be important because they've received a Nobel Prize." Therefore, these people know what they're doing, and when it comes to choosing between using ChatGPT or Gemini, it's very likely that will weigh into my decision, and therefore, it's a value for them.
So, it's great that those of you who are involved in this say, "Gentlemen, this isn't as rosy as you paint it." It started as a very independent company doing very adventurous things, playing Go, playing X, playing Y, and it's very interesting. He is a very interesting person, and his way of seeing all this is very interesting, and his progression, why they've done certain things and others, is super interesting. But in the end, they've become so visible and public that they've been practically completely absorbed by Google, and they're now facing other obligations and other types of...
What's your opinion of Hassabis, because after all, this type of person is presented to us as a genius? He is [clears throat] a genius. I believe, I honestly believe, having spoken with him many times, that he is a genius, without a doubt. I think so, he's a super smart guy. Inspired. I believe he is inspired, at least up to what he's doing now at Google, I don't know, but until then, he was inspired by the desire to understand how the human brain works. And since he wasn't inclined to dissect human brains and found it too complicated to make progress that way, he dedicated himself to how computational brains work. And his entire progression was to face problems that couldn't be solved by calculation but could be approached by induction, by strategy, by other means.
The origin of AlphaFold, at least as he tells it, is a game called Folding at Home, in which you were given a structure of a protein on your home computer and had to produce the three-dimensional structure, to fold it. [exhales] Surprisingly, people at home, without any knowledge of physics or calculation tools, started folding it and achieving very realistic folds. Because it was intuitive in some way. Also, no one could fold it, but they did. So, he thought there was an element of intuition, and therefore, after playing a strategy game, they played an intuition game. That's why they entered the protein problem. That's what he explains. I think it's true and I find it a fascinating idea. Absolutely fascinating, isn't it? Moving away from Go to delve into all of this.
Hey, speaking of AlphaFold, which undoubtedly has an incredible impact and everything that has come from it, which I think is tremendous. More recently, AlphaGenome has been published. What's the difference? Well, training with proteins, as we said, there are millions of proteins, the structures are well-determined, it covers a space consistently, so it's a great training dataset. Genomes are a terrible training dataset; they are enormously large. Small variations in any letter of the code can have consequences. We don't have that many genomes whose consequences we know. It's a very difficult training dataset.
And what you want to predict isn't something like a very defined structure that you can evaluate exactly. You want to predict consequences for diseases, which is a much less defined theory. So, it's a much more difficult point, and predicting the consequences of mutations is much harder than with these tools to predict protein structure. Therefore, what we can get from predicting the consequences of mutations is much less useful and works much less well because the problem is less defined, you quickly get into things you've never seen and are outside the prediction range, and you can't predict them. So, it's less useful, less applicable, and a more difficult problem.
And do you think that if they solve it? Because AlphaFold, what DeepMind presented, maybe it was a bit of hype, but they said it was as transformational as AlphaFold itself. So, do you think there's less chance of success for it to be useful to the scientific community? Yes, because for starters, our diseases depend very little on our genome. There are cases where a specific mutation causes a disease, like a rare disease, you find a single mutation that causes the phenotype of the disease. But generally, a disease is a very complex set of mutations, of changes, and not everything...
And the genetic part of a disease is only a small percentage, although external factors, the environment, development, and such also play a role. A lung disease is very determined by how the child's lungs developed. That's not in the child's genome, although there is a part that does, so it's only a proportion. Therefore, knowing the genomes is not enough to explain all diseases, far from it. This is where the disappointment comes from when genomes were sequenced and we thought we would find solutions.
And why are they trying to create AlphaGenome? What's the objective? Well, because the objective is more ambitious. The objective is not just to use genome information but to try to decode all genomic regulation. That is, ultimately, from our early embryonic development, you have an embryo with one cell initially, and then it expands to more cells where there are initial factors, and everything will happen, everything is encoded there, the entire development will be encoded there. So, the idea is that we can find the rules of biology and should be able to reproduce the development process, the disease process, the rest of the processes.
Do we have enough information to do this computationally solely with genome sequences? No, we cannot reproduce development solely with genome sequences because we don't know all the intrinsic signals involved. The signals within the sequence. Our genome is a sequence that contains signals. This activates that, this activates that, this does something else, plus initial components that set it all in motion. It's a very complex problem. Are we capable of decoding all these signals? In the case of proteins, are we capable of decoding, understanding the relationships between things so they fold? In the case of the genome, are we capable of finding all the signals and how they correlate to result in such a complex process as development? Not yet. And we lack elements because there are things we don't know. Can we understand all these things in a machine learning system without knowing why they happen? Can we find all the correlations between everything? Well, it shouldn't be impossible because there's no limit saying it's impossible, but right now, we don't have enough information to do it.
And foundational models are starting to emerge, trying to explain part of that process. I can explain how genes will be expressed under these conditions. Given these conditions, this genome, and these initial conditions, which genes will be expressed? In which tissue? We have foundational models that are starting to be able to do those things, but I think complete development is a very complex process; we are still a long way from achieving it. Having the sequences, the series of information, and predicting whether this will lead to a disease or not can be useful in a clinical context. It has some clinical utility, but it's not a tool that, by itself, changes medicine.
Because, of course, I imagine, looking at the evolution that has occurred with AlphaFold, that if we were in 2019, this also seemed impossible, right? [laughter] That is, it's true that sometimes it feels like, as far as you can see, this couldn't happen, and then suddenly it did. And suddenly, in just a few years, this has changed so radically that it's a complete turnaround. Do you think that due to the exponential advance of artificial intelligence, it could go faster than it seemingly appears, or how are you experiencing that exponential advance? Because, you know, sometimes people I talk to say, "Well, it hasn't changed that much." And others say, "Well, I look every day, and it's completely different." I understand it depends on the field, right? For now, but in your case...
I wouldn't know how to answer that very well. I think no. I mean, I believe that complex problems like development and diseases won't be easily solved with current technology. "Easily" means there won't be a paper tomorrow saying we've solved it, because there's a lack of data, a lack of things. It's not just a problem of having better AI; it's that you won't have enough information to create the papers. At least for the first three years, and beyond that, perhaps there will be a technological change that is so transformative, or the accumulation of data might happen faster than we're seeing. But that's at least what I think. I understand there might be specific areas where a very big change can occur within biology, where suddenly a very big change can happen, but globally, the fundamental problems, I don't think so.
So, that statement by Demis Hassabis that in the next ten years we will cure all diseases, how do you see it? Well, as he is the vice president of Google, I don't know. I think first, curing all diseases would mean having biological and chemical processes to intervene in all diseases. We have drugs for all diseases, and it's not that easy to generate drugs that are adequate, not toxic, that can be demonstrated in a population, and so on. Right now, we have a large battery of... Cancer treatment has improved enormously in many areas. There are still cancers where we don't, or much less so, and we don't have drugs for everything, even if we have very fine diagnoses, we don't have the drug.
So, drug development is complex, even if there's artificial intelligence, even if we can synthesize new drugs, the process of testing them and seeing if they work takes time. So, due to the dynamics of things, we don't have enough knowledge, nor enough instruments, and we have a regulated process. The culmination of these three things makes it very difficult to think you'll have the cure for all diseases in such a short time. That said, also, the cure for all diseases for whom and where? Because this is a second, very interesting, added problem.
Yes, yes. Let's see, who has access to this? Because, you know, I understand that there are different ways to treat diseases. One is to prevent them from appearing, another is to cure them with medication. Which of the two paths do you think brings us closer? Because we're thinking about developing drugs, but perhaps if we manage to prevent the next generation from getting cancer, the need for medication will end, right? Of course, this seems logically evident, doesn't it? If we can prevent diseases, we can stop smoking, for example, which would be a good way to prevent many diseases. We can reduce city pollution; that would be a good way to reduce diseases, indeed. And being able to make not just general recommendations but specific recommendations is a good way to ward off diseases.
The other way, curing them, is more complicated because things happen that are then difficult to resolve through chemical methods. We're talking about drug development and its limitations. Surgery has advanced a lot, but it also has its limitations regarding who can be operated on and who cannot. So, obviously, prevention would be better. The idea of learning with artificial intelligence on patient trajectories, that is, what diseases you've had, what medications you've taken, but that also influences what environments you've been in. And based on this information, which is like your life trajectory, being able to predict what will happen to you next. There are already papers on this, discussing diseases and predicting what will happen next. These are very data-driven papers and other things, and they work well. The results that have been checked with other trials and so on seem very reasonable, and more papers will come out from Google and similar entities in this sense of predicting disease trajectories.
This is very important because if you have one, it's not the same to say, "You should exercise more because you're going to have X," as to say, "Your probability of having this disease within five years is Y." This makes it much easier to foresee, take measures, follow up, and eventually develop treatments. Therefore, there is a lot of potential in studying medical records, accumulated medical records. Of course, what we are seeing, and it's a fact, is that in the United States, they announced a long time ago that they would halve the time to market for a drug, mainly due to all these benefits that artificial intelligence is providing, which is tremendous. But, of course, there is a series of bureaucracy, especially that part of drug testing, like phase one, phase two, phase three, where there's a period that is relatively non-negotiable, right?
It's relatively... Well, there are also two things: one is to reduce the time before human trials, clinical trials. The first drug approved for human testing without animal experiments took two years, right? What happened? Ana approved the paper first, then it was already in phase two trials. Last year it was in phase two trials, so this is already happening. And more and more drugs being produced incorporate a part of this preclinical part, which are models, methods, predictions, computational things. We work, like other groups, on replacing animals used in experiments with synthetic animals because you save money, you avoid animal sacrifice, you go faster. Then you get to clinical trials, and obviously, in clinical trials, there are increasingly novel models attempting to reduce experimentation times, use existing data, retrospective data, use smaller numbers, combine, open up the branches of the trial so you don't have to follow just two branches, but adapt to what's happening. A lot of things, and especially in this context, both the preclinical and clinical parts play a fundamental role. All this part that has a lot to do with generative AI, with creating synthetic data, synthetic animal data, synthetic human data that can be used to complement real data.
Of course. And here, undoubtedly, comes the part of the digital twin, right? That famous digital twin, which is supposed to allow us to do tests digitally without having to do them on a human specifically, or on the individual concerned, right? Because I understand there will also be digital twins of animals. Let's delve into this because I think the topic of digital twins is something people are accustomed to hearing. It's something that... I always think there are like two versions of the digital twin, right? There's a lot of talk in the corporate world, at the company level, etcetera, about what a digital twin is within the company. But the initial point of all this comes from the biological part, from the science part, where we try to create a simulator, right? Of someone's body or something. And from there, we can do tests on how what we're doing will affect that organism. The Digital Twin is this, right?
Yes. When I'm asked what we do at the Barcelona Supercomputing Center, the easiest way to explain it is: we make digital twins, right? Because it's true. Our engineering colleagues make digital twins of physical systems: the digital twin of a rotor, the digital twin of a wind turbine blade. It's a system that represents, that simulates the behavior of the physical system in the computer using equations and simulations. In the purest spirit of the term "digital twin," the physical system constantly sends data to the simulated system, which adapts to this new data and reproduces it. This works well. There are commercial digital twins, and you can buy... You mentioned the corporate realm; you can ask Siemens to make a digital twin of your factory, and they'll make a digital twin of it.
This is done to, for example, optimize that digital twin and those things, those types of things, right? Which are amazing, but you understand it, don't you? They can map every function, what it does, where it does it, how it does it, and make the entire system function with a simulation. Our case is more complicated because, obviously, digitalizing a factory is relatively easy; digitalizing a human is impossible, and we return to the topic of data. We don't have all the data for everything, and today, a digital twin of a human organism functioning at the same level as a factory's digital twin, meaning all elements, knowing how all elements operate, is impossible. We don't even know how many cell types we have, so we don't even know how many different cell types.
What we do, as is obvious, is try to have digital twins of parts of the system. One way to see it is atomic parts. These proteins we talked about earlier, you can think of each of these structures acting as a digital twin that simulates how that protein behaves, and that's what you do by doing molecular dynamics to generate new drugs or analyze new drugs. At the other extreme, you can make digital twins of an organ at a certain level of representation, a certain level of resolution, a lung where you see how air enters through the airways, how it expands, how it contracts, how oxygen reaches the alveoli, how it distributes. Of course, this is very useful. You can personalize it to a specific person and consider how a virus entering will affect a specific person's lung configuration, how it will reach different parts. It could be very different from how it will affect another person.
That is, the beauty of the digital twin here is not to create a digital twin of a lung, but of my lung. I hope you are enjoying the podcast. You know there's a tool I use very often, which is Plot. This device allows me to take a recorder wherever I go and have an interaction with another person. It allows me not only to record those conversations with a single button but also to have an interactive conversation with those conversations afterward. This is truly useful because when we have in-person meetings, we're always playing the game of "who said what and when." From that, we find ourselves in a situation where we don't have the certainty, like with an email, of the information we've captured. Having this allows us to have that interactivity with those meetings afterward, and at any time, I can access the Plot app and ask about that conversation I had. But not only that, but I also have a sort of ChatGPT with the full contextual memory of all those meetings we've had in the past.
This allows me, for example, if I need to create a budget, to ask Plot. If I need to write an email to a client, I can ask Plot because Plot is part of all those interactions we have between humans. The reality is that it's a super useful tool and fits in your pocket, so for me, it's a no-brainer. It's like a tool that costs very little to carry, and at the moment you have an interaction, you put it to record, you ask the person's permission, and from then on, you have it recorded to refer to at any time. If you want to summarize your in-person meetings, all that has to do with the physical world, Plot is the best option. You have a link in the description for both the app and the new Plot Pro. And if you want the old one, it's still available, and the Plot PIN, that smaller one you can wear as a necklace or a watch, is also available. You have the links in the description.
Or at least many types of lungs simultaneously that you can adjust to yours. So, this is the kind of thing that's being done and has potential practical utility in medicine. I'm going to study this person's heart, or this prototype of a person's heart, to which I'll apply arrhythmia data, and I'll see how it reacts to arrhythmia surgery, surgery in a specific region, before performing the surgery. I can plan the surgery better with a digital twin. We are interested in digital twins that are somewhere between these two parts, between the atomic part and the part of organs or systems, which have to do with cell behavior. What happens inside a cell, or what happens in the interaction between cells? In a tumor. How do the cells within the tumor behave? Each of them is a different cell type: the immune system too, normal cells, tumor cells, the distribution, and how they react to a drug. Can I reproduce the behavior of a tumor on the computer? Well, no, because a tumor is very complex. Can I have a sufficient simplification to predict the reaction of a drug? Well, that's what we're doing, and we're starting to get positive results when we make predictions and then do experiments.
So, basically, for me to understand, the ultimate goal of the digital twin is to have a copy of me on a computer, specifically me, not any human, but John specifically. And from there, scientists, doctors, can experiment with that digital twin without having to do things to me, so that when they find the key, they can use it. Many years ago, we had a slide of a girl who is in the real world and has her data, her medical history stored. She has a watch that takes data, and then she has an accident. So, there's a copy of the girl in the virtual world where all that data is. When the accident happens, the new accident data arrives, and the system very quickly prepares what the solutions are. In this case, the best intervention is this, or the second best is this, and it gives these recommendations for monasticism for a disease, that type of thing, which then applies to the patient.
Then they can be applied to the patient. They make recommendations to the real patient. Again, it depends on the level. At a very general level, this would be like having medical information and creating a digital twin of your medical information. This is about predicting future diseases. At a deeper level, we can look at things about organs. For an intervention for a lung infection, I take the X-rays of that person's lung, personalize it, see where viruses might have reached, what would be better to do, better to intervene to collapse the lung or not. These would be interventions at the organ or system level. At a deeper level, you look at what's happening at the level of... In the lung example, I don't just look at how viruses reach a part of the lung, but how the lung cells of this specific person, who has a mutation, react and where they will react or not react. I am simulating at a deeper level. Even deeper, I have all the proteins and genes and their interactions. These systems with many layers are very complex, and for now, we are at the level of resolving individual layers and are not capable of integrating them much. So, to have that complete image on the computer and for it to work for all types of medicine, we won't see it this year or in the coming years, but we will start to see systems that at least in the experimental part allow us to tackle specific problems.
What things do you see artificial intelligence making possible that you didn't think you'd see? Proteins, we've already discussed that; I find them absolutely incredible. The first models for extracting information from medical reports seem incredible. A medical report written by doctors, containing incredible amounts of information that you have to read to understand, if you understand it at all, and now we have systems capable of extracting all that information. This is incredible. We've worked on this topic for many years, and the ability to detect entities like disease names, patient professions, symptoms, and drugs has been limited. It's extremely simple with LLMs, I don't understand. Well, relatively simple, right? There's a very particular language, they need retraining, they have work, but it can be done, and this seems absolutely transformative because this is humanity's greatest experiment: accumulating all that information. So, this seems absolutely incredible. The first systems we have for... these genomes we talked about, predicting how they are expressed, what is expressed, in which tissue it is expressed, how it is expressed in different cell types, is something that was unthinkable. We didn't even have the data six or seven years ago to start doing it. We have the data, and we're starting to have systems that are entering that system.
This part about digital twins, combined or not with artificial intelligence, is fascinating. It's what fascinates me the most right now, and being able to say, based on a person's genomic data, what their response to a drug will be through an AI-powered model is fascinating. All the part where we generate synthetic data and use it not only to have more data and be able to train a system, but we generate synthetic data to explain reality. This is truly science fiction. That is, we have genomic information from children with a disease, with childhood cancer, and in some cases, we can't differentiate if they are one group or two groups. And now we use synthetic data to try to explain this reality, and with synthetic data, suddenly we flood a biology that we wouldn't understand, and synthetic data. This was also science fiction; we couldn't generate synthetic data, we couldn't do any of this. So, these are varied topics where progress is incredible and was totally unexpected. And now, in the last two years, we have the whole story of agents; agent systems are like total madness. Everyone has agent systems running to do all sorts of things. But of course, for us, they represent not only a potential time saving and so on but a different way of understanding processes that we used to do, but perhaps not by putting them in the context of agents, we are starting to understand what they meant or how they worked differently. It's not just the potential to do more, or better, or faster, or more, whatever, but the potential to decouple systems that were previously very coupled, allowing us to understand them differently. So, this is also unthinkable. I can't imagine five years ago that our lab would be creating agents to do things. And yet, here we are.
We've seen Hinton, who is a physicist and not closely related to biology, but he made that famous statement: "In the next ten years, we will advance more in science than in the last hundred." Do you think the trend is towards that, towards multiplying because I understand that in recent years, there has been tremendous progress in science? If we go back to the early 1900s, perhaps progress wasn't so fast. So, it's not so far-fetched to think we'll progress much faster, but that much?
Well, what does that mean? Will we solve the basic problems of science in the next ten years? Basic scientific problems, not curing diseases, but the origin of life, how life originates and how the first cells originate, protein functions, how new ones emerge, and how... Of course, we have larger, more complex proteins, but the first ones in a bacterium, how did the first proteins emerge? How does a system develop? How did the complexity of cellular systems arise? Development, how development begins, how development evolves, how development functions. In essence, we have more instruments to do it and can go back. Tony Gabaldón published a paper a few days ago that goes further back in time to the origin of cells, [exhales] because we have new models, because we have computational capacity, because we can do these kinds of things. But still, the frontier is very far from how these origins are.
So, I believe that the really interesting problems, scientifically deep scientific problems, will continue to be difficult to develop, and it will take time. Will we do it faster than ten years ago or twenty years ago? My guess is possibly yes. Perhaps we can answer some of them, if not completely, then approximate some of them, but I think we underestimate what it represents. Of course, scientists will always defend themselves, always say they are indispensable. I mean, it's hard for me to imagine that there isn't a part that remains very difficult to replace with the technology we know, because all this about digital twins, all this about treating diseases, involves something we are beginning to see and that I believe could have a very big impact in changing how we understand science, medicine, and biology in general in the coming years, which is precisely that part of personalized medicine we are seeing, right? Knowing the individual better could allow us to develop specific drugs. Right now, I have an allergy, and I take a tablet just like everyone else, right? But we might find that for me, a tablet isn't the solution for my allergy because it has different components. And so, we're developing personalized medications for individuals, but that requires understanding the individual, which is precisely what artificial intelligence can give us. This has already begun, right? And in fact, there are already certain cancer treatments that are personalized.
Where is the world now, and where are we heading? Thanks to AI in the area of personalized treatment, the idea of trying to personalize diagnoses and treatments, that is, not giving everyone the same diagnosis, but a diagnosis that is more suited to your progression, what will happen to you specifically, in general terms, and making possible a treatment. It's this idea of moving from a very general classification of diseases to increasingly specific classifications, dividing each disease into smaller and smaller groups, and ultimately doing it in a more personalized way. And obviously, this is happening. We are increasingly seeing more personalized diagnoses, with more personalized treatment plans, but we have a long way to go.
I believe that in this, we clearly have a long way to go. On the one hand, we need personal information that is more treatable. Let's get back to having patient trajectories, predicting what will happen next. This has a lot to do with being able to have those personal trajectories in a computational system that helps us know specifically what will happen next. And the second part is having special treatments for each person. Obviously, we won't have a drug for each person; we will have a drug for each group of people, which will become smaller and smaller over time. Synthesizing drugs and so on. We return to the previous discussion. While technically it's becoming easier and easier to explore the enormous space of molecules, finding a suitable molecule for that protein, and we know the process and so on, and synthesizing them, etc. Then you have to go through the entire regulatory and approval process. What will be the limiting factor? The limiting factor will be the price. Personalized treatments are obviously much more expensive. For an industry that produces aspirin, it's very convenient because it works for many people, and it's a single formula. If you had to make one for each person, it would be unaffordable.
So, we will face a problem: we need more information and better systems to have personalized trajectories. The capacity to design and synthesize drugs. We are getting more, but it's not infinite. And obtaining molecules against the RAS21 protein, which is a fundamental cancer target, has taken decades, tens of years, and only now are molecules starting to become more effective. So, it's not that all key molecules, not all key proteins, are easy to obtain, nor do we know all the processes. And once you achieve it, getting it approved and making it economically viable is a whole other story, isn't it? So, while it's a laudable goal, it's a somewhat utopian goal to have a drug for every disease, for every person. We will have drugs that will be increasingly suitable for groups of people, at best.
Of course, in the end, what I'm getting from this conversation is that AI is damn good, it's absolutely incredible, but be careful, because the problem is more complex than it seems, right? That perhaps AI won't solve it, and this is something we're finding in many other fields, right? We see that, for example, no matter how much AI can fix, say, the electoral process, ultimately, until there's an electoral cycle where someone pushes for it, four years pass, you encounter human limitations, bottlenecks, so even if you find the solution with AI, it might not be that easy to apply it. I have a feeling that in this... this is exactly it. When we reach medical application, the system is highly regulated, partly for good reasons, right? Because it has to do with human health and so on. And for other reasons, not so good, because there are also many bureaucratic layers that aren't so interesting.
With that, you enter a process that is still more regulated than many other industries, which makes the capacity to adopt new technologies slower. If you walk around hospitals, you'll see a lot of developments. They develop a new test to detect X cancer or whatever, it makes the newspaper, and so on. Expanding that system so that it is actually used beyond the hospital's use is extremely difficult and happens very slowly and in few places. Of course, you're right, because we saw this, for example, in Charles Darwin University in Australia, right? Where they did that thing where with a photo of a biopsy, it told you if that type of cancer was present or not with 99% accuracy, and those doctors achieved 79% the night before having this. But we haven't followed up on the case if that came out of Charles Darwin University, right? If that is now being used worldwide because it's no use diagnosing much better if it's not in all hospitals. The approval process for a medical device, and software as a medical device, is very expensive and complex. So, many of these things don't go anywhere.
But this is the less positive part, right? Even though there are super interesting developments, it's difficult for them to be implemented in a short time, and then not all of them will be implemented. In fact, few are implemented. The positive part of all this is that suddenly you can make diagnoses and so on accessible to a population with very few resources otherwise. We were talking earlier about Africa, for example, where you find places with no doctors, no specialists, and suddenly you might find a device on your phone that is very useful. So, depending on your level, it can be more or less useful, right? Totally, right? Totally. I think people aren't aware of this, right? That often they say, "No, but it's better to have, in education, for example, a human teacher who uses the guide and so on." I say, "Yes, yes, but in the middle of Cambodia, there might not be a human teacher, you know? So, it's better to have ChatGPT through a crappy 3G on a mobile phone, and while that might seem insufficient for your standards, for the standard of that poor kid in rural Cambodia, this is... information for patients or information for the doctors themselves in medicine, because it's incredible, or in science. Everyone, I always say that if AI is doing nothing else, at least it's going to make a part of humanity that has been on the sidelines of the conversation able to enter the conversation because all of this has this potential to allow entry. We will discover new researchers, new doctors who will emerge from places where they couldn't have emerged in the past due to education, for example, right?
Speaking of which, about the future, how do you see the world? I mean, what scenario do you think we will face in 10, 20 years? Well, [laughter] I think if there's no major disaster, and today, no one can say there won't be a major disaster, right? In the finer points. We will see that we have overcome this stage. We will see that there is another, different technology, more powerful than language models and generative AI, with fewer explainability problems, with a way of integrating knowledge: acquired knowledge, case-based knowledge, probabilistic knowledge, mechanistic knowledge in our case, so that we have systems we can trust more. Moreover, they will be integrated systems, not just "I ask ChatGPT, and so on," but something that incorporates the entire chain: "I ask, there's a model, it answers me, it validates it, it tells me how it's desirable, it tells me how it should be implemented." And we will see that systems have adapted to use it in a way that it's part of the system. If, by then, I had a lab, it would be... I think there will still be students, engineers, and postdocs, but there will also be AI systems, and we will work together with them, seeking solutions, and it will be more of a dialogue with colleagues who will give opinions, and we will think, and together we will reach solutions. Perhaps less responsibility solely on the human, but something we can take more collectively. But they will be systems that are different from what we have now, that will be more reliable, and will function with a logic that will have more to do with human reasoning chains, than what is functioning now, which doesn't have to do with human reasoning chains. That's what I would like to happen in our...
Of course. And do you consider that we will be able to keep pace for a long time? Because there's a big debate right now between the term AGI, right? Artificial General Intelligence, where it can do productive work, which is somewhat what you're proposing, right? One that can participate at the human level. But lately, we're talking more and more about ASI, Artificial Superintelligence, right? That artificial intelligence that far surpasses human capacity and is a bit like... I don't know who said this, I think it was TechCrunch, one of those who said it's like if we were to go back to the year 1800 in a time machine and try to explain the concept of an air conditioner to a good citizen of 1800. You say, "Well, the physics is there, but they'll say it's witchcraft, right?" So, to what extent will we be able to keep pace with AI, and won't we have to say, "Well, if that's the case, give me the solution, but I'll accept it," right? Like today, for example, I understand biology; I have no clue; I accept the medicine they give me or accept the treatment because people like you have done that work, you're much smarter than me, but I don't intend to understand what you've done. Do you understand? I mean, I don't know if I have the mental capacity to do it, but I assume there will come a point where if AI continues to improve and humans continue on our line, because ultimately, one thing humans have is a very poor capacity to improve our cognitive abilities. That is, we have evolved over the last thousands of years; we are smarter than our ancestors were, but not that much, right? I mean, it hasn't been that crazy. So, if there's a trend where it's getting smarter and smarter, and we're on a trend of getting a little smarter each time, there will come a point where we won't be able to keep pace.
So, I have a feeling that all that part of wanting it to be explained, which makes a lot of sense from a human perspective, is something we will have to give up at some point, right? We will have to give up control, so to speak, or limit it, meaning either we give up control to obtain greater benefits, or we say, "Okay, this is the limit; I don't want you to go beyond this because I stop understanding you." Again, I think there isn't a single answer to that super interesting problem. I think it will depend a lot on the type of application and the environment we are talking about. I can imagine environments where that is true, where you have sufficiently tested the systems, you don't have total validation even with current models, and there will be better things in the future, but even with current models, you haven't been able to test it exhaustively for all cases, but you've set some boundary conditions. You say, "Look, if you don't go outside these boundary conditions, I'll believe it." You will say, because we've always done that.
And there will be environments where you will believe it and won't try to understand how it does it because why bother if it does it very well, and you don't need to know why. And there will be other environments where neither the answer is black and white, nor is the data as clear and reliable, where it will continue to be a research topic, and you will need it to be explainable because otherwise, you won't be able to interact with them. We will continue to open up parts that are black boxes because it will predict X, and I don't want to know how, but for the specific problem, you will still need it to be explainable because otherwise, you won't be able to interact with them. In a problem where it won't find the solution alone. We must consider that we are talking about problems where everything happens through things that can be put into written text. As long as we can keep the problems in written text, perhaps it's true that they can solve all problems, and we don't need to know how. As soon as we start having problems that begin to have other dimensions, these systems are not so good at doing it. We have to wait for another type of solution. No, no, we can't even say right now that apart from solving problems of a certain type, they have been very good at solving problems of another type. Mm.
So, I hold out hope that for the problems that interest me, it will remain necessary to have adequate technology for explainability, because otherwise, we won't be able to interact with them, and we won't be able to solve the problems. I don't think they are the solution for all problems.
Hey, well, I hope you can solve many more problems and that we can advance much further. I don't know if there's anything in the AI field that you have great hope for in the short term, something that could surprise as much as AlphaFold. Is there anything you see in your work, or are we currently in a much more experimental and slower advance phase? I believe that what is going to happen most rapidly now, what will emerge in the coming months, will be systems, these systems for predicting diseases. What's happening is that there are already some publications that are good, but they need to be applied to certain datasets and so on, which will allow us not only to say what disease you will have based on your record, but to start thinking backward: this is because you had this disease, or because you had these conditions, or because you took this medication, or because you have this genetic composition. That is, prediction systems will allow us to make predictions, which is good because you can do prevention and so on, but they will also allow us to go backward, to the causes. I think you will see unprecedented change in this, because we have never had this amount of information available from medical reports, nor these prediction systems, and it's something transformative. Suddenly, it's a door to understanding diseases that we haven't had before.
Well, let's hope so. Alfonso, thank you very much for coming to shed some light and coherence on all the hype that is sometimes sold to us. Because, from my perspective, I think it's necessary that we also adopt a dose of realism and realize that this is powerful, but it's not black magic, right? This is something we also need to see. Well, if we have managed, at least, to shed a minimum of light in this fascinating, so changing, and so difficult-to-predict environment. Delighted. Thank you very much for the conversation.
Thank you for listening to this Podhoc podcast.
