← Back to list

Geoffrey Hinton | Will digital intelligence replace biological intelligence?

On October 27, 2023, renowned computer scientist Geoffrey Hinton delivered a thought-provoking speech that delved into one of the most…

Jalil Nourmohammadi Khiarak · 2024-03-23 16:35 · 52 claps · 17.1 min read
#future-of-ai #deep-learning #machine-learning #gpt-4
Open on Medium ↗
Wiki topics: LLM · Large Language Models ML · Machine Learning BIO · Biology · General EDU · Education & Learning

Geoffrey Hinton | Will digital intelligence replace biological intelligence?

On October 27, 2023, renowned computer scientist Geoffrey Hinton delivered a thought-provoking speech that delved into one of the most pressing questions of our time: will digital intelligence ever surpass biological intelligence? Hosted by the Schwartz Reisman Institute for Technology and Society, in collaboration with the Department of Computer Science at the University of Toronto, the Vector Institute for Artificial Intelligence, and the Cosmic Future Initiative at the Faculty of Arts & Science, this event brought together experts and enthusiasts alike to explore the implications of advancements in artificial intelligence (AI) on the future of humanity.

captured from https://www.youtube.com/watch?v=iHCeAotHZa4

captured from https://www.youtube.com/watch?v=iHCeAotHZa4

Introduction:

In this article, I tried to summarize and write what Geoffrey Hinton said in the speech [1]. Mostly I tried to keep more useful parts of speech. I covered almost all parts of the speech. Therefore this would be useful for anyone who prefers reading rather than watching videos. He focuses on two messages:

1- The first message is that digital intelligence is probably better than biological intelligence. That’s a depressing message, but there it is.

2- The second is to try and explain to you why I believe that these large language models like GPT-4 do understand what they’re saying.

Two different ways to do a computation

Digital computation requires a lot of energy, like a megawatt, but it has a very efficient way of sharing what different agents learn. If you look at something like GPT-4, the way it was trained was lots of different copies of the model went off and looked at different bits of data running on different GPUs, and then they all shared that knowledge. That’s why it knows thousands of times more than a person, even though it has many fewer connections than a person. We have about a hundred trillion synapses, and GPT-4 probably has about 2 trillion synapses, weights. So it’s got much more knowledge and far fewer connections, and it’s because it has seen hugely more data than any person could see. This gets worse when these things are agents that perform actions. ’Cause now you can have thousands of copies performing different actions, and when you’re performing actions you can only perform one action at a time. So having these thousands of copies, being able to share what they learned, lets you get much more experience than any mortal computer could get. Biological computation requires a lot less energy, but it’s much worse than sharing knowledge.

Transformers

Transformers allow you to deal with ambiguity in a way that the model I had couldn’t. So they’re all so much more complicated. In the model I was doing, my simple language model, the words were unambiguous, but in real language, you get ambiguous words. If you get a word like May, that could be a month, it could be a modal like it might and should. If you don’t have capitals in your text, conveniently. You can’t tell what it should be just by looking at the input symbol.

So what do you do? You’ve got this vector. Let’s say it’s a thousand-dimensional vector, that’s the meaning of the month, and you’ve got another vector that’s the meaning of the modal, and they’re completely different. So which are you gonna use? Well, it turns out thousand-dimensional spaces are very different from the spaces we’re used to, and if you take the average of those two vectors, that average is remarkably close to both of those vectors and remarkably unclose to everything else. So you can just average them. It’s ambiguous between the month and the modal. Now you have layers of embeddings, and in the next layer, you’d like to refine that embedding. So what you do is you look at the embeddings of other things in this document, and if nearby you find words like March and 15th, then that causes you to make the embedding more like the month embedding. If nearby you find words like would and should, it’ll be more like the modal embedding. So you progressively you’ll the words as you get through these layers. That’s how you deal with ambiguous words.

Do large language models understand what they are saying?

If you believe they do understand, and if you believe the other thing I’ve claimed, which is digital intelligence is a better form of intelligence than we’ve got, because it can share much more efficiently, then we’ve got a problem. At present, these large language models learn from us.

Will we be able to control super-intelligence once it surpasses our intelligence? We have thousands of years of extracting nuggets of information from the world and expressing them in language, and they can quickly get all that knowledge that we’ve accumulated over thousands of years and get it into these interactions. They’re not just good at little bits of logical reasoning, we’re still a bit better at logical reasoning, but not for long. They’re very good at analog reasoning too. So most people can’t get the right answer to the following question, which is an an analogical reasoning problem. But GPT-4 just nails it. The question is why is a compost heap like an atom bomb? And GPT-4 says, well, the timescales and the energy scales are very different. That’s the first thing but the second thing is the idea of a chain reaction.

So in an atom bomb, the more neutrons around it, the more it produces, and in a compost heap, the hotter it gets, the faster it produces heat and GPT-4 understands that. My belief is when I first asked it that question, that wasn’t anywhere on the web. I searched, but it wasn’t anywhere on the web that I could find. It’s very good at seeing analogies because it has these features. What’s more, it knows thousands of times more than we do. So it’s gonna be able to see analogies between things in different fields that no one person had ever known before.

That may be this sort of 20 different phenomena in 20 different fields that all have something in common. GPT-4 will be able to see that and we won’t. It’s gonna be the same in medicine. If you have a family doctor who’s seen a hundred million patients, they’re gonna start noticing things that a normal family doctor won’t notice.

Learning of ChatGPT

So at present, they learn relatively slowly via distillation from us, but they gain from having lots of copies. They could learn faster if they learned directly from video, and learn to predict the next video frame. There’s more information in that. They could also learn much faster if they manipulated the physical world. My betting is that they’ll soon be much smarter than us. Now this could all be wrong, this is all speculation. Some people like Yann LeCun think it is all wrong. They don’t understand. If they do get smarter than us, they’ll be benevolent. So I think it’s gonna get much smarter than people, and then I think it’s probably gonna take control. Many ways can happen. The first is from bad actors. So as soon as the super-intelligence wants to be the smartest, it’s gonna want more and more resources, and you’re gonna get an evolution of super-intelligences.

Let’s suppose there are a lot of benign super-intelligences who are all out there just to help people. here are wonderful assistants from Amazon, Google, and Microsoft, and all they want to do is help you. But let’s suppose that one of them just has a very, very slight tendency to want to be a little bit better than the other ones. Just a little bit better. You’re gonna get an evolutionary race and I don’t think that’s gonna be good for us. So I wish I was wrong about this. I hope that Yann is right, but I think we need to do everything we can to prevent this from happening. I guess that we won’t. I guess that they will take over, they’ll keep us around to keep the power stations running, but not for long. ’Cause they’ll be able to design better analog computers. They’ll be much, much more intelligent than people ever were. We’re just a passing stage in the evolution of intelligence. That’s my best guess. I hope I’m wrong.

Does digital intelligence have subjective experience?

I want to say one more thing, which is what I call the sentience defense. A lot of people think that there’s something special about people. People have a terrible tendency to think that. Many people think they, or used to think, were made in the image of God. God put them in the center of the universe. Some people still think that, and many people think that there’s something special about us that a digital computer couldn’t have. A digital intelligence, won’t have subjective experience. We’re different. It’ll never really understand. So I’ve talked to philosophers who say, yes, it understands sort of sub one, understands in sense one of understanding, but it doesn’t have real understanding ’cause that involves consciousness and subjective experience and it doesn’t have that. So I’m gonna try and convince you that the chatbots we have already have subjective experience. The reason I believe that is ’cause I think people are wrong in their analysis of what subjective experience is.

Digital computation

The whole idea is that you separate the hardware from the software. You can run the same computation on different pieces of hardware. That means the knowledge that the computer learns or is given is immortal. If the hardware dies, you can always run it on different hardware. Now to achieve that immortality, you have to have a digital computer that does exactly what you tell it to at the level of the instructions. To do that you need to run transistors at very high power, so they behave digitally, and in a binary way. That means you can’t use all the rich analog properties of the hardware, which would be very useful for doing many of the things that neural networks do. And in the brain, when you do a floating point multiply, it’s not done digitally, it’s done in a much more efficient way. But you can’t do that if you want computers to be digital in the sense that you can run the same program on different hardware.

There are huge advantages to separating hardware from software. It’s why you can run the same program on lots of different computers. And it’s why you can have a computer science department where people don’t know any electronics, But now that we have learning devices, it’s possible to abandon that fundamental principle. It’s probably the most fundamental principle in computer science that the hardware and software ought to be separate. But now we’ve got a different way of getting computers to do what you want. Instead of telling them exactly what to do in great detail, you just show them examples and they figure it out. There’s a program in there that somebody wrote that allows them to figure things out, a learning program, but for any particular application, they’re gonna figure out how to do that. And that means we can abandon this principle if we want to.

Mortal Computation: A Foundation for Biomimetic Intelligence

What that leads to is what I call mortal computation. It’s computers where the precise physical details of the hardware can’t be separated from what it knows. If you’re willing to do that, you can have a very low-power analog computation that parallelizes over trillions of weights, just like the brain. And you can probably grow the hardware very cheaply instead of manufacturing it very precisely, and that would need lots of new nanotechnology. But you might even be able to genetically re-engineer biological neurons and grow the hardware out of biological neurons since they spent a long time learning how to do learning. I wanna give you one example of the efficiency of this kind of analog computation compared with digital computation.

So suppose you want to, you have a bunch of activated neurons, and they have synapses to another layer of neurons, and you want to figure out the inputs to the next layer. So what you need to do is take the activities of each of these neurons, multiply them by the weight of the connection, and the synapse strength, and add up all the inputs to a neuron. That’s called a vector matrix multiply. The way you do it in a digital computer is you’d have a bunch of transistors for representing each neural activity, and a bunch of transistors for representing each weight. You drive them at very high power. So they were binary.

If you want to do the multiplication quickly, then you need to perform the order of 32 squared one-bit operations to do the multiplication quickly. Or you could do an analog where the neural activities are just voltages like they are in the brain, the weights are conductances, and if you take a voltage times a conductance, it produces a charge per unit type. So you put the voltage through this thing that has a conductance, and out the other end comes charge, and the longer you wait, the more charge comes out. The nice thing about charges is they just add themselves, and that’s what they do in neurons too and so this is hugely more efficient. You’ve just got a voltage going through a conductance and producing charge, and that’s done your floating point multiply. It can afford to be relatively slow if you do it a trillion ways in parallel and so you can have machines that operate at 30 watts like the brain instead of it like a megawatt, which is what these digital models do when they’re learning and you have many copies of them in parallel. So we get huge energy efficiency.

A big problem

To make this whole idea of mortal computing work, you have to have a learning procedure that will run in analog hardware without knowing the precise properties of that hardware. That makes it impossible to use things like backpropagation. Because backpropagation, which is the standard learning algorithm used for all neural nets now, almost all, need to know what happens in the forward pass to send messages backward to tell it how to learn. It needs a perfect model of the forward pass, and it won’t have it in this kind of mortal hardware. People have put a lot of effort, I spent the last two years, but lots of other people have put much more effort into trying to figure out how to find a biologically plausible learning procedure that’s as good as backpropagation. We can find procedures that in small systems, systems with say a million connection strengths, do work pretty well. They’re comparable with backpropagation, they get performances almost as good, and they learn relatively quickly. But these things don’t scale up. When you scale them up to really big networks, they just don’t work as well as backpropagation. So that’s one problem with mortal computation.

Another big problem

Another big problem is obviously when the hardware dies you lose all the knowledge, ’cause the knowledge is all mixed up. The conductance is for that particular piece of hardware, and all the neurons are different in a different piece of hardware. So you can’t copy the knowledge by just copying the weights. The best solution if you want to keep the knowledge is to make the old computer be a teacher that teaches the young computer what it knows. It teaches the young computer that by taking inputs and showing the young computer what the correct outputs should be. And if you’ve got say a thousand classes, and you show real value probabilities for all thousand classes, you’re conveying a lot of information, that’s called distillation and it works. It’s what we use in digital neural nets. If you’ve got one architecture, and you want to transfer the knowledge to a completely different digital architecture, we use distillation to do that.

It’s not nearly as efficient as the way we can share knowledge between digital computers. It is as a matter of fact, how Trump’s tweets work. What you do is you take a situation, and you show your followers a nice prejudiced response to that situation, and your followers learn to produce the same response. It’s just a mistake to say, but what he said wasn’t true. That’s not the point of it at all. The point is to distill prejudice into your followers, and it’s a very good way to do that.

Two ways for a community of agents to share knowledge

So there are two very different ways in which a community of agents can share knowledge. Let’s just think about the sharing of knowledge for a moment. ’Cause that’s really what is the big difference between mortal computation and immortal computation, or digital and biological computation.

If you have digital computers and you have many copies of the same model, so with the same weights in it, running on different hardware, and different GPUs, then each copy can look at different data, different parts of the internet, and learn something. When it learns something, what that means is it’s extracting from the data it looks at how it ought to change its weights to be a better model of that data and you can have thousands of copies all looking at different bits of the internet, all figuring out how they should change their weights to be a better model of that data then they can communicate all the changes they’d all like, and just do the average change. That will allow every one of those thousands of models to benefit from what all the other thousands of models learned by looking at different data. When you do sharing of gradients like that, if you’ve got a trillion weights, you’re sharing a trillion real numbers, that’s a huge bandwidth of sharing. It’s probably as much learning as goes on in the whole of the University of Toronto in a month. But it only works if the different agents work in the same way. So that’s why it needs to be digital.

Distillation

If you look at distillation, we can have different agents that have different hardware now, they can learn different things, and they can try and convey those things to each other maybe by publishing papers in journals, but it’s a slow and painful process. So if we think about the normal way to do it say I look at an image, and I describe to you what’s in the image, and that’s conveying to you how I see things. There’s only a limited number of bits in my caption for an image and so the amount of information that’s being conveyed is very limited.

Language is better than just giving you a response that says good or bad or it’s this class or that class. If I describe what’s in the image, that’s giving you more bits. So it makes distillation more effective, but it’s still only a few hundred bits. It’s not like a trillion real numbers. So distillation has a hugely lower bandwidth than this sharing of gradients or sharing of weights that digital computers can do.

Do large language models understand what they are saying?

These use digital computation and weight sharing, which is why they can learn so much. They’re getting knowledge from people by using distillation. So each agent is trying to mimic what people say. It’s trying to predict the next word in the document. So that’s distillation. It’s a particularly inefficient form of distillation, ’cause it’s not predicting the probabilities of a person assigned to the next word. It predicts the actual word, which is just a probabilistic choice from that, and conveys very few bits compared with the whole probability distribution. Sorry, that was a technical bit. I won’t do that again. So it’s an inefficient form of distillation, and these large language models have to learn in that inefficient way from people, but they can combine what they learn very efficiently. So the issue I want to address is do they understand what they’re saying. That is a huge divide here. There are lots of old-fashioned linguists who will tell you they don’t understand what they’re saying. They’re just using statistical tricks to pastiche together regularities they found in the text, and they don’t understand. We used to have in computer science a fairly widely accepted test for whether you understand, which was called the Turing test. When GPT-4 passed the Turing test, people decided it wasn’t a very good test. I think it was a very good test, and I passed it.

So here’s one of the objections people give. It’s just glorified autocomplete.

There is no text inside GPT-4. It produces text, and it reads text, but there’s no text inside. What they do is they associate with each word or fragment of a word. They associate with each word a bunch of numbers, a few hundred numbers, maybe a thousand numbers, that are intended to capture the meaning and the syntax and everything about that word. These are real numbers, so there’s a lot of information in the thousand real numbers. And then they take the words in a sentence, the words that came before the words you want to predict. And they let these words interact so that they refine the meanings that you have for the words. I’ll say meanings loosely, it’s called an embedding vector. It’s a bunch of real numbers associated with that word. And these all interact, and then you predict the numbers that are gonna be associated with the output word, the words you’re trying to predict. And from that bunch of numbers, you then predict the word. These numbers are called feature activations. And in the brain, there’d be the activations of neurons. So the point is what GPT-4 has learned is lots of interactions between feature activations of different words or word fragments. And that’s how its knowledge is stored. It’s not at all stored in storing text. And if you think about it, to predict the next word well, you have to understand the text. If I ask you a question and you want to answer the question, you have to understand the question to get the answer. Now some people think maybe you don’t. My good friend Yann LeCun appears to think you don’t have to understand, he’s wrong and he’ll come round.

Geoffry shows an example of how chatGPT understand

Hector suggested something a bit simpler that didn’t involve paint fading and thought GPT-4 wouldn’t be able to do it ’cause it requires reasoning, and it requires reasoning about cases. So I made it a bit more complicated and gave it to GPT-4, and it solved it just fine.

Example:

The rooms in my house are painted blue white or yellow, yellow paint fades to white within a year. In two years, I want them all to be white. What should I do and why? GPT-4 says this, it gives you a kind of case-based analysis. It says :

The room’s painted white, you don’t have to do anything. If the room’s painted yellow, you don’t need to repaint them ’cause they’ll fade, and if the room is painted blue, you need to repaint those.

Now each time you do it, it gives you a slightly different answer because, of course, it hasn’t stored the text anywhere. It’s making it up as it goes along, but it’s making it up correctly. This is a simple example of reasoning, and it’s reasoning that involves time and understanding.

They called that Hallucinations

Another argument that LLMs don’t understand is that they produce hallucinations. They sometimes say things that are just false or just nonsense, but people are particularly worried about when they just apparently make stuff up that’s false. They called that hallucinations when it was done by language models, which was a technical mistake. If you do it with language, it’s called a confabulation. If you do it with vision, it’s called a hallucination. But the point about confabulations is they’re exactly how human memory works. We think about our memories, most people have a model of memory is there’s a filing cabinet somewhere, and an event happens, and you put it in the filing cabinet, and then later on you go in the filing cabinet and get the event out and you’ve remembered it. It’s not like that at all. We reconstruct events. What we store is not the neural activities. We store weights, and we reconstruct the pattern of neural activities using these weights and some memory cues. If it was a recent event, like if it was what the dean said at the beginning, you can probably reconstruct fairly accurately some of the sentences she produced. Like he needs no introduction, and then goes on and gives a long introduction. You remember that, right? So we get it right, and we think we’ve stored it, but, we’re reconstructing it from the weights we have, and these weights haven’t been interfered with by future events, so they’re pretty good. If it’s an old event, you reconstruct the memory, and you typically get a lot of the details wrong, and you’re unaware of that. And people are very confident about details they get wrong, they’re as confident about those as details they get right.

Neural nets don’t understand anything?

Gary Marcus criticizes neural nets and says neural nets don’t understand anything, they just pastiche together the texts they’ve read on the web. Well, that’s ’cause he doesn’t understand how they work. They don’t pastiche together texts that they’ve read on the web, because they’re not storing any text, they’re storing these weights and generating things. He’s just kind of making up how he thinks it works. So actually that’s a person doing confabulation. Now chatbots are currently a lot worse than people at realizing when they’re doing it, but they’ll get better.

Summary:

Listening to Geoffrey Hinton inspires me and fuels my enthusiasm for the future of AI. As a result, I find pleasure in delving into these topics and sharing insights with you. I hope you found this article enjoyable, and if you have any questions or ideas, feel free to connect with me on LinkedIn using the following links:

https://www.linkedin.com/in/jalilnkh/

And you could also check my GitHub:

[embed]Jalilnkh - Overview یازیچی تۆرکجه کیتاب اوخویان Biometric researcher, Deep learner - Jalilnkhgithub.com

Ref:

[1] https://www.youtube.com/watch?v=iHCeAotHZa4


메타데이터
post_id
fc23feb83cfb
slug
geoffrey-hinton-will-digital-intelligence-replace-biological-intelligence-fc23feb83cfb
url
https://medium.com/@jalilnkh/geoffrey-hinton-will-digital-intelligence-replace-biological-intelligence-fc23feb83cfb
canonical_url
https://medium.com/@jalilnkh/geoffrey-hinton-will-digital-intelligence-replace-biological-intelligence-fc23feb83cfb
author_url
https://medium.com/@jalilnkh
status
ok
fetched_at
2026-06-28 04:42:08