World Models Next Wave of AI? What Are Investors Actually Buying for $3.5 Billion?
The purpose of this article is not to tell you what to think about the latest billion-dollar AI bet. It is to make you think about it. Most…
World Models Next Wave of AI? What Are Investors Actually Buying for $3.5 Billion?

Is it Really Real Intelligence?
The purpose of this article is not to tell you what to think about the latest billion-dollar AI bet. It is to make you think about it. Most of what follows is publicly available, but the implications are rarely spelled out in plain English. Read it, push back on it, work out where you land.
In March 2026, Yann LeCun’s new startup AMI Labs raised $1.03 billion at a $3.5 billion pre-money valuation. It is the largest seed round in European history. The company has no product and no revenue. What it has is a thesis: large language models like ChatGPT are a dead end, and the path forward runs through something called “world models.”
That pitch is being narrated everywhere as a fundamental break from the ChatGPT paradigm. It is worth slowing down and asking how fundamental the break actually is.
The pitch
The story investors are buying goes like this. ChatGPT-style systems are trained on text. They learn to predict the next word in a sequence. This is why they hallucinate, why they confidently invent citations, why they cannot really reason about physical cause and effect. They have read about the world but they have not understood it.
World models, the pitch continues, fix this. Instead of training on text, you train on video and physical interaction. The system learns to predict not the next word but the next state of the world. It develops something closer to a real understanding of how things work.
A billion dollars says this is the future.
What has actually changed
Underneath the marketing, almost nothing about the machinery is different. It is still the same kind of neural network, trained the same basic way (gradient descent), running on the same kind of chips made by the same companies. The architecture is transformer-adjacent. The math is the math.
The only thing that has changed is what you feed it.
That is like claiming you have invented a new kind of car because you switched from gasoline to diesel. The engine is identical. The fuel is different. Whether that difference matters depends on whether the engine was the problem in the first place.
The clever bit
The technical move that makes JEPA (Joint Embedding Predictive Architecture, the approach behind AMI Labs) feel different is this. The model does not try to predict the actual world directly. It predicts a compressed summary of the world, called a “latent representation,” and does its reasoning inside that summary. The promise is that this is more efficient and more abstract, closer to how humans seem to think.
This sounds like a real advance, and in a narrow technical sense it is. But it depends on the compression being faithful to what it is compressing.
A quick word on loss functions
The rest of this hinges on a concept worth pausing on. When you train a neural network, you need some way to tell it whether it is doing well or badly so it can adjust itself. The loss function is that scorecard. It takes the model’s prediction, compares it to the right answer, and produces a number representing how wrong the prediction was. Training is just the process of nudging the model’s billions of internal settings to make that number smaller and smaller.
The catch is that the model only cares about what the loss function measures. If your scorecard is checking the wrong thing, the model will get very good at that wrong thing and you will never know, because by every metric you are tracking it is improving.
Hold that thought.
The hole in the floor
A recent academic survey (“Critiques of World Models,” arXiv 2507.05169) essentially proves a clean theorem against JEPA’s central claim. For the latent-space trick to work, the compression process has to be effectively reversible. You have to be able to squeeze the world down into the summary and then unsqueeze it back out without losing anything important. In real systems, that never actually happens. Some information always gets lost or distorted in the squeeze.
What that means in practice is that the model’s internal summary of the world and the actual world can quietly drift apart. The training process will not catch this, because the loss function is only checking whether the summary is consistent with itself, not whether it still matches reality. You end up with a system that is getting “better” by every metric the trainers care about, while its picture of the world is slowly going sideways. It looks confident. It is wrong.
This is the same fundamental problem ChatGPT has. The hallucinations, the confident nonsense, the absence of any real grip on what is actually true. JEPA does not solve this problem. It moves it one layer deeper, into a space that is harder to inspect, and gives it a more sophisticated name.
What is being bought
So here is the question worth sitting with. The pitch is “we fixed the problem with LLMs.” The actual technical move looks more like “we relocated the problem to somewhere harder to see and put a more impressive label on it.” The wall that ChatGPT-style systems are hitting is probably the same wall waiting for video-trained systems. The marketing has not caught up yet, but it will.
A billion dollars at $3.5 billion pre-money is a lot of money to spend finding that out.
This is not a prediction that AMI Labs will fail. LeCun is one of the most important figures in modern AI, and his team is real. They will publish interesting papers. They will produce impressive demos. Some of the work will be genuinely useful in domains like robotics and industrial automation where the distribution of problems is narrow enough that the consistency gap matters less.
But the framing being sold to investors, and through them to the public, is that this is a paradigm break. That framing is doing a lot of work. The substrate has not changed. The training method has not changed. The hardware has not changed. What has changed is the data and the loss function. If those changes turn out to be sufficient to get to general intelligence, that is a significant scientific result, but it is also evidence that the substrate-neutral assumption was right all along, which is the assumption LeCun has spent years arguing against.
The interesting question is not whether world models are better than LLMs at certain tasks. They probably are. The interesting question is whether they are different enough, in the right ways, to escape the deeper problem. The honest answer right now is that nobody knows, and a billion dollars is being spent to find out.
That is worth thinking about.
메타데이터
- post_id
- 554fcdc5126c
- slug
- world-models-next-wave-of-ai-what-are-investors-actually-buying-for-3-5-billion-554fcdc5126c
- url
- https://medium.com/@Gbgrow/world-models-next-wave-of-ai-what-are-investors-actually-buying-for-3-5-billion-554fcdc5126c
- canonical_url
- https://medium.com/@Gbgrow/world-models-next-wave-of-ai-what-are-investors-actually-buying-for-3-5-billion-554fcdc5126c
- author_url
- https://medium.com/@Gbgrow
- status
- ok
- fetched_at
- 2026-06-10 12:26:30