05 / What’s next
What comes next?
Today’s AI learned about the world mostly by reading. The next wave is learning the way you did as a baby: by watching, moving and predicting what happens next. Here are the new kinds of AI, and the big questions they raise.
Where the frontier is heading
Not just one kind of AI
The model
family tree
“AI” covers several different kinds of models. Most new products combine a few of them.
Language models
Predict the next word. Great at writing, summarizing, translating and coding.
ChatGPT, Claude, Gemini, Llama
Reasoning models
Write out a long chain of thought before answering, checking and fixing their own work. Much better at math, science and code.
OpenAI o-series, DeepSeek-R1, “thinking” modes
Agents
Models that take actions: click, type, search, run code, and keep going for hours toward a goal.
Coding agents, computer-use agents
Generative media
Create images, video, music and voices, often by starting from random noise and “cleaning” it into a picture (diffusion).
Image, video and music generators
World models
Learn how the world works by predicting what happens next in video and 3D. They can generate whole worlds you can walk around in.
Genie 3, World Labs Marble, NVIDIA Cosmos
Embodied AI (robots)
Models that see, understand an instruction and move a robot body. Often trained in simulated worlds first.
Humanoid and warehouse robots, self-driving cars
New designs beyond the transformer
Some researchers think next-word prediction alone won’t reach human-level intelligence. Yann LeCun, a deep-learning pioneer, left Meta to start AMI Labs, betting on JEPA: models that predict the meaning of what comes next instead of every pixel or word. Others explore memory that lasts, learning continuously, and much more efficient “brain-like” chips.
Deep dive
What is a
world model?
You can catch a ball without doing physics homework because your brain predicts where it will be. A world model gives AI that same kind of “inner simulator.”
A language model predicts…
the next word
It has read millions of descriptions of balls, but it has never seen one bounce. Its “understanding” of physics comes secondhand, from text.
A world model predicts…
the next moment
Trained on huge amounts of video (and the actions that caused it), it learns gravity, momentum and cause-and-effect directly. Add a “what if I press left?” and it can simulate the result.
Google DeepMind
Genie 3
Aug 2025. Turns a text prompt into a world you can walk through in real time (720p, 24 frames a second), drawing it frame by frame like a video rather than building a 3D model. Worlds stay consistent for a few minutes, and its visual memory reaches back about a minute. Opened to some subscribers as Project Genie in Jan 2026.
World Labs
Marble
Nov 2025. From the startup founded by Fei-Fei Li (of ImageNet). Builds downloadable 3D worlds from text, photos or video that game engines can use.
NVIDIA
Cosmos
World models made for “physical AI.” Robot and self-driving-car companies use them to generate realistic practice scenarios instead of crashing real cars.
AMI Labs
JEPA
Yann LeCun’s Paris-based startup (founded Dec 2025, about $1 billion in seed funding). Aims for AI that understands physics, remembers and plans, by predicting the meaning of what happens next.
Why it matters: a robot can practice a million times in an imagined world for every one try in the real one. Many researchers see world models as a missing piece between today’s chatbots and machines that can safely act in our homes, roads and labs.
Sources: Google DeepMind (Genie 3) · Project Genie · World Labs (Marble) · NVIDIA Cosmos · AMI Labs funding · World models (overview)
The road ahead
The further out,
the foggier
These are directions labs say they are working toward, not promises. The further ahead, the less anyone really knows.
Next 1–2 years
AI coworkers
- Agents that handle projects lasting days, not hours
- Assistants that remember you and work across all your apps
- Strong models small enough to run on a phone or laptop
Next 3–5 years
AI in the physical world
- Robots trained in world models doing real jobs in warehouses and homes
- AI that runs its own experiments and speeds up science and medicine
- Possibly the first systems most experts would call AGI
Further out
Beyond human?
- Machines that out-think the best humans in most fields
- Or a slowdown, if scaling hits limits in data, energy or new ideas
- Either way, the choices people make now will shape it
Big open questions
What should
we decide?
The technology isn’t the only thing that matters. These questions don’t have settled answers yet, and your generation will help decide them.
Safety
Can we keep it on our side?
How do we make sure powerful AI does what we intend, can’t be misused for things like cyberattacks, and gets tested before release?
Work
What happens to jobs?
AI may take over some tasks, change most jobs and create new ones. Who benefits, and how do people retrain?
Energy
Who pays the power bill?
Giant data centers need as much electricity as cities. How do we power AI without harming the climate or raising local bills?
Power
Who controls it?
A handful of companies and countries build the frontier models. Should the most powerful systems be open, limited or regulated?
Truth
What’s real?
When anyone can generate a realistic photo, voice or video, how do we know what to trust?
You
Skills that last
Asking sharp questions, checking sources, understanding how AI works and where it fails, and knowing a subject deeply enough to spot mistakes.
Bring this to class →Put curiosity to work
Read it. Try it. Question it.
Explore new AI research, or take a paper-first investigation into your classroom.