02 / The engine inside modern AI
What is a transformer?
The “T” in ChatGPT stands for transformer. It isn’t a machine or a chip. It’s a design: a recipe for how to arrange the math inside an AI program. Almost every big AI model today (Claude, Gemini, GPT, Llama) is built from this recipe. Here’s how it works in five steps, no math required.
The whole idea in one picture
First, clear up the name
Is a transformer
a machine?
No. There’s no transformer box, chip or robot. A transformer is a method: a specific set of math steps, written as code, for turning text into a prediction.
Think of long division. Long division isn’t a calculator; it’s a method. You can do it with a pencil or a calculator can do it for you, but the method is the same.
The transformer is the method. Computer chips are what carry it out. And a model like GPT or Claude is what you get after running that method on trillions of words: a giant file of learned numbers that follows the transformer’s steps every time you ask it something.
Three things people mix up
So why is it called a “transformer”?
Because it transforms
It takes a list of word-numbers in and, layer by layer, transforms each one into a better-informed version that carries meaning from the words around it. It was first built for translation: transform an English sentence into a German one. It has nothing to do with the robot toys or the gray boxes on power lines.
Why it matters that it’s a method
Anyone can use it
Because the transformer is just published math, every lab could build its own version. That’s why OpenAI, Google, Anthropic, Meta and Chinese labs all make competing models from the same basic recipe. What differs is their data, their computing power and their training choices.
2017 · “Attention Is All You Need”
The idea that
changed everything
In 2017, eight researchers at Google published a paper with a catchy title: “Attention Is All You Need.” They were trying to improve language translation.
Older AI programs worked through a sentence one word at a time, in order, like reading through a straw. Each step had to wait for the one before it, and by the end of a long sentence the start was mostly forgotten.
When people say a transformer “looks at every word at once,” nothing is actually looking. It means the program does one big calculation that compares every word with every other word in the same moment, instead of in a line. That turned out to be more accurate, and much faster on chips that can do thousands of multiplications side by side.
Faster training meant you could feed it far more text. That turned out to be the key that unlocked today’s AI.
Old way vs. transformer
What “looking” really is: a table of numbers
The scoreboard behind every word
For “The dog chased the ball,” the program fills in a table. Each row is a word, each column is a word, and each box gets a score for how much those two words relate. “Chased” scores high with “dog” (who chased?) and “ball” (chased what?).
That’s the whole trick. An old-style program filled the table row by row, in order. A transformer is designed so a chip can fill all 25 boxes at the same time, like 25 students each solving one problem at once instead of one student solving 25 in a row.
Real models do this for thousands of words, hundreds of times over, in a fraction of a second.
Step 1 · Tokens
Chop text into pieces
Computers only understand numbers. So the first job is to split text into small chunks called tokens (whole words, or parts of long words) and give each one an ID number.
Try it · type a sentence
A simplified tokenizer for illustration: it chops up every long word. Real ones learn their chunks from data, so they keep most common words whole and split rarer ones; a typical English word is about 1.3 tokens. Big models read trillions of tokens during training.
Step 2 · Embeddings
Turn words into places on a map
Each token becomes a list of numbers, like map coordinates. Words with similar meanings end up close together. Real models use thousands of dimensions; this map shows just two.
The famous trick: the direction from man → woman is about the same as king → queen. Meaning becomes geometry.
A meaning map (simplified)
Step 3 · Attention (the big idea)
Every word asks:
who matters to me?
Read this sentence. What does “it” mean? You know instantly, because of the last word. Attention lets the model make the same connection. Switch the last word and watch the arrows move.
Real models run dozens of these attention “heads” side by side, each looking for a different kind of connection (grammar, who did what, rhymes, facts) and repeat it in layer after layer. This example comes from Google’s 2017 translation model, which reads the whole sentence at once. Chatbot models read left to right, so for them the link is made when the model reaches “tired” or “wide” and looks back at “animal” or “street”.
Step 4 · Stack the layers
Repeat it dozens of times
One attention layer plus a “thinking” network (the bigger of the two parts) makes one transformer block. Big models stack dozens of these blocks, and some of the largest stack more than 100 (Meta’s Llama 3.1 405B has 126).
Early layers notice simple things like grammar. Middle layers track meaning and facts. Late layers plan what to say next. Nobody programs these jobs; they emerge during training.
All of the model’s knowledge lives in its parameters: the adjustable numbers inside these layers. Today’s biggest models are thought to have trillions of them.
Inside the stack
Step 5 · Predict the next token
Guess the next word.
Then do it again.
At the top of the stack, the model scores every possible next token. It picks one, adds it to the text, and runs the whole thing again. A long answer is thousands of these guesses in a row.
A setting called temperature controls how often it takes a less-likely word. Higher = more creative, and more likely to wander off.
Watch it write
Illustrative probabilities, not from a real model. Only the top 5 choices are shown, so they add up to less than 100%; the rest is spread over thousands of other tokens.
From predictor to assistant
How does that
make an AI?
A raw next-word predictor is like a brilliant parrot that has read the whole library. Turning it into a helpful assistant takes four more stages of training.
01 · Pretraining
Read everything
Predict the next word across trillions of words of books, websites and code. Takes months on tens of thousands of chips.
Learns: language, facts, patterns
02 · Instruction tuning
Learn to help
Train on examples of good questions and answers so it responds like an assistant instead of just continuing text.
Learns: follow instructions
03 · Feedback
Learn what people want
People (and AI judges) rate answers. The model is nudged toward helpful, honest, safe ones. Often called RLHF.
Learns: helpful & safe behavior
04 · Reasoning training
Learn to think it through
Practice on problems with checkable answers (math, code) and get rewarded for getting them right, so it learns to work step by step.
Since 2024: o1, R1, Claude, Gemini…
05 · Tools & agents
Learn to act
Give it tools (web search, code, a computer) and train it to use them over long tasks.
2025–26: AI agents
Why bigger worked
Scaling laws
In 2020 researchers found that adding more data, more parameters and more computing power made transformers better in a smooth, predictable way. That discovery set off the race to build giant data centers.
A real weakness
Hallucinations
The model is built to produce likely-sounding text, not to check facts. When it doesn’t know, it can still produce a confident, fluent, wrong answer. Always verify claims that matter.
Beyond words
Not just text
The same design works on anything you can chop into tokens: image patches, sound, video frames, DNA, robot movements. That’s why one architecture now powers chatbots, image generators and more.
Put curiosity to work
Read it. Try it. Question it.
Explore new AI research, or take a paper-first investigation into your classroom.