What is a transformer?

arqiv✳

02 / The engine inside modern AI

What is a transformer?

The “T” in ChatGPT stands for transformer. It isn’t a machine or a chip. It’s a design: a recipe for how to arrange the math inside an AI program. Almost every big AI model today (Claude, Gemini, GPT, Llama) is built from this recipe. Here’s how it works in five steps, no math required.

The whole idea in one picture

Text goes in, gets split into tokens, passes through stacked attention layers, and a next-word prediction comes out “The cat sat on the …”INPUT ThecatsatontheTOKENS ATTENTION + THINKING LAYER 1 LAYER 2 … LAYER 100+ next word → “mat” (41%)PREDICTION
A transformer is a very large next-word predictor. Everything else is built on top of that one skill.

First, clear up the name

Is a transformer
a machine?

No. There’s no transformer box, chip or robot. A transformer is a method: a specific set of math steps, written as code, for turning text into a prediction.

Think of long division. Long division isn’t a calculator; it’s a method. You can do it with a pencil or a calculator can do it for you, but the method is the same.

The transformer is the method. Computer chips are what carry it out. And a model like GPT or Claude is what you get after running that method on trillions of words: a giant file of learned numbers that follows the transformer’s steps every time you ask it something.

Three things people mix up

Three layers: the hardware is the kitchen, the transformer is the recipe, and the trained model like GPT or Claude is the finished dish THE TRAINED MODEL · THE FINISHED DISH GPT, Claude, Gemini, Llama The recipe plus billions of learned numbers. A file, not an object. 0.21-1.7 THE TRANSFORMER · THE RECIPE A method (math steps, written as code) Published in a 2017 research paper anyone can read. Instructions, not an object. THE HARDWARE · THE KITCHEN AI chips (GPUs, TPUs) in data centers The only physical part. They do the math, about a thousand trillion times per second, per chip.
“Transformer” names the recipe, not the computer and not the finished model.

So why is it called a “transformer”?

Because it transforms

It takes a list of word-numbers in and, layer by layer, transforms each one into a better-informed version that carries meaning from the words around it. It was first built for translation: transform an English sentence into a German one. It has nothing to do with the robot toys or the gray boxes on power lines.

Why it matters that it’s a method

Anyone can use it

Because the transformer is just published math, every lab could build its own version. That’s why OpenAI, Google, Anthropic, Meta and Chinese labs all make competing models from the same basic recipe. What differs is their data, their computing power and their training choices.

2017 · “Attention Is All You Need”

The idea that
changed everything

In 2017, eight researchers at Google published a paper with a catchy title: “Attention Is All You Need.” They were trying to improve language translation.

Older AI programs worked through a sentence one word at a time, in order, like reading through a straw. Each step had to wait for the one before it, and by the end of a long sentence the start was mostly forgotten.

When people say a transformer “looks at every word at once,” nothing is actually looking. It means the program does one big calculation that compares every word with every other word in the same moment, instead of in a line. That turned out to be more accurate, and much faster on chips that can do thousands of multiplications side by side.

Faster training meant you could feed it far more text. That turned out to be the key that unlocked today’s AI.

Old way vs. transformer

Comparison: a recurrent network processes words one after another; a transformer processes all words at once with connections between every pair BEFORE 2017 · ONE WORD AT A TIME Thedogchasedtheball step 1 → step 2 → step 3 … slow, forgetful TRANSFORMER · ALL PAIRS IN ONE CALCULATION Thedogchasedtheball every word is scored against every other word

What “looking” really is: a table of numbers

The scoreboard behind every word

For “The dog chased the ball,” the program fills in a table. Each row is a word, each column is a word, and each box gets a score for how much those two words relate. “Chased” scores high with “dog” (who chased?) and “ball” (chased what?).

That’s the whole trick. An old-style program filled the table row by row, in order. A transformer is designed so a chip can fill all 25 boxes at the same time, like 25 students each solving one problem at once instead of one student solving 25 in a row.

Real models do this for thousands of words, hundreds of times over, in a fraction of a second.

A 5 by 5 table of attention scores for the sentence The dog chased the ball; chased has high scores with dog and ball
Illustrative scores. Each row shows how much that word pays attention to every word in the sentence, and each row adds up to 1.00. Brighter green = more attention.
1

Step 1 · Tokens

Chop text into pieces

Computers only understand numbers. So the first job is to split text into small chunks called tokens (whole words, or parts of long words) and give each one an ID number.

Try it · type a sentence

A simplified tokenizer for illustration: it chops up every long word. Real ones learn their chunks from data, so they keep most common words whole and split rarer ones; a typical English word is about 1.3 tokens. Big models read trillions of tokens during training.

2

Step 2 · Embeddings

Turn words into places on a map

Each token becomes a list of numbers, like map coordinates. Words with similar meanings end up close together. Real models use thousands of dimensions; this map shows just two.

The famous trick: the direction from man → woman is about the same as king → queen. Meaning becomes geometry.

A meaning map (simplified)

Two-dimensional map where animal words cluster together, food words cluster together, and royalty words form a parallelogram with man and woman ANIMALS catdogkitten FOOD pizzatacoburger PEOPLE & ROYALTY man woman king queen same arrow =same relationship
3

Step 3 · Attention (the big idea)

Every word asks:
who matters to me?

Read this sentence. What does “it” mean? You know instantly, because of the last word. Attention lets the model make the same connection. Switch the last word and watch the arrows move.

Attention diagram

Real models run dozens of these attention “heads” side by side, each looking for a different kind of connection (grammar, who did what, rhymes, facts) and repeat it in layer after layer. This example comes from Google’s 2017 translation model, which reads the whole sentence at once. Chatbot models read left to right, so for them the link is made when the model reaches “tired” or “wide” and looks back at “animal” or “street”.

4

Step 4 · Stack the layers

Repeat it dozens of times

One attention layer plus a “thinking” network (the bigger of the two parts) makes one transformer block. Big models stack dozens of these blocks, and some of the largest stack more than 100 (Meta’s Llama 3.1 405B has 126).

Early layers notice simple things like grammar. Middle layers track meaning and facts. Late layers plan what to say next. Nobody programs these jobs; they emerge during training.

All of the model’s knowledge lives in its parameters: the adjustable numbers inside these layers. Today’s biggest models are thought to have trillions of them.

Inside the stack

Stack of transformer blocks: early layers handle grammar, middle layers meaning and facts, late layers planning the answer BLOCK 1: attention + think BLOCK 2 BLOCK 3 ⋮ BLOCK 99 BLOCK 100 grammar,word partsmeaning,factsplanning thenext word tokens in ↑↑ prediction out
5

Step 5 · Predict the next token

Guess the next word.
Then do it again.

At the top of the stack, the model scores every possible next token. It picks one, adds it to the text, and runs the whole thing again. A long answer is thousands of these guesses in a row.

A setting called temperature controls how often it takes a less-likely word. Higher = more creative, and more likely to wander off.

Watch it write

Illustrative probabilities, not from a real model. Only the top 5 choices are shown, so they add up to less than 100%; the rest is spread over thousands of other tokens.

From predictor to assistant

How does that
make an AI?

A raw next-word predictor is like a brilliant parrot that has read the whole library. Turning it into a helpful assistant takes four more stages of training.

01 · Pretraining

Read everything

Predict the next word across trillions of words of books, websites and code. Takes months on tens of thousands of chips.

Learns: language, facts, patterns

02 · Instruction tuning

Learn to help

Train on examples of good questions and answers so it responds like an assistant instead of just continuing text.

Learns: follow instructions

03 · Feedback

Learn what people want

People (and AI judges) rate answers. The model is nudged toward helpful, honest, safe ones. Often called RLHF.

Learns: helpful & safe behavior

04 · Reasoning training

Learn to think it through

Practice on problems with checkable answers (math, code) and get rewarded for getting them right, so it learns to work step by step.

Since 2024: o1, R1, Claude, Gemini…

05 · Tools & agents

Learn to act

Give it tools (web search, code, a computer) and train it to use them over long tasks.

2025–26: AI agents

Why bigger worked

Scaling laws

In 2020 researchers found that adding more data, more parameters and more computing power made transformers better in a smooth, predictable way. That discovery set off the race to build giant data centers.

A real weakness

Hallucinations

The model is built to produce likely-sounding text, not to check facts. When it doesn’t know, it can still produce a confident, fluent, wrong answer. Always verify claims that matter.

Beyond words

Not just text

The same design works on anything you can chop into tokens: image patches, sound, video frames, DNA, robot movements. That’s why one architecture now powers chatbots, image generators and more.

Put curiosity to work

Read it. Try it. Question it.

Explore new AI research, or take a paper-first investigation into your classroom.