The history of AI

arqiv✳

03 / History

80 years of thinking machines

AI didn’t appear overnight with ChatGPT. It’s a story of big dreams, two “winters” when funding dried up, and a surprise comeback powered by data, graphics chips and one 2017 idea. Here’s the whole ride.

The big picture · hype, winters and booms

Timeline of AI eras from 1943 to 2026 with a hype curve showing peaks in the 1960s and 1980s, winters in the 1970s and early 1990s, and a steep rise after 2012
The curve shows excitement and funding, not a measured number. Notice the two crashes: each came when AI promised more than the computers of the day could deliver.

Milestones

The moments
that mattered

Filter by era. Highlighted cards are the true turning points.

Why AI took off after 2012

The scaling story

The ideas behind neural networks are decades old. What changed was the fuel: far more data, and far more computing power. The biggest AI training runs now use over a billion times as much compute as in 2012.

Computing power used to train landmark AI models

Bar chart on a log scale: training compute rising from AlexNet in 2012 to frontier models in 2025 and 2026, an increase of more than a billion times
Measured in FLOP, the total number of calculations. Each gridline is 100× the one below. Values are rounded estimates from Epoch AI; labs don’t publish exact numbers for the newest models. Dashed bars are the roughest: GPT-6 Astra’s figure is an early estimate, and published estimates for AlphaGo Zero differ by more than 100×. Training compute for frontier models has grown roughly 4–5× per year.

Fuel 1

Data

The internet created a giant library of text, images and video. ImageNet (launched in 2009) grew to more than 14 million labeled photos; today’s models train on tens of trillions of words.

Fuel 2

Chips

Graphics chips (GPUs) built for video games turned out to be perfect for the math neural networks need. A single 2026 training run can use over 100,000 of them.

Fuel 3

Better recipes

Backpropagation (1986), deep networks (2012), transformers (2017) and reasoning training (2024) each let the same chips get more out of the same data.

Put curiosity to work

Read it. Try it. Question it.

Explore new AI research, or take a paper-first investigation into your classroom.