Act 05 · Architectures & the model zoo 2:45 The road to the transformer

ImageNet, and the spark that lit the era.

In 2012, AlexNet — a deep convolutional net trained on GPUs — won the ImageNet challenge by a landslide, the collision of big data (ImageNet), compute (GPUs), and old connectionist ideas that lit the deep-learning era and led straight to 2017's transformer; the winters had been waiting, not failing.

Video rendering soonThe cinematic render for this supplement is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — In 2012, AlexNet — a deep convolutional net trained on GPUs — won the ImageNet challenge by a landslide, the collision of big data (ImageNet), compute (GPUs), and old connectionist ideas that lit the deep-learning era and led straight to 2017's transformer; the winters had been waiting, not failing.
  • How it is shown — Three streams (data, compute, old ideas) converge at a threshold; Fei-Fei Li builds an ocean of labeled images; a shared arena barely moves for two years; AlexNet's gauge plunges by a canyon-wide margin; the two winters re-light from within as a dam bursts toward the transformer.
  • The trap to avoid — Seeing AlexNet as a brand-new genius idea rather than a fifty-year-old idea finally fed — the breakthrough was scale meeting theory, not novelty.

Two winters. Fifty years of an idea that worked in theory and starved in practice.

The one idea

In 2012, AlexNet — a deep convolutional net trained on GPUs — won the ImageNet challenge by a landslide, the collision of big data (ImageNet), compute (GPUs), and old connectionist ideas that lit the deep-learning era and led straight to 2017's transformer; the winters had been waiting, not failing.

This episode is where the ingredients finally show up all at once — and the whole thing catches fire. This is the payoff. Start with data. Around two thousand nine, most researchers were polishing algorithms on tiny datasets. Fei-Fei Li bet the opposite: the models weren't too simple, they were too hungry. So she built ImageNet — millions of images, hand-labeled into thousands of categories through crowdsourcing. A dataset enormous enough to feed a deep network. To prove data mattered, you needed a scoreboard. From twenty-ten, ImageNet ran a public contest: one thousand categories, over a million images, and one number — how often your model's top guesses were wrong. Every team, same test. For two years, the winners were conventional, and progress crawled.

How it works — the demo

Three streams (data, compute, old ideas) converge at a threshold; Fei-Fei Li builds an ocean of labeled images; a shared arena barely moves for two years; AlexNet's gauge plunges by a canyon-wide margin; the two winters re-light from within as a dam bursts toward the transformer.

Then twenty-twelve. A team from Toronto — Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton — entered a deep convolutional neural network called AlexNet. It didn't edge ahead. It won by roughly ten percentage points, a margin nobody had ever seen. The trick that made it trainable: gaming GPUs — the chips Act four explained. Look at what happened there. AlexNet was not a bolt from nowhere. It was a convolutional network — LeCun's idea — trained by backpropagation — Rumelhart and Hinton's idea — on Fei-Fei Li's data, using gaming hardware. Three things never ready at once, finally colliding: old ideas, big data, real compute. And here's the payoff of this whole supplement. It's easy to call AlexNet a sudden breakthrough.

The trap to avoid

Seeing AlexNet as a brand-new genius idea rather than a fifty-year-old idea finally fed — the breakthrough was scale meeting theory, not novelty.

Why it matters — and what’s next

It wasn't. It was a fifty-year-old idea finally fed. That reframes the winters: they weren't failures, they were waiting — the right theory, patient in the cold, until data and compute caught up. After twenty-twelve, the dam broke. The pace turns dizzying. Within three years, error rates fall past human level. Speech, vision, language — it all works. The same recipe keeps paying out, until two thousand seventeen, when a paper titled Attention Is All You Need introduces the transformer. That's where the main spine begins — and now you know what it stands on.

This is a supplement in AI: Zero → Frontier — a side-trip that deepens the act it sits beside, one file and one loop at a time.

Full transcript 2:45 of narration

Two winters. Fifty years of an idea that worked in theory and starved in practice. This episode is where the ingredients finally show up all at once — and the whole thing catches fire.

This is the payoff. Start with data. Around two thousand nine, most researchers were polishing algorithms on tiny datasets.

Fei-Fei Li bet the opposite: the models weren't too simple, they were too hungry. So she built ImageNet — millions of images, hand-labeled into thousands of categories through crowdsourcing. A dataset enormous enough to feed a deep network.

To prove data mattered, you needed a scoreboard. From twenty-ten, ImageNet ran a public contest: one thousand categories, over a million images, and one number — how often your model's top guesses were wrong. Every team, same test.

For two years, the winners were conventional, and progress crawled. Then twenty-twelve. A team from Toronto — Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton — entered a deep convolutional neural network called AlexNet.

It didn't edge ahead. It won by roughly ten percentage points, a margin nobody had ever seen. The trick that made it trainable: gaming GPUs — the chips Act four explained.

Look at what happened there. AlexNet was not a bolt from nowhere. It was a convolutional network — LeCun's idea — trained by backpropagation — Rumelhart and Hinton's idea — on Fei-Fei Li's data, using gaming hardware.

Three things never ready at once, finally colliding: old ideas, big data, real compute. And here's the payoff of this whole supplement. It's easy to call AlexNet a sudden breakthrough.

It wasn't. It was a fifty-year-old idea finally fed. That reframes the winters: they weren't failures, they were waiting — the right theory, patient in the cold, until data and compute caught up.

After twenty-twelve, the dam broke. The pace turns dizzying. Within three years, error rates fall past human level.

Speech, vision, language — it all works. The same recipe keeps paying out, until two thousand seventeen, when a paper titled Attention Is All You Need introduces the transformer. That's where the main spine begins — and now you know what it stands on.

AI HistoryDeep LearningComputer Vision