History · Read 7 min

A brief history of AI: from Turing to the transformers

None of the machines that talk to us today would exist without an uncomfortable question posed in 1950, an optimistic summer in 1956 and two winters that nearly buried it all. This is the story of that stubbornness.

In October 1950, in the pages of the philosophy journal Mind, Alan Turing did something a mathematician of his standing was not supposed to do: he changed the question. Rather than debating whether machines can think, an argument he judged doomed to die of definitions, he proposed a game. An interrogator types questions through a screen; on the other side sit a person and a machine. If the interrogator cannot tell which is which, then what are we left with? Turing called it the "imitation game". We call it, perhaps unfairly, the Turing Test. That paper, "Computing Machinery and Intelligence", described no real machine. It described a wager.

01 · the namingA summer in New Hampshire

The wager took six years to find a name. In the summer of 1956, a young mathematician named John McCarthy summoned a handful of colleagues (among them Marvin Minsky, Claude Shannon and Nathaniel Rochester) to a two-month workshop at Dartmouth College. In the funding proposal, McCarthy wrote for the first time the two words that concern us here: artificial intelligence. It was not a flash of collective genius; it was, rather, a pragmatic label to sell a project. But labels found disciplines. From that summer, messy, ambitious, richer in promises than in results, was born the field that today moves trillions.

McCarthy did not invent artificial intelligence at Dartmouth. He invented its name, and the name was enough to found the field.

Figura 1 · línea de tiempo
19501956196920122017 Turingthe test Dartmouth"AI" is born Minsky &Papert AlexNetdeep learning "Attention"transformers
Seventy years condensed into five dates. Each marks a turn on which the ones that followed depended.

02 · the promise and the coldMachines that would think "in twenty years"

The early years brimmed with euphoria. Herbert Simon and Allen Newell built programs that proved theorems; there were automatic translators, fledgling chess, systems that solved algebra problems. In 1965 Simon went so far as to forecast that within twenty years machines would do "any work a man can do". The enthusiasm ran so high that public money flowed with few questions asked.

The cold arrived twice over. In the early 1970s, a report commissioned from the mathematician James Lighthill in the United Kingdom concluded that AI had failed to keep its grandiose promises; the agencies slashed funding. A first AI winter set in. A second would come in the 1980s, when the costly "expert systems" that companies had bought proved brittle and hard to maintain. The word "AI" became, for a time, almost embarrassing to write in funding applications.

Figura 2 · los dos inviernos de la IA
195619741987today 1st winter 2nd winter INTEREST AND FUNDING ↑
The field never advanced in a straight line: two "winters" (in the mid-1970s and the late 1980s) nearly buried it, and each had been preceded by an excess of promises.

03 · the neuron that nearly diedThe perceptron and its gentle executioner

Meanwhile, a parallel and subterranean story was unfolding. In 1958, the psychologist Frank Rosenblatt had unveiled the perceptron: a model inspired by neurons that learned to classify patterns by adjusting weights, instead of following rules written by a human. The press grew excited; there was talk of machines that would walk and speak. But in 1969, Marvin Minsky (the same man from Dartmouth) and Seymour Papert published Perceptrons, a book that proved with mathematical elegance that a single-layer perceptron could not solve problems as simple as the XOR logical function.

The book did not kill neural networks. But it laid on them a slab that took nearly twenty years to lift.

The critique was technically true and yet incomplete: the limitations dissolved once you stacked several layers. What was missing was an efficient way to train those layers. It would arrive in the 1980s with the popularization of backpropagation. But the prestige of the symbolic approach and the weight of Minsky's book had steered nearly an entire generation of researchers away from the connectionist path.

04 · the thaw2012: the night the machines began to see

The definitive thaw has a date and a contest. In 2012, at the annual ImageNet image-recognition challenge, a team from the University of Toronto (Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton) presented a deep neural network we now call AlexNet. It did not win by a hair: it crushed the competition, cutting the error to nearly half that of its rivals. The combination was explosive: plenty of labeled data, deep networks and the muscle of graphics cards (GPUs), designed for video games and repurposed into computation engines.

That night legitimized what had for decades been almost a heresy. Deep learning ceased to be an academic curiosity and became the dominant current. Hinton, one of those who had kept the connectionist flame alive through the long winter, would receive the Turing Award in 2018 alongside Yoshua Bengio and Yann LeCun.

05 · the sentence that changed everythingAttention is all you need

Five years later, in 2017, a group of researchers at Google published a paper with an almost provocative title: "Attention Is All You Need". Vaswani, Shazeer, Parmar and their coauthors proposed abandoning the recurrent architectures that processed text word by word, single file, and replacing them with an attention mechanism able to look at all the words in a sentence at once and decide which matter to which. They called it the Transformer.

Almost everything you call "AI" today (the chatbots, the text and image generators) descends, in a direct line, from that 2017 paper.

The architecture turned out to be astonishingly scalable: the more data and the more parameters you fed it, the better it worked, without soon hitting a ceiling. From there to the large language models, the LLMs, was a single conceptual step but a gigantic leap in scale. Trained on enormous swaths of human text, those models learned to predict the next word so well that, almost as a byproduct, they began to translate, summarize, program and converse. Turing's wager, seventy years on, had stopped being a parlor game.

What we learned

The history of AI is not an ascending line but a pendulum between euphoria and disillusion. The winters were no accidents: they were born of promising too much. And the current revolution did not spring from nothing but from an old idea (neural networks) that survived contempt until data and compute made it viable. None of the machines that dazzle us today would have arrived without those who persisted when no one was betting on them.

Sources

  1. Turing, A. M. (1950). "Computing Machinery and Intelligence". Mind, 59(236), 433-460.
  2. McCarthy, J., Minsky, M., Rochester, N., Shannon, C. (1955). "A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence".
  3. Rosenblatt, F. (1958). "The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain". Psychological Review, 65(6).
  4. Minsky, M. y Papert, S. (1969). Perceptrons: An Introduction to Computational Geometry. MIT Press.
  5. Lighthill, J. (1973). "Artificial Intelligence: A General Survey". Science Research Council, Reino Unido.
  6. Krizhevsky, A., Sutskever, I., Hinton, G. (2012). "ImageNet Classification with Deep Convolutional Neural Networks". NeurIPS.
  7. Vaswani, A. et al. (2017). "Attention Is All You Need". NeurIPS.

Learn more about AI

See all Learn AI