How It Works
Transformer
The neural network design underlying nearly all modern LLMs; processes an entire sequence of text at once using “attention” instead of reading word-by-word, making it faster to train and better at long-range context.
Origin
Introduced by Ashish Vaswani and colleagues at Google in the 2017 paper “Attention Is All You Need.”
