Skip to content

architectures

Transformer

A neural network architecture introduced in the 2017 paper 'Attention Is All You Need' that processes sequential data using self-attention mechanisms instead of recurrence. Transformers enable massive parallelization during training, making them the foundation of modern LLMs and vision models.

In practice

Nearly all modern language models, including GPT-4, Claude, and Gemini, are built on the transformer architecture.