Skip to content

architectures

GPT Architecture

A decoder-only transformer architecture that generates text autoregressively, predicting one token at a time from left to right. The GPT architecture uses masked self-attention to prevent the model from looking ahead, and scales effectively to billions of parameters.

In practice

The GPT architecture has been scaled from GPT-1 (117M params) to GPT-4 (estimated trillions of params).