architectures
GPT Architecture
A decoder-only transformer architecture that generates text autoregressively, predicting one token at a time from left to right. The GPT architecture uses masked self-attention to prevent the model from looking ahead, and scales effectively to billions of parameters.
In practice
The GPT architecture has been scaled from GPT-1 (117M params) to GPT-4 (estimated trillions of params).