Skip to content

architectures

Vision Transformer

A transformer architecture adapted for image processing that divides images into patches and processes them as a sequence, similar to how text transformers process tokens. ViTs have matched or exceeded CNNs on many computer vision benchmarks.

In practice

ViTs power image understanding in multimodal models like GPT-4V and Claude, enabling them to analyze uploaded photos.