architectures
Vision Transformer
A transformer architecture adapted for image processing that divides images into patches and processes them as a sequence, similar to how text transformers process tokens. ViTs have matched or exceeded CNNs on many computer vision benchmarks.
In practice
ViTs power image understanding in multimodal models like GPT-4V and Claude, enabling them to analyze uploaded photos.