architectures
Cross-Attention
An attention mechanism where queries come from one sequence while keys and values come from a different sequence. Cross-attention enables models to relate information across different inputs, such as an image and text in multimodal models.
In practice
In image captioning, cross-attention allows the text decoder to attend to different regions of the encoded image.