Skip to content

architectures

Cross-Attention

An attention mechanism where queries come from one sequence while keys and values come from a different sequence. Cross-attention enables models to relate information across different inputs, such as an image and text in multimodal models.

In practice

In image captioning, cross-attention allows the text decoder to attend to different regions of the encoded image.