Back to Roadmap
Transformers 5 min read

Attention Mechanism

The core innovation behind modern LLMs, allowing models to weigh the importance of different words in a sequence dynamically.

1. What is it?

The Attention Mechanism allows a neural network to focus on specific parts of the input sequence when producing a specific part of the output, rather than treating all inputs equally.

2. Why do we need it?

Before attention, models squashed entire sentences into a single fixed-size vector, causing them to "forget" earlier words in long sentences. Attention solves this by looking at the entire sequence at once.

3. Visual Explanation

4. Mathematical Explanation

The standard Scaled Dot-Product Attention is defined as:

Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V

Where:

  • QQ is the Query matrix (what we are looking for).
  • KK is the Key matrix (what the input represents).
  • VV is the Value matrix (the actual content).
  • dk\sqrt{d_k} scales the dot product to prevent vanishing gradients during the softmax step.