1. What is it?
The Attention Mechanism allows a neural network to focus on specific parts of the input sequence when producing a specific part of the output, rather than treating all inputs equally.
2. Why do we need it?
Before attention, models squashed entire sentences into a single fixed-size vector, causing them to "forget" earlier words in long sentences. Attention solves this by looking at the entire sequence at once.
3. Visual Explanation
4. Mathematical Explanation
The standard Scaled Dot-Product Attention is defined as:
Where:
- is the Query matrix (what we are looking for).
- is the Key matrix (what the input represents).
- is the Value matrix (the actual content).
- scales the dot product to prevent vanishing gradients during the softmax step.