Attention (Bahdanau): Stop Compressing, Start Consulting
The old machine-translation models squeezed a whole sentence into one fixed vector before writing the translation — and long sentences got mangled. Bahdanau attention says: keep every encoder state, and at each decoding step take a learned, softmax-weighted look-back at all of them. Score → softmax → weighted sum: that pipeline is the ancestor of all cross-attention.
01.The Problem: A Whole Sentence Through One Small Bag
Picture the machine-translation setup from Topic 116 (in one plain sentence: an encoder RNN reads the source sentence word by word, then a decoder RNN writes the translation word by word).
The catch was how the handoff worked.
The decoder only ever saw one vector: the encoder's final hidden state, c = h_n.
So the encoder had to squeeze:
- 3 words or 300 words
- every noun, verb, tense, and agreement
- into a fixed-size vector (say 512 numbers)
What if the sentence is long?
Then 512 numbers is a lossy photocopy of a 300-page book.
And the experiments agreed. Translating with a single fixed context vector, BLEU scores (the standard translation-quality metric) fell off a cliff once the source passed roughly 30 words.
There was a second pain too.
How does the error signal for the last French word reach the first English word?
Through that one squeezed vector. Long dependency = long, vanishing gradient path (the Topic 113 disease).
So the question becomes
Why compress at all?
Why throw away every encoder state except the last one, when they were just computed and could simply be kept?
That single "why not keep them?" is Bahdanau attention.
Bahdanau’s Per-Step Context: Score → Softmax → Weighted Sum 🎯
Bahdanau’s Per-Step Context: Score → Softmax → Weighted Sum 🎯
The decoder never loses access to the source: at each step it "asks" every encoder state for help, softmax turns raw affinity scores into a probability mixture, and c_t is re-derived from the full encoder memory.
Unlock Topic #118: Attention (Bahdanau): Stop Compressing, Start Consulting
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?