Encoder vs. Decoder vs. Encoder-Decoder: Choosing the Transformer Shape
One attention block, three wiring choices: let every token see everything (encoder, a champion reader), let each token see only the past (decoder, a champion writer), or read the input fully and write the output while glancing back (encoder-decoder, a champion translator). The masking policy alone decides which tasks each family can win — understanding, generation, or transformation.
01.The Problem: Same Bricks, Three Different Jobs
You have met the Transformer twice now.
- In BERT (Concept 134): a stack that turns words into context vectors and reads.
- In GPT (Concept 137): a stack that predicts the next word and writes.
Wait — aren't they the same building blocks?
They are. Same self-attention, same feed-forward layers, same residual connections — all from the original 2017 paper "Attention Is All You Need" (arXiv:1706.03762).
The difference is one policy decision:
Who is allowed to look at whom?
That single question — the masking policy — produces three model families, and it quietly decides which jobs each family can win:
- a hiring manager who must understand a full résumé (read everything, decide),
- a novelist producing one word at a time with no draft in front of her (write forward, never re-plan),
- a conference interpreter: one person listens to the whole speech, another speaks the translation while the listener passes notes.
Get this three-way split straight and every architecture debate in NLP becomes a bookkeeping exercise.
Three Ways to Wire Attention 🔀
Three Ways to Wire Attention 🔀
Masking policy is the whole difference: encoder = no mask, decoder = causal mask, encoder-decoder = encoder reads the source, a causal decoder writes the target while cross-attending back.
Unlock Topic #139: Encoder vs. Decoder vs. Encoder-Decoder: Choosing the Transformer Shape
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?