TOPIC #139Intermediate 16 min read

Encoder vs. Decoder vs. Encoder-Decoder: Choosing the Transformer Shape

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

One attention block, three wiring choices: let every token see everything (encoder, a champion reader), let each token see only the past (decoder, a champion writer), or read the input fully and write the output while glancing back (encoder-decoder, a champion translator). The masking policy alone decides which tasks each family can win — understanding, generation, or transformation.

01.The Problem: Same Bricks, Three Different Jobs

You have met the Transformer twice now.

  • In BERT (Concept 134): a stack that turns words into context vectors and reads.
  • In GPT (Concept 137): a stack that predicts the next word and writes.
Insight

Wait — aren't they the same building blocks?

They are. Same self-attention, same feed-forward layers, same residual connections — all from the original 2017 paper "Attention Is All You Need" (arXiv:1706.03762).

The difference is one policy decision:

Insight

Who is allowed to look at whom?

That single question — the masking policy — produces three model families, and it quietly decides which jobs each family can win:

  • a hiring manager who must understand a full résumé (read everything, decide),
  • a novelist producing one word at a time with no draft in front of her (write forward, never re-plan),
  • a conference interpreter: one person listens to the whole speech, another speaks the translation while the listener passes notes.

Get this three-way split straight and every architecture debate in NLP becomes a bookkeeping exercise.

Three Ways to Wire Attention 🔀

PRO Architecture Blueprint

Three Ways to Wire Attention 🔀

Masking policy is the whole difference: encoder = no mask, decoder = causal mask, encoder-decoder = encoder reads the source, a causal decoder writes the target while cross-attending back.

Three Ways to Wire Attention 🔀
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...
PRO & LIFETIME CURRICULUM

Unlock Topic #139: Encoder vs. Decoder vs. Encoder-Decoder: Choosing the Transformer Shape

You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.

Production Deep Dive

Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.

Interactive Blueprints

Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.

Knowledge Assessment

Staff-level multiple-choice quiz questions with instant feedback and answer explanations.

Cross-Device Progress Sync

Firebase Google authentication automatically syncs your completed topics and quiz scores.

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?