Masked Language Modeling: Learning from the Blanks
Masked language modeling is the fill-in-the-blank trick that lets a model train on unlabeled text in both directions at once: corrupt about 15% of tokens with an 80/10/10 scheme, predict only the holes, and learn context from every side. With span variants like T5's sentinel tokens, it shapes how every modern encoder learns.
01.The Problem: How Do You Teach Meaning With No Answer Key?
You want a model that understands words.
But "understanding" has no label you can pay annotators for.
Where do the answer keys come from?
Think how humans actually learn vocabulary as kids. Not from definitions. From sentences:
"The fisherman cast his ___ into the river."
You have never seen "net" defined. But the blank is obvious.
Sentence context teaches the word.
Now the hard question for a machine:
Can the model just predict the next word, over and over?
It can — but that only lets it read left to right. Remember the BERT problem (Concept 134): if a model reads both directions and you train it to predict every word from its context, each position can cheat by copying its own input embedding. The answer is already sitting in the input. Circular. Useless.
So the question becomes:
How do we make full bidirectional reading honest?
The trick is embarrassingly simple: erase the word before asking about it. Hide a slice of the input, and demand predictions only for the hidden positions. Now both directions contain zero copies of the answer — the model must actually use context.
That trick is Masked Language Modeling (MLM).
BERT's 80/10/10 Corruption Scheme 🕳️
BERT's 80/10/10 Corruption Scheme 🕳️
MLM creates supervised holes in unsupervised text. The 80/10/10 policy deliberately mixes what the model sees ([MASK], wrong word, or the true word) with what it must predict (always the original).
Unlock Topic #135: Masked Language Modeling: Learning from the Blanks
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?