Next Sentence Prediction: BERT's Most Controversial Objective
Next sentence prediction was BERT's second training task: show the model two sentences and ask, "does B really follow A?" It was bolted onto masked language modeling to give BERT document-level coherence — and it became the cautionary tale of the field, when RoBERTa's experiments showed the task was doing more harm than good.
01.The Problem: Blanks Taught Sentences, Not Stories
Recall the tool from the last two topics. Masked language modeling (MLM) — Concept 135 — is endless fill-in-the-blank practice inside a single sentence.
It works beautifully for one sentence at a time.
But look at the tasks BERT was aimed at:
- Question answering: Does this sentence answer that question?
- Natural language inference: Does this passage entail that claim?
The essential signal there does not live inside a sentence. It lives between sentences.
Can fill-in-the-blank alone teach cross-sentence understanding?
Not really. MLM never forces the model to ask "do these two sentences belong together?" Each blank is answered from one window of text.
So Devlin et al. wanted a second exercise, one that requires comparing two complete sentences. Cheap. Automatic. No annotators — the same principle as MLM: the document itself must supply the answer key.
Documents have a free answer key lying around: order. In a real book, sentence B really did follow sentence A. And a random sentence from another book did not.
That observation became Next Sentence Prediction (NSP).
How NSP Training Pairs Are Built 🎲
How NSP Training Pairs Are Built 🎲
Half the pairs are true adjacent sentences from the same document; half are random corpus sentences. The [CLS] representation is trained to tell them apart.
Unlock Topic #136: Next Sentence Prediction: BERT's Most Controversial Objective
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?