Exploration vs. Exploitation
The fundamental trade-off: exploit the known-best action or explore to reduce uncertainty. Epsilon-greedy, upper-confidence bound, Boltzmann softening, Thompson sampling, and intrinsic-reward/RLHF-era exploration strategies.
01.The Problem: Your New City Has 500 Restaurants
You just moved to a new city. Tonight you need dinner.
Option 1: go back to the taqueria you tried twice. Average rating in your head: solid. Safe, decent, known.
Option 2: try the new place down the street you have never visited. It might become your favorite. It also might be terrible, and you spend the evening annoyed.
Every meal is this choice. That is the whole dilemma — and an RL agent faces it on every single step.
The two urges, in plain words:
- Exploitation: do the thing currently believed best, to harvest reward now.
- Exploration: do something else, to improve your estimate of its value, which may pay off later.
Both extremes fail:
- Exploit only → you commit to whatever looked best from two visits. A better restaurant you barely sampled loses to a mediocre one that got lucky. You eat at the taqueria for eight years.
- Explore only → you ignore everything you have learned and never bank any reward. You have never eaten a good meal twice in a row.
So the real question is
How do I split my limited number of tries between the known good and the possibly-better?
The Explore/Exploit Dilemma 🎰
The Explore/Exploit Dilemma 🎰
A gambler at many slot machines must decide each pull: play the machine that has paid best so far (exploit) or try under-sampled machines that might be better (explore). Pure exploitation locks in a suboptimal choice; pure exploration never banks reward. Algorithms try to minimize cumulative regret over time.
Unlock Topic #207: Exploration vs. Exploitation
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?