Hidden State: The Compressed Memory of the Past
The hidden state h_t is one fixed-size vector — the model’s entire memory of everything it has seen. Learn what a few hundred floats can store, what they provably cannot, and why the "summarize vs re-read" choice still decides 2024-2026 serving economics.
The Hidden State Is a Lossy Funnel 🧠
No matter how many tokens arrive, the recurrent model squeezes them into one d_h-dimensional vector. That single design choice grants O(1) memory and imposes the capacity ceiling that attention later escapes.
01.The Problem: Answering About Hour 3 During Hour 7
Picture a meeting that has run for four hours.
Someone asks you: "What did Priya say about the budget in hour 3?"
You have two options:
- Re-read the transcript of everything said so far. Exact — but slow, and the transcript grows all day.
- Use your running summary: what you believe matters, kept in your head. Fast, cheap — and lossy.
A recurrent model must do one of these every single time step. Topic 111 showed an RNN chooses the second option: it updates one vector per step and throws the raw past away.
But what actually fits inside that one vector? And who decides?
That question — the memory itself — is this topic. It is the reason RNNs feel "streaming and cheap," and the reason they eventually lost to architectures that re-read.
02.The Idea in Plain Words: h_t Is a Lossy, Learned Summary
Formally, h_t ∈ R^{d_h} is the value of the recurrent activation after consuming tokens x_1, ..., x_t. Functionally, it is a lossy, learned summary of the entire prefix:
h_t = f(h_{t-1}, x_t)
For a vanilla RNN, f is tanh(W_h·h_{t-1} + W_x·x_t + b).
Two details in that sentence do enormous work:
- Lossy. Nothing stores the original tokens. The past exists only as compressed into the current vector.
- Learned. Nobody wrote a spec for what to keep. The readout
y_tis computed fromh_talone — so whatever the loss function rewarded surviving. The hidden state is the model’s entire memory, and every sequence claim the model makes is mediated by what made it into this vector.
Two intuitions that stick:
- Pointer register: like a CPU register updated every cycle. The sequence is the program tape; h is the accumulator.
- Sufficient statistic: a good
h_ttries to capture exactly the features needed to predict the future, discarding the rest — compression target set by the loss, not by a human.
03.A Tiny Worked Example: The Memory That Halves the Past
Make memory comically small: the hidden state is one number, and each step updates it by averaging with the new input:
h_t = (h_{t-1} + x_t) / 2
Feed the stream 0, 100, 0, 0, 0, 0, 0, 0 — one loud event at step 2, then silence.
h_1 = (0+0)/2 = 0h_2 = (0+100)/2 = 50— the event is "in memory."h_3 = 25,h_4 = 12.5,h_5 = 6.25…h_8 = 0.78
The number 100 never disappears completely, but after six silent steps its echo is roughly 130x fainter. Anyone reading h_8 and guessing "something big happened once" would need luck.
Now scale the same arithmetic up: a real RNN squashes with tanh and multiplies by the same matrix W_h every step (Topic 113 shows why that is exponential forgetting, not linear). The lesson survives at any size:
An update-rewritten memory is a leaky memory.
Retention here is a property of the mechanism, not a bug or a tuning failure — which is exactly what the gates of Topics 114-115 will be invented to override.
04.Capacity: What a Few Hundred Floats Cannot Hold
A 512-dimensional float32 hidden state holds ~2 KB — and it must hold everything: grammar state, discourse entities, quoted text 400 tokens back, numeric running totals. The consequences are measurable, not philosophical:
- Vanilla RNNs reliably preserve only ~5-10 steps of exact history before information is scrambled by repeated nonlinear projections (Topic 113 explains the mechanism).
- LSTMs stretch faithful ranges to hundreds of steps, but still fail hard on exact copy/retrieval at length: a classic probe task, "repeat the sequence you just read," which no finite-state-style compression solves gracefully as T grows.
- Memorizing arbitrary content scales parameters linearly in capacity for state size — impractical; this is precisely the gap that made attention (arbitrary-size exact recall) decisive.
05.The Analogy: Minutes on a Single Sticky Note
Carry this analogy through the topic: you are taking meeting minutes, but you are only allowed one sticky note.
- Each hour, you rewrite the note: the best summary of "note so far + what just happened." That rewrite is
h_t = f(h_{t-1}, x_t). - The note never grows. Hour 7 is the same size as hour 1 — that is the fixed dimension
d_h. - What survives the rewriting? Decisions, open conflicts, who owes what — the things likely to come up again. That is the learned part: your boss’s questions (the loss function) train your judgment about what to keep.
- What leaks out? Exact quotes, who said them, the joke in minute 40. Nobody planned the loss; it is the price of one note.
- When a new meeting starts, you toss the note. That is state reset — and "save the note for the follow-up meeting" is state carry.
Then why accept one sticky note at all?
Because it costs nothing to carry, update, and store. The whole architecture debate of this phase is: sticky note (cheap, lossy) versus full transcript re-read (exact, growing). Keep this note in mind — LSTM adds a filing cabinet, Transformers carry the transcript, and 2024-2026 hybrids decide per layer.
06.State Mechanics in Practice: Init, Carry, Reset
Engineering with hidden states is engineering around three lifecycle decisions:
- Initialization: conventionally zeros. In production (speech, translation) the final state of a segment becomes the initial state of the next, keeping discourse context across chunk boundaries.
- Carrying through training: truncated BPTT detaches the graph every k steps but carries the value forward — memory flows, gradients do not. This is the quiet reason RNNs "forget" beyond the truncation horizon.
- Resetting: a session boundary, a new document, or a privacy/session cache eviction discards state. In multi-tenant serving, state is user data: it must be isolated per session and bounded. An LLM’s equivalent — the KV cache, i.e. the stored "transcript" per request — is 10-100x heavier: per-request state of 1-40 GB at long context in 2024-2026 deployments is a first-order capacity-planning number.
import torch.nn as nn
lstm = nn.LSTM(hidden_size=512, num_layers=2, batch_first=True)
h0 = torch.zeros(2, 1, 512) # per-layer initial state
c0 = torch.zeros(2, 1, 512) # LSTM cell state (Topic 114)
out, (hT, cT) = lstm(chunk_1, (h0, c0)) # process a segment...
out2, (hT2, cT2) = lstm(chunk_2, (hT, cT))# ...and carry its final state into the next
# hT IS the "memory checkpoint" of everything seen so far — serializable, resumable.07.Two Directions, Several Depths: BiRNNs and Stacked States
One direction of state is often not enough:
- Bidirectional RNNs (BiRNN): run one forward state and one backward state (over the same sequence, opposite directions), then concatenate them at each position. Every token gets a "past-state + future-state" summary. Standard for tagging and early speech encoders (DeepSpeech-2 used BLSTMs); unusable for online generation because the future doesn’t exist yet.
- Stacked states: layers give each depth its own sequence of hidden vectors; lower layers track surface/local patterns, upper layers track slow, abstract features (a consistent finding in 2010s probing studies).
- The encoder’s final state = the context vector of Topic 116: encoder-decoder architectures are exactly "one model’s hidden state becomes another model’s initial condition" — the sticky note of the understanding phase, handed to the generation phase.
08.What Breaks the Fixed-State Model — and Why It Still Matters
Three structural failure modes of "compress the past into one vector":
- Exact recall — copying a phone number from 300 tokens back requires the content to survive every subsequent update
f. It usually doesn’t. - Interference — writing a new fact overwrites old geometry in
h_t; capacity is shared across all memories simultaneously. - Gradient starvation — even if information persists, learning to use it needs gradients to flow back to the moment it arrived (Topic 113).
The historical arc of this entire phase is an escape plan from these limits:
↓ gates slow the forgetting (LSTM/GRU, Topics 114-115) ↓ then attention replaces compression with re-reading (Topic 118) ↓ Transformers make re-reading the default (Topic 120) …while recurrence keeps resurfacing for its unique O(1)-state economics (Mamba, RWKV).
So the honest summary is a trade, not a ranking: the sticky note is not a worse transcript — it is a different product. If your workload pays per GB of stored context or needs milliwatt streaming, the note wins and the phase is not over.
Architectural Trade-offs & Production Realities
Architectural Advantages
- O(1) memory and compute per step regardless of history length — unbeatable for streaming and edge devices.
- State is serializable: sessions, checkpoint-resume, and cross-segment context fall out naturally.
- Forces useful abstraction: the model must learn task-relevant summaries, not brute-force replay.
Trade-offs & Constraints
- Hard capacity cap: a few hundred floats cannot losslessly store tens of thousands of tokens.
- Exact retrieval/copy at distance is systematically unreliable in vanilla RNNs.
- Interference: concurrent memories contend for the same vector subspace.
State-space and gated-recurrent hybrids keep a fixed-size state per layer instead of a growing KV cache, making per-token decode memory independent of session length. The pitch is literally the hidden-state property from this topic: bounded memory per conversation, at the cost of exact arbitrary-distance recall that Transformers get from re-reading.
Staff+ Engineering Takeaways
- h_t = f(h_{t-1}, x_t) is the entire memory of a recurrent model: a lossy, learned summary of the prefix.
- Fixed-size state buys O(1) per-step memory and compute, and caps faithful retention at hundreds to thousands of floats.
- Exact recall and interference are structural limits of compression, not tuning failures.
- The encoder’s final hidden state becomes the decoder’s initial condition — the seed of the context-vector bottleneck.
- 2024-2026 recurrent revival (Mamba, RWKV) is a bet that constant-size state economics still matter at LLM scale.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Why is the hidden state called a "lossy" summary of the sequence?
How clear and actionable was this distributed systems breakdown?