Vector Operations: Addition, Scaling, and Linear Combinations
Adding vectors means adding position by position. Multiplying by a scalar means stretching every position. From word arithmetic (king − man + woman ≈ queen) to residual connections and every optimizer update, these two tiny moves are the workhorses of ML.
Where Vector Addition Shows Up in a Network
The same element-wise sum powers residual connections (u + v) and optimizer updates (w - alpha * g).
01.The Problem: You Have Two Lists of Numbers. Now What?
In the previous topic you learned that a vector is just a list of numbers — one address per entry.
Say a model hands you two lists:
a = [1.0, 2.0, 3.0]
b = [0.5, -1.0, 2.0]
Now you want to combine them. But "combine" is vague, so you face two small questions
What does it even mean to add two lists of numbers?
What does it mean to "multiply a list by a plain number"?
Every fancy ML operation — word arithmetic, residual connections, optimizer updates — is built from just these two answers. So get them rock solid here.
02.The Idea in Plain Words: Two Primitive Operations
Given two vectors a, b ∈ Rⁿ and a scalar α ∈ R, there are exactly two primitive moves.
Addition means: add matching components, position by position.
a + b → take entry 1 of a plus entry 1 of b, entry 2 plus entry 2, and so on. Geometrically it is a translation — sliding one arrow tip-to-tail. Three facts you can trust:
- Addition is commutative (
a + b = b + a) and associative ((a + b) + c = a + (b + c)). a + 0 = a: the zero vector changes nothing — it's the identity.
Scalar multiplication means: multiply every component by the same number.
α · a → stretch or shrink the arrow. If α < 0 it also flips the direction.
Everything else in linear algebra is built from these two. From them you get the one big word
A linear combination is just scaled additions chained together.
c₁v₁ + c₂v₂ + … + cₖvₖ
And one more term you'll meet constantly:
- The span of a group of vectors is the set of all linear combinations you can build from them. This is the intuition behind "a basis is a minimal set of vectors whose span is the whole space."
import numpy as np
u = np.array([1.0, 2.0, 3.0])
v = np.array([0.5, -1.0, 2.0])
print(u + v) # [1.5 1. 5. ] element-wise sum
print(3.0 * u) # [ 3. 6. 9. ] scalar multiplication
print(2.0*u - 0.5*v) # linear combination
bias = np.array([1.0, 1.0, 1.0])
batch = np.array([[1., 2., 3.], [4., 5., 6.]])
print(batch + bias) # bias added to every row: broadcasting03.A Tiny Worked Example, With Real Numbers
Let's do u = [1, 2, 3] and v = [0.5, -1, 2] by hand, one entry at a time.
Addition (u + v):
- position 1:
1 + 0.5 = 1.5 - position 2:
2 + (-1) = 1 - position 3:
3 + 2 = 5 - result:
[1.5, 1, 5]
Scaling (3 · u): multiply every entry by 3.
[3·1, 3·2, 3·3] = [3, 6, 9]
A linear combination (2u − 0.5v): scale each, then add.
2u = [2, 4, 6]0.5v = [0.25, -0.5, 1]2u − 0.5v = [2 − 0.25, 4 − (-0.5), 6 − 1] = [1.75, 4.5, 5]
Notice the shape rule hiding here: you can only add vectors of the same length, entry lined up with entry.
But the code above did something sneaky — it added a length-3 bias to a 2×3 batch. That's broadcasting: the small vector is conceptually repeated to match each row, so bias rides onto every row without a copy. The rules match sizes from the rightmost axis first (3 vs 3 → OK; a missing axis counts as size 1).
04.Visual Intuition: Arrows, Tip-to-Tail
Picture each vector as an arrow drawn from the origin.
codeaddition u + v : slide v so its tail sits on u's tip ▲ u starts at origin │ ╱╲ v then v continues from u's tip │ ╱ ╲ │╱ ╲ ◄ final tip ●──────╲────────► u+v = straight line origin → final tip └─────────────────► `` scaling α·u : the same line, only the length changes α = 2 ●──────────────► stretch: twice as long α = 0.5 ●─────► shrink: half as long α = -1 ◄─────● flip: same length, opposite way
- Adding = "go along
u, then from there go alongv." The shortcut from start to finish isu + v. - Scaling = "turn the volume knob on the whole arrow." Longer, shorter, or reversed — never bent.
That's also why hardware loves these two moves: both are O(n) and perfectly parallel — every GPU core and SIMD lane can handle one position at the same time.
05.The Analogy: A Mixing Desk with Sliders
Carry one image through the rest of this topic: an audio mixing desk.
A vector is the set of slider positions. [1, 2, 3] means "channel 1 at 1, channel 2 at 2, channel 3 at 3."
- Addition = laying two settings on top of each other. Blend track A with track B; every slider moves by the matching amount.
- Scaling = a master volume knob.
α = 2doubles every slider at once;α = 0.5halves them;α = -1flips the signal. - Linear combination = a remix:
2·trackA − 0.5·trackB— set each track's level, then blend.
Keep this desk in mind, because
Embedding arithmetic, residual connections, and optimizer updates are all "add some settings, scale some settings, blend them."
06.Meaning in Embedding Space
The desk stops being a toy the moment a model maps words, images, or users into vector space — because then vector operations become semantic operations. The classic (approximate) embedding arithmetic king − man + woman ≈ queen works because addition and scaling move points along directions that models learn to correlate with attributes like gender or tense. In desk terms: "maleness" happens to live along one group of sliders, so subtracting man and adding woman nudges the mix the right way.
Modern systems rely on the same idea at scale:
- Retrieval pipelines add a query vector to a small "reformulation" vector to broaden semantic search.
- Diffusion models literally add scaled Gaussian noise to a latent vector during the forward process and learn to reverse that sum during denoising.
- Model merging (2023–2026 practice, e.g. merging fine-tuned checkpoints) is element-wise scalar-weighted addition of whole parameter vectors:
θ_merged = α·θ₁ + (1−α)·θ₂.
07.Residual Connections: Addition as Architecture
ResNet (2015) and every modern Transformer block make vector addition a first-class architectural primitive: instead of computing f(x), the block computes f(x) + x, adding the input vector element-wise to the transformed output.
On the mixing desk: the block reads the current mix, produces its own small adjustment, then adds that adjustment on top instead of replacing the whole track.
Why it matters:
- The added
xpath guarantees gradient can flow directly backwards, fixing the vanishing-gradient problem in 100+ layer stacks (see Topic 12, Chain Rule). - Each block performs a small translation of the representation rather than replacing it, so depth accumulates refinements — a literal linear combination along the residual stream.
- Mechanistic interpretability (2024–2026) analyzes transformers as a residual stream: all attention heads and MLPs just read from and add to one shared vector per token.
08.In Practice: Optimizers Are Vector Arithmetic
The parameter update of every optimizer is a linear combination of vectors. On the desk, the gradient says which sliders to move and by how much, the learning rate η is the master knob that scales that instruction, and adding it to w is one training step.
- SGD:
w ← w − η·g(add the scaled gradient). - Momentum:
v ← μ·v + g; w ← w − η·v— an exponentially weighted linear combination of past gradient vectors. - Adam (2024+ standard variant AdamW): builds updates from element-wise combinations of running mean and RMS of gradients, plus a decoupled weight-decay term
−λ·w.
Understanding vector addition and scaling is therefore a prerequisite to understanding training itself, not just "math warm-up."
Architectural Trade-offs & Production Realities
Architectural Advantages
- Element-wise ops are perfectly parallelizable across GPU cores and SIMD lanes.
- Broadcasting removes explicit loops and copies when adding biases across batches.
- Residual addition stabilizes deep-network training and enables model-merging tricks.
Trade-offs & Constraints
- Broadcasting can silently add dimensions and hide shape bugs.
- Embedding-space arithmetic is only approximately meaningful — directions are not guaranteed orthogonal attributes.
- Adding in-place during the forward pass can break autograd ("modified by an operation after it was saved for gradient").
Each LLaMA layer computes attn(x) + x and mlp(x) + x — pure vector addition on (batch, seq, d_model) tensors. Fine-tuning methods like LoRA likewise *add* a low-rank update matrix to the frozen weight: W' = W + BA.
Staff+ Engineering Takeaways
- Vector addition = element-wise sum (a translation); scalar multiplication = scaling; both are O(n) and fully parallel.
- A linear combination c1v1 + ... + ckvk and its span underpin bases, rank, and overparameterization.
- Embedding arithmetic (king − man + woman) works because learned directions approximately encode attributes.
- Residual connections (f(x) + x) are vector addition as architecture and cure vanishing gradients in deep stacks.
- Every optimizer update is a scaled linear combination of gradient vectors.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
In a Transformer block, what operation implements the "residual connection"?
How clear and actionable was this distributed systems breakdown?