Scalars, Vectors, Matrices, and Tensors
A tensor is just a box of numbers. One number = scalar, a row = vector, a grid = matrix, a stack of grids = higher-rank tensor. The "rank" is how many coordinates you need to grab one number. In real ML, getting the shape right is most of the debugging.
The Rank Ladder: Scalar to Tensor
Each level is simply an ordered collection of the level below it. ML frameworks generalize this to arbitrary rank with the single "tensor" type.
01.The Problem: Numbers Come in Different Shapes
You open a machine learning tutorial. The very first line is a number.
loss = 2.31
That one is easy. It's just a single number.
But a few lines later you see something stranger.
embedding = [0.12, -0.87, 0.44, ...]
That's a list of numbers.
Then this:
weights = [[1, 2], [3, 4], [5, 6]]
A grid of numbers.
And the code keeps talking about "shape", "rank", "dimensions", "axes".
So the question becomes
How do a single number, a list, a grid, and a stack-of-grids all fit into one idea?
And the follow-up that every beginner hits
Why do my errors keep saying "shape mismatch"?
The whole of deep learning moves numbers around. To do it without confusion you need one clean vocabulary for "how many coordinates does it take to point at a single number". That vocabulary is scalars, vectors, matrices, and tensors — and one umbrella word for all of them: the tensor.
02.The Idea in Plain Words: A Tensor Is Just a Box of Numbers
The core definition
A tensor is a container of numbers, laid out on an n-dimensional grid of coordinates.
And the one word you must learn
Rank (or order) = how many indices (coordinates) you need to pick out a single number.
Think of it as "how many addresses does one number need?" Here is the whole ladder, from simplest to richest.
- Scalar — rank 0. A single number. No coordinates at all. Example: the temperature in a city, or the final cross-entropy loss
2.31. To grab it you point once; you need 0 indices. - Vector — rank 1. A list of numbers in a row. Example: the 3,072-dimensional embedding of the word "cat" in a modern LLM. To grab one number you say "position 7" — 1 index.
- Matrix — rank 2. A grid with rows and columns. Example: a fully-connected layer's weight matrix, or a grayscale image. To grab one number you say "row 3, column 2" — 2 indices.
- Higher-rank tensor — rank 3+. A stack of grids. Example: an RGB image is a 3-D tensor (height × width × 3 color channels), so you say "row, column, channel" — 3 indices. A video is rank-4: time × height × width × channels — 4 indices.
Now the liberating part
In NumPy, PyTorch, and TensorFlow there is essentially one data type.
It's called ndarray / torch.Tensor / tf.Tensor. Scalars, vectors, and matrices are not separate types — they are just tensors of rank 0, 1, and 2. So you never learn four systems; you learn one container and ask two questions about it:
- shape — the size along each axis, like
(32, 3, 224, 224). - dtype — what kind of number lives inside (float32, float16, bfloat16, int8).
import numpy as np
scalar = np.array(3.14) # shape () rank 0
vector = np.array([1.0, 2.0, 3.0]) # shape (3,) rank 1
matrix = np.array([[1, 2], [3, 4], [5, 6]])# shape (3, 2) rank 2
batch = np.zeros((32, 768, 768)) # shape (32, 768, 768) rank 3
print(vector.ndim, matrix.shape) # 1 (3, 2)03.A Tiny Worked Example: Reading a Matrix by Hand
Take the matrix from the code above.
matrix = [[1, 2], [3, 4], [5, 6]]
Its shape is (3, 2). Read that as: 3 rows, 2 columns.
To point at one number you need exactly two coordinates — a row index and a column index.
- Row 0, column 0 →
1 - Row 0, column 1 →
2 - Row 2, column 1 →
6
Count the coordinates: 2. So this tensor has rank 2. That's all "rank" ever means — the number of coordinates in the address.
Now strip one axis away.
- Drop the row axis: you have a single number, like
4. That's a scalar (rank 0). - Keep only rows, drop columns: you get a list, like
[1, 2]. That's a vector (rank 1).
The ladder is just this
- rank 0 → one number
- rank 1 → a list of numbers →
stack N numbers - rank 2 → a grid of lists →
stack N rows - rank 3+ → a stack of grids →
stack N slices
Each level is simply an ordered collection of the level below it. That single sentence is the whole rank ladder.
04.Visual Intuition: The Rank Ladder
Picture the ladder going up. Each box is one number; stacking boxes builds the next rank.
coderank 0 SCALAR [ 3.14 ] one number, no address rank 1 VECTOR [ 1 | 2 | 3 ] one row -> needs 1 index rank 2 MATRIX [ 1 | 2 ] a grid -> needs 2 indices [ 3 | 4 ] (row, col) [ 5 | 6 ] rank 3 TENSOR a STACK of matrices -> needs 3 indices slice 0 slice 1 ... (i, j, k)
Or, the other direction: flatten anything back down the ladder.
codebatch of RGB images rank 4 (time, H, W, channel) | squeeze one axis v one RGB image rank 3 (H, W, channel) | take one channel v grayscale image rank 2 (H, W) <- a matrix | take one row v a scanline rank 1 (W,) <- a vector | take one pixel v one brightness rank 0 <- a scalar
The arrow → going up means "stack more". The arrow going down means "fix one coordinate". Every operation in ML is doing one of these two things.
05.The Analogy: A Blocking Facility
Carry one analogy through everything below: a building that stores parcels in rooms.
- A single parcel sitting on the floor = a scalar. No address needed; it's right there.
- A corridor of parcels, numbered 1..N = a vector. To name one you give one number: "corridor slot 7".
- A floor of corridors arranged in a grid = a matrix. One parcel now needs two numbers: "row 3, aisle 2".
- A whole tower of such floors = a higher-rank tensor. Now you need three or four numbers: "floor 2, row 3, aisle 2, shelf 1".
So
rankhow many numbers are in a parcel's address.
And
shapehow big each part of the building is.
A [32, 3, 224, 224] image batch is a tower with 32 identical floors, each floor a 3×224×224 arrangement. Keeping all parcels in one labeled building (one contiguous tensor) — instead of loose bags scattered across the lot (a Python list of separate vectors) — is what lets workers (GPUs) grab everything at once. We'll return to that exact point in section 7.
06.Shape Conventions in Real ML Systems
Every tensor in a neural network gives its axes semantic meaning — each axis "means" something. Misreading a convention is the single most common source of broadcasting and matmul bugs, so learn the layouts people actually ship.
- Computer vision (images): PyTorch uses
NCHW— (batch, channels, height, width); TensorFlow/Keras defaults toNHWC. A batch of 32 color 224×224 images is[32, 3, 224, 224]in PyTorch. Same data, different axis order. - Transformers / LLMs (2024–2026): activations are typically
(batch, seq_len, d_model), e.g.[8, 4096, 4096]. Inside attention, the projection is split into heads:(batch, num_heads, seq_len, head_dim)— a rank-4 tensor. - Weight matrices: a linear layer mapping
d_in → d_outstores a matrix of shape(d_out, d_in)in PyTorch (nn.Linear), which surprises people coming from math textbooks that write(d_in, d_out).
None of this is new math — it is the "address" idea from the analogy, now wearing production names.
07.Why AI Cares: Batch Dimension and Numerical Footprint
Two practical facts turn "shapes are cute" into "shapes decide whether my GPU runs".
Why the batch dimension exists. A single training example is rarely processed alone. GPUs are throughput machines: they are fastest when the same operation is applied to thousands of rows in parallel. Prepending a batch axis lets one tensor hold 32 or 4,096 examples, and frameworks use broadcasting so a bias vector of shape (d,) can be added to an activation tensor of shape (B, T, d) without being copied.
Concretely: an embedding lookup for a batch of 8 sequences, each 4,096 tokens long, with hidden size 4,096, produces a rank-3 tensor of shape (8, 4096, 4096) ≈ 134 million floats (~537 MB in float32). Keeping that data as one contiguous tensor — the "one building" from the analogy, instead of a Python list of 32,768 separate vectors — is what makes GPU kernels and KV-cache management tractable at all.
Dtypes and numerical footprint. The scalar type has become a first-class systems concern in AI:
- float32: the classic default; 4 bytes per scalar.
- bfloat16 / float16: 2 bytes per scalar; standard for LLM pre-training and inference in 2024–2026, roughly doubling throughput on modern accelerators.
- int8 / FP4–FP8: used for quantized inference, packing 4× or 8× more weights into the same memory bandwidth.
A model with 70 billion parameters stores 70×10⁹ scalars. At bfloat16 that is 140 GB — a rank-N tensor problem wearing a hardware disguise. Do it in one line of intuition
70 × 10⁹ params × 2 bytes/param = 140 GB
That is why interviewers ask you to compute tensor memory: elements × bytes-per-element.
Architectural Trade-offs & Production Realities
Architectural Advantages
- A single tensor abstraction covers data, weights, gradients, and activations uniformly.
- Contiguous high-rank tensors enable vectorized SIMD/GPU execution and memory-mapped checkpoints.
- Shape + dtype metadata makes whole models analyzable (parameter counts, FLOPs, VRAM) without touching values.
Trade-offs & Constraints
- Broadcasting is powerful but hides silent bugs when a missing size-1 axis auto-expands.
- Axis-order conventions (NCHW vs NHWC, d_out×d_in) differ across frameworks and cost real debugging time.
- Large dense tensors can force copies/reshapes that stress memory bandwidth more than compute.
A GPT-style forward pass is a sequence of shape-preserving tensor ops: token IDs (batch, seq) → embeddings (batch, seq, d_model) → attention reshaped to (batch, heads, seq, head_dim) → logits (batch, seq, vocab). Checkpoint files (Safetensors) are literally named tensors plus their shape and dtype metadata.
Staff+ Engineering Takeaways
- Rank = number of indices needed to address one scalar: 0 = scalar, 1 = vector, 2 = matrix, 3+ = tensor.
- Scalars, vectors, and matrices are just special cases of the tensor type used by every ML framework.
- Shape is semantics: vision uses NCHW/NHWC, transformers use (batch, seq, d_model) and (batch, heads, seq, head_dim).
- The batch dimension exists to feed GPUs throughput-parallel work; broadcasting avoids copying biases and scalars.
- dtype (float32, bfloat16, int8) determines memory and speed as much as shape does in modern inference.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
What is the rank of the activation tensor entering an attention block in a standard transformer, before head splitting?
How clear and actionable was this distributed systems breakdown?