PidokuInfra

Vectors and Matrices

Foundations Beginner 1h Difficulty 1/5 Topic 01 of 12

Prerequisites I.04


1. What is it?#

  • A vector is an ordered list of numbers: [3, -1, 4].
  • A matrix is a rectangular grid of numbers.
  • Both are just arrays. What makes them useful is the operations defined on them.

In inference, essentially every number you touch is an element of a vector or matrix, and essentially every operation is a matrix multiply or an elementwise function.


2. Why does it exist?#

Because they let you express “apply this transformation to all of these things” as a single object, which (a) is compact to write and (b) maps directly onto hardware that does many identical operations at once.

A neural network layer is literally “multiply by a matrix, add a vector, apply a function elementwise.” If you understand those three operations you can read any model.


3. Simple analogy#

A vector is a recipe’s ingredient list; a matrix is a book of recipes.

A vector [2, 1, 3] might mean 2 eggs, 1 cup flour, 3 tbsp sugar.

A matrix’s rows are different recipes, its columns the same ingredients. Multiplying the matrix by a price vector gives you the cost of every recipe at once — one operation, many answers. That “one operation, many answers” property is exactly what neural networks exploit and exactly what GPUs are built for.


4. Tiny example#

Vector:        v = [3, -1, 4]                   shape (3,)

Matrix:        M = [ 1  2  3 ]                  shape (2, 3)
                   [ 4  5  6 ]                  2 rows, 3 columns

Element:       M[1][2] = 6      (row 1, column 2, zero-indexed)
Row:           M[0]    = [1, 2, 3]
Column:        M[:,1]  = [2, 5]

Operations you need:

Addition (elementwise, same shape):
  [1,2,3] + [4,5,6] = [5,7,9]

Scalar multiply:
  2 * [1,2,3] = [2,4,6]

Dot product (two vectors → one number):
  [1,2,3] · [4,5,6] = 1·4 + 2·5 + 3·6 = 4+10+18 = 32

Hadamard / elementwise product (two vectors → vector):
  [1,2,3] ⊙ [4,5,6] = [4,10,18]

Transpose (flip rows and columns):
  [1 2 3]ᵀ  =  [1 4]
  [4 5 6]      [2 5]
               [3 6]

Norm (length):
  ||[3,4]|| = sqrt(9+16) = 5

The dot product is the single most important operation in this curriculum. It is a similarity measure: large and positive when the vectors point the same way, zero when perpendicular, negative when opposed. Attention (file 08) is built entirely out of dot products between query and key vectors.


5. Technical explanation#

What a matrix does#

A matrix is a linear transformation: it takes a vector in and produces a vector out.

M is (2, 3);  v is (3,);  Mv is (2,)
   → M maps 3-dimensional vectors to 2-dimensional ones

For neural networks, read W @ x as: “project x from its space into a new space where each output dimension is a different learned weighted combination of the inputs.” Row i of W is the recipe for output feature i.

This is why the shapes tell you so much:

W_q : (d, d)         → same-dimensional projection (query space)
W_up: (d, d_ff)      → project up to a wider space
W_down: (d_ff, d)    → project back down
embedding: (V, d)    → each of V tokens gets a d-dimensional vector
lm_head: (d, V)      → project a hidden state onto vocabulary scores

Dot product as similarity#

For unit-length vectors, a · b = cos(θ). In attention:

q · k  large  →  "this query is looking for what this key offers"
q · k  small  →  irrelevant

Scaling matters: with random vectors of dimension d, the dot product has variance proportional to d, so it grows like sqrt(d). That’s why attention divides by sqrt(d_head) — without it, the softmax input has huge magnitude and saturates (file 07).

Notation conventions you’ll meet#

Math papers:      y = Wx        (column vectors, weight on the left)
Code (PyTorch):   y = x @ W.T   or   nn.Linear stores weight as (out, in)
Batched code:     Y = X @ W.T   where X is (batch, in), Y is (batch, out)

nn.Linear(in_features, out_features) stores weight with shape (out_features, in_features) and computes x @ weight.T + bias. Knowing this saves you from a category of confusing bugs when reading or writing custom kernels.


6. Under the hood#

A matrix in memory is a flat buffer plus a shape and strides (Section I.04):

M = [[1,2,3],[4,5,6]]  stored row-major:  [1, 2, 3, 4, 5, 6]
strides = (3, 1)   → M[i][j] is at offset 3i + j

Row-major (C order) is PyTorch/NumPy default. Column-major (Fortran order) is what BLAS historically expects, which is why cuBLAS calls in framework code look transposed. This is a frequent source of confusion when reading kernel code — the transposes are usually layout bookkeeping, not mathematics.


7. Performance implications#

  • Dot products of length d cost 2d FLOPs and read 2×d×bytes. Intensity ≈ 1/bytes_per_elem. Pure memory-bound. A single dot product is the worst possible use of a GPU.
  • Batching dot products into a matrix multiply is what makes them efficient (Section I.06).
  • Transpose is free in metadata, expensive in memory if a kernel then requires contiguity.
  • Norms and elementwise ops are memory-bound — always fuse them into neighbors when you can.

8. Production implications#

You will read model code and kernel code constantly. Fluency with shapes is the difference between “this file is inscrutable” and “oh, this projects to the KV space with 8 heads.”

Practical habit: when reading unfamiliar model code, annotate every line with its output shape. Do it in a comment. It takes two minutes and turns confusion into understanding.


9. Common mistakes#

Confusing dot product with elementwise product. a · b is a scalar; a ⊙ b is a vector.

Getting the transpose wrong in nn.Linear. Weight is (out, in), computation is x @ W.T.

Forgetting the sqrt(d) scaling in attention. Produces saturated softmax and, at inference, near-deterministic garbage.

Assuming matrix multiplication commutes. AB ≠ BA, and often the shapes don’t even allow both.


10. Hands-on exercise#

A. By hand, no computer. Given

A = [1  2]     B = [ 3  0  1]     v = [2]
    [0 -1]         [-1  2  4]         [5]

compute: Av, AB, Bᵀ, v · v, ||v||. Check with a few lines of Go afterwards.

B. Shape detective. Load any small model in PyTorch and print [(n, tuple(p.shape)) for n, p in model.named_parameters()]. For each parameter, say in words what transformation it performs and between which spaces.

C. Dot product as similarity. Take 5 word embeddings from any pretrained model. Compute all pairwise cosine similarities. Do the results match your intuition about the words?

D. The sqrt(d) argument. Generate random vectors of dimension d ∈ {64, 512, 4096} with entries ~N(0,1). Measure the standard deviation of their dot products. Confirm it scales as sqrt(d). Now explain why attention divides by sqrt(d_head).


11. Interview questions#

  1. What does a matrix multiplication represent geometrically?
  2. Why does attention scale by 1/sqrt(d_head)?
  3. What shape does nn.Linear(768, 3072) store its weight in, and what does it compute?
  4. Why is a single dot product a terrible use of a GPU?
  5. What is the difference between row-major and column-major, and where does it bite you?

12. Further reading#

  • [FUNDAMENTAL] 3Blue1Brown, Essence of Linear Algebra — watch videos 1-4 and 9
  • [REFERENCE] NumPy and PyTorch broadcasting/linear algebra docs
  • Next: 02 — Matrix multiplication by hand

↑↓ navigate↵ openesc close