PidokuInfra
03Machine Learning Fundamentals
On this page

Machine Learning Fundamentals

FoundationsModule 0312 topics~15h 30m

Topics, in order

01 Vectors and Matrices- A vector is an ordered list of numbers: [3, -1, 4]. - A matrix is a rectangular grid of numbers. - Both are just arrays. What makes them useful is the operations defined on them. Beginner 1h 02 Matrix Multiplication by HandMatrix multiplication is 90-95% of the arithmetic in LLM inference. Everything in Sections VI, VII, and IX is, ultimately, about making this one operation faster. Do the exercises by hand. Beginner 1h 15m 03 Tensors, Shapes, and BroadcastingA tensor is an n-dimensional array. Broadcasting is the rule that lets operations between different-shaped tensors work without explicit copying. Beginner 1h 04 Neural Networks from ScratchA neural network is a stack of alternating linear transformations and nonlinear functions: Beginner 1h 30m 05 Activation FunctionsAn elementwise nonlinear function applied to every value in a tensor. No parameters (usually), no mixing between elements. Beginner 45 min 06 EmbeddingsAn embedding is a learned vector that represents a discrete thing — a token, a word, a user, a product. The embedding table is a matrix of shape (V, d): one row per vocabulary item. Beginner 1h 07 Softmax and Numerical StabilityThe "online softmax" trick in section 6 is the mathematical core of FlashAttention. If you understand it here, Section VII.10 will be easy. Intermediate 1h 15m 08 Attention from First PrinciplesThis is the most important file in Section III. Everything in Section V depends on understanding it mechanically, not just conceptually. Intermediate 2h 09 The TransformerA transformer is attention and an MLP, stacked with residual connections and normalization, repeated L times. Intermediate 2h 10 Training vs Inference: The MathA precise accounting of what training does that inference doesn't, in FLOPs and bytes. File I.02 gave the operational differences; this gives the arithmetic. Intermediate 45 min 11 Number Formats: FP32, FP16, BF16, FP8, INT8, INT4How a number is encoded in bits. The choice determines memory footprint, memory bandwidth, arithmetic throughput, and how much precision you lose. Intermediate 1h 30m 12 Quantization FundamentalsThis file teaches the mechanism. Section VII covers the specific production methods (GPTQ, AWQ, SmoothQuant, FP8) and when to use each. Intermediate 1h 30m

About this module

Goal: teach exactly the ML you need to reason about inference — no more, no less.

This is not an ML course. There is no training theory, no optimization, no generalization bounds. What is here: the linear algebra you will do by hand, the architecture you will trace through, and the numerics you will trade away for speed.

If you already know ML, do not skip files 07, 11, and 12 — softmax stability, number formats, and quantization are where ML people most often have gaps that matter for inference.

Files#

#FileLevelTime
01Vectors and matricesBeginner60 min
02Matrix multiplication by handBeginner75 min
03Tensors, shapes, and broadcastingBeginner60 min
04Neural networks from scratchBeginner90 min
05Activation functionsBeginner45 min
06EmbeddingsBeginner60 min
07Softmax and numerical stabilityIntermediate75 min
08Attention from first principlesIntermediate120 min
09The transformerIntermediate120 min
10Training vs inference: the mathIntermediate45 min
11Number formats: FP32, FP16, BF16, FP8Intermediate90 min
12Quantization fundamentalsIntermediate90 min

The thread#

flowchart TD
  N0["Numbers in arrays<br/><b>01, 03</b>"]
  N1["Multiplied by learned arrays<br/><b>02</b>"]
  N2["Stacked with nonlinearity between them<br/><b>04, 05</b>"]
  N3["Text becomes vectors<br/><b>06</b>"]
  N4["Scores become probabilities<br/><b>07</b>"]
  N5["Tokens look at each other<br/><b>08</b>"]
  N6["Assembled into a transformer<br/><b>09</b>"]
  N7["Which we run forward only<br/><b>10</b>"]
  N8["In as few bits as we can get away with<br/><b>11, 12</b>"]
  N0 --> N1 --> N2 --> N3 --> N4 --> N5 --> N6 --> N7 --> N8

  class N0,N1 neutral
  class N2,N3 io
  class N4,N5 queue
  class N6,N7 compute
  class N8 memory

Checkpoint B#

  1. Multiply a 2×3 by a 3×2 matrix by hand, showing every dot product.
  2. Explain attention in three sentences with no equations.
  3. Why is softmax computed with a max subtraction? What breaks without it?
  4. What does BF16 trade away relative to FP16, and why does that trade matter?
  5. Write the parameter count formula for a decoder-only transformer and apply it to a real model.
  6. What is the difference between per-tensor and per-channel quantization, and why does the latter usually win?

↑↓ navigate↵ openesc close