PidokuInfra

Model Anatomy: Parameters, Weights, Activations

Foundations Beginner 1h Difficulty 2/5 Topic 03 of 11

Prerequisites 01, 02


1. What is it?#

Open a model file and you find three categories of numbers. Two live on disk; one exists only while the model runs.

ThingLives whereChanges at inference?Size scales with
Parameters / weightsdisk → GPU memory, permanently residentNomodel size
ActivationsGPU memory, transientYes, every requestbatch × sequence
KV cache (generative models)GPU memory, per-sequenceGrows every tokenbatch × context × model shape
  • Parameter — any learned number in the model. Umbrella term.
  • Weight — a parameter used to multiply an input. The bulk of parameters.
  • Bias — a parameter that is added. A small minority; many modern LLMs drop them.
  • Activation — a value computed during the forward pass. Not stored in the model file.
  • KV cache — stored activations from attention, kept deliberately across decode steps (Section V; the most important special case in the whole curriculum).

Getting these straight matters because each one has a completely different memory and bandwidth behavior, and confusing them makes capacity planning impossible.


2. Why does it exist?#

Because “how much memory does my model need?” has three answers that add up, and engineers who only know the first one get paged at 3 a.m.

Total GPU memory = weights + activations + KV cache + framework overhead
                    (fixed)   (∝ batch)   (∝ batch×tokens)   (~1-2 GB)

Weights are the number everyone quotes. KV cache is the number that actually causes outages, because it is the only term that grows during a request and is unbounded by anything except your context limit.


3. Simple analogy#

A restaurant kitchen.

  • Weights are the recipes and the installed equipment. Fixed. Take up a known amount of space. Present whether or not anyone orders.
  • Activations are the dishes being prepared right now on the counters. They appear, get used, and are cleared away immediately. More simultaneous orders → more counter space.
  • KV cache is the running tab for each table: everything they’ve ordered so far, which you must keep because the next course depends on it. Grows all evening. If tables stay long enough, you run out of clipboards — and you can’t seat anyone new. Sound familiar? That’s exactly how an LLM server dies under load.

4. Tiny example#

A 2-layer network, fully written out.

Go
package main

import "fmt"

// matVec computes W·x + b for a rows×cols matrix stored as [][]float64.
func matVec(W [][]float64, x, b []float64) []float64 {
	out := make([]float64, len(W))
	for i, row := range W {
		sum := b[i]
		for j, w := range row {
			sum += w * x[j]
		}
		out[i] = sum
	}
	return out
}

func relu(v []float64) []float64 {
	out := make([]float64, len(v))
	for i, x := range v {
		out[i] = max(0, x)
	}
	return out
}

func main() {
	// ---------- WEIGHTS (learned, fixed) ----------
	W1 := [][]float64{ // 3x3 = 9 params
		{0.5, -0.2, 0.1},
		{0.3, 0.8, -0.4},
		{-0.1, 0.6, 0.2},
	}
	b1 := []float64{0.1, -0.1, 0.0}     // 3 params
	W2 := [][]float64{{1.0, -1.0, 0.5}} // 1x3 = 3 params
	b2 := []float64{0.2}                // 1 param
	// total: 16 parameters

	// ---------- INPUT ----------
	x := []float64{1.0, 2.0, 3.0}

	// ---------- FORWARD PASS ----------
	z1 := matVec(W1, x, b1)  // ACTIVATION (pre-nonlinearity)
	a1 := relu(z1)           // ACTIVATION (post-ReLU)
	z2 := matVec(W2, a1, b2) // ACTIVATION (output)

	fmt.Printf("z1 %.2f\n", z1)
	fmt.Printf("a1 %.2f\n", a1)
	fmt.Printf("z2 %.2f\n", z2)
}

Now count memory in FP32 (4 bytes each):

Weights:      16 params × 4 B = 64 bytes        ← same for every request, forever
Activations:  (3 + 3 + 1) × 4 B = 28 bytes      ← per request, freed immediately

Send 100 requests at once (batch 100):

Weights:      still 64 bytes        ← THIS IS THE WHOLE POINT OF BATCHING
Activations:  2,800 bytes

The weights are read once and used for all 100 inputs. Remember this; it is the mechanical reason batching works, and it recurs in file 06, file 08, and every batching discussion for the rest of the curriculum.


5. Technical explanation#

Where the parameters actually are in an LLM#

Take a decoder-only transformer with:

  • L layers, hidden size d, h attention heads, FFN intermediate size d_ff (often 4d), vocabulary V.

Per layer:

Attention:
  W_q : d × d          d²
  W_k : d × d          d²
  W_v : d × d          d²
  W_o : d × d          d²
                     ----
                      4d²

FFN (standard):
  W_up   : d × d_ff    d·d_ff
  W_down : d_ff × d    d_ff·d
                     ----------
                      2·d·d_ff   = 8d²  when d_ff = 4d

(SwiGLU FFN has 3 matrices: gate, up, down, with d_ff ≈ 8d/3, giving ~8d² again)

Norms: ~2d  (negligible)
                     ----------
Per layer total:     ~12d²

Plus embeddings:

Token embedding:  V × d
Output head:      V × d   (often tied with the embedding, i.e. shared)

So:

P ≈ 12·L·d² + 2·V·d

Check it against a real model — Llama 3 8B: L=32, d=4096, V=128256, GQA with 8 KV heads (which reduces W_k and W_v), SwiGLU with d_ff=14336.

Attention per layer: W_q 4096×4096 = 16.8M
                     W_k 4096×1024 = 4.2M   (8 KV heads × 128 = 1024)
                     W_v 4096×1024 = 4.2M
                     W_o 4096×4096 = 16.8M
                                    -------
                                     41.9M
FFN per layer:       3 × 4096×14336 = 176.2M
                                    -------
Per layer:                           218.1M
× 32 layers:                         6.98B
Embeddings: 2 × 128256×4096 =        1.05B
                                    -------
Total:                               8.03B   ✓

That is the actual composition of “8B.” Notice: the FFN is ~80% of the parameters. That is why FFN quantization and MoE (which touches the FFN) are such high-leverage optimizations.

Bytes per parameter#

PrecisionBytes/param8B model70B model405B model
FP32432 GB280 GB1.6 TB
FP16 / BF16216 GB140 GB810 GB
FP818 GB70 GB405 GB
INT40.54 GB35 GB203 GB

Memorize the FP16 row. params × 2 = GB is the estimate you will do in your head a hundred times.

Activations at inference#

For one token through one layer, the significant activations are O(d) and O(d_ff) values. Across the batch and the layer being executed, peak transient activation memory is roughly:

activation_peak ≈ batch × seq_processed × (a few × d_ff) × bytes

During decode, seq_processed = 1 per sequence, so activations are tiny — kilobytes to a few MB. During prefill with a long prompt, seq_processed can be 8,000, and activations become gigabytes. This is why prefill can OOM a server that decodes fine, and why chunked prefill (Section XIII.05) exists.

KV cache — introduced here, developed in Section V#

KV bytes = 2 (K and V) × L × n_kv_heads × head_dim × bytes_per_elem × tokens × batch

For Llama 3 8B (L=32, n_kv=8, head_dim=128, FP16):

per token per sequence = 2 × 32 × 8 × 128 × 2 = 131,072 bytes = 128 KiB

At 8,192 context: 1 GB per sequence. Batch 32: 32 GB — more than double the weights. This is the number that ends up governing your entire architecture.


6. Under the hood#

On disk. A safetensors file is a JSON header (tensor names, dtypes, shapes, byte offsets) followed by raw tensor bytes. You can inspect it without loading:

Go
// inspect.go — go run inspect.go model.safetensors
package main

import (
	"encoding/binary"
	"encoding/json"
	"fmt"
	"os"
	"sort"
)

type tensorInfo struct {
	DType string  `json:"dtype"`
	Shape []int64 `json:"shape"`
}

func main() {
	f, err := os.Open(os.Args[1])
	if err != nil {
		panic(err)
	}
	defer f.Close()

	// The first 8 bytes are a little-endian uint64: the length of the JSON header.
	var n uint64
	if err := binary.Read(f, binary.LittleEndian, &n); err != nil {
		panic(err)
	}
	header := make([]byte, n)
	if _, err := f.Read(header); err != nil {
		panic(err)
	}

	raw := map[string]json.RawMessage{}
	if err := json.Unmarshal(header, &raw); err != nil {
		panic(err)
	}
	delete(raw, "__metadata__")

	names := make([]string, 0, len(raw))
	for k := range raw {
		names = append(names, k)
	}
	sort.Strings(names)

	var total int64
	for _, k := range names {
		var t tensorInfo
		json.Unmarshal(raw[k], &t)
		count := int64(1)
		for _, s := range t.Shape {
			count *= s
		}
		total += count
		fmt.Printf("%-60s %-6s %-20v %12d\n", k, t.DType, t.Shape, count)
	}
	fmt.Printf("TOTAL %d parameters\n", total)
}

Run this on any model you can download. It converts “8B” from a marketing number into a concrete inventory, and it is the fastest way to learn a new architecture’s shape.

In memory. Weights are laid out as contiguous 2D arrays optimized for the GEMM kernel that will consume them (Section IV.07 covers layouts). They are read from HBM into the SMs’ registers and shared memory every time a matmul runs — which, in decode, is every token. That repeated reading is the bandwidth cost that dominates decode.

Reference counting. Activations are freed as soon as no operation needs them. The framework’s caching allocator (PyTorch’s CUDACachingAllocator) does not return the memory to the driver; it keeps it in a pool. That is why nvidia-smi shows high memory even when your tensors are freed, and why torch.cuda.empty_cache() exists (Section X.05).


7. Performance implications#

Weights dominate bandwidth in decode. Every decode step reads all weights from HBM. For a 70B FP16 model that is 140 GB read per step. On a 3.35 TB/s H100:

140 GB / 3350 GB/s = 41.8 ms per token, minimum, before any other cost

That is ~24 tokens/sec, and no amount of compute optimization improves it, because the arithmetic is not the bottleneck — the reading is. Quantizing to FP8 halves that time immediately. This calculation is the entire economic argument for quantization, and you should be able to reproduce it from memory.

Activations dominate memory in prefill. Long prompts, large batches → activation spikes. Mitigations: chunked prefill, FlashAttention (which removes the N² attention matrix from memory entirely), activation-light architectures.

KV cache dominates memory in steady state, and therefore dominates concurrency.


8. Production implications#

  • Capacity planning is a memory budget exercise. Write it down explicitly: 80 GB (H100) − weights − 2 GB overhead − activation headroom = KV budget, then KV budget ÷ bytes-per-token ÷ avg context = max concurrent sequences. Do this before choosing instance types, not after.
  • Quantization buys concurrency, not just speed. Going FP16 → FP8 on a 70B model frees 70 GB, which on an 8×H100 node roughly doubles the number of concurrent users you can hold.
  • Model surgery is legitimate. Tied embeddings, smaller vocabularies, GQA — these are weight-count and KV-count reductions with real production value.
  • Always report memory as a breakdown, never a single number. “We need 200 GB” is not actionable. “140 GB weights + 48 GB KV at our p99 context + 12 GB headroom” is.

9. Common mistakes#

“Weights are the memory requirement.” They are the floor. KV cache commonly exceeds them.

Counting parameters but not embeddings. For small models with large vocabularies, embeddings can be 20-30% of parameters. A “1B” model with a 256k vocabulary is mostly embedding table.

Assuming activations are negligible. True in decode, false in long-context prefill. Teams get bitten by exactly this: their service is stable for weeks, then someone pastes a 100-page document and the prefill activation spike OOMs the box.

Forgetting framework overhead. CUDA context, cuBLAS workspaces, NCCL buffers, the allocator’s fragmentation — budget 1-3 GB per GPU and more with tensor parallelism.

Confusing “model size on disk” with “memory needed.” A 4-bit GGUF file is 4 GB on disk but needs working buffers, dequantization scratch space, and KV cache on top.


10. Hands-on exercise#

A. Inventory a real model. Download any small open model (e.g. a 1B-class model) and run the safetensors script from section 6. Produce a table: attention params, FFN params, embedding params, norm params, and percentages. Does it match the 12Ld² + 2Vd formula?

B. Derive the memory budget. For Llama 3 70B (L=80, d=8192, n_kv_heads=8, head_dim=128) on 8×H100 (640 GB total):

  1. Weights at FP16, FP8, INT4.
  2. KV bytes per token per sequence.
  3. For each precision, how many concurrent sequences at 8k context fit in the leftover memory?
  4. Which precision would you choose and why?

C. Watch activations move. Run a small transformer with a 16-token prompt and then a 4096-token prompt, printing torch.cuda.max_memory_allocated() after each. Explain the difference quantitatively.


11. Interview questions#

  1. Break down where the parameters of a decoder-only transformer live. Which portion is largest?
  2. A 70B model at FP16 on 8×A100 (40 GB each). Does it fit? Show your work including overhead.
  3. Why does prefill sometimes OOM a server that decodes fine?
  4. Give the formula for KV cache size and explain every term.
  5. Your monitoring shows GPU memory at 78/80 GB constantly, even at low traffic. Is that a problem? How would you find out?

12. Further reading#

  • [REFERENCE] safetensors format spec
  • [ESTABLISHED] Llama 3 model card / paper for real architecture numbers
  • [ESTABLISHED] Kwon et al., PagedAttention (SOSP 2023), Section 2 — the memory breakdown figure
  • Next: 04 — Tensors and the forward pass

↑↓ navigate↵ openesc close