PidokuInfra

Activation Functions

Foundations Beginner 45 min Difficulty 1/5 Topic 05 of 12

Prerequisites 04


1. What is it?#

An elementwise nonlinear function applied to every value in a tensor. No parameters (usually), no mixing between elements.

ReLU(x)  = max(0, x)
GELU(x)  = x · Φ(x)                 Φ = standard normal CDF
SiLU(x)  = x · sigmoid(x)           also called Swish
SwiGLU(x, y) = SiLU(x) · y          a *gated* variant, uses two inputs

2. Why does it exist?#

To break linearity (file 04). Beyond that, the specific choice affects trainability and, for us, kernel cost.

For inference the practical questions are: how expensive is it, can it be fused, and does its numerical range cause problems in low precision?


3. Simple analogy#

A dimmer switch versus an on/off switch.

ReLU is a hard switch: below zero, nothing passes; above, everything passes proportionally. Simple, fast, and information below zero is destroyed entirely.

GELU/SiLU are dimmers: near zero they pass a fraction, smoothly. Slightly more expensive, empirically better for transformers because the smooth transition preserves small signals.

Gated variants (SwiGLU) add a second control: one signal decides how much of the other passes. Like a valve whose opening is itself computed from the data.


4. Tiny example#

Go
package main

import (
	"fmt"
	"math"
)

func relu(x float64) float64    { return math.Max(0, x) }
func sigmoid(x float64) float64 { return 1 / (1 + math.Exp(-x)) }
func silu(x float64) float64    { return x * sigmoid(x) }
func geluTanh(x float64) float64 {
	return 0.5 * x * (1 + math.Tanh(math.Sqrt(2/math.Pi)*(x+0.044715*x*x*x)))
}

func main() {
	xs := []float64{-2.0, -0.5, 0.0, 0.5, 2.0}
	for name, f := range map[string]func(float64) float64{"relu": relu, "silu": silu, "gelu": geluTanh} {
		fmt.Printf("%-5s", name)
		for _, x := range xs {
			fmt.Printf(" %7.3f", f(x))
		}
		fmt.Println()
	}
	// relu    0.000   0.000   0.000   0.500   2.000
	// silu   -0.238  -0.189   0.000   0.311   1.762
	// gelu   -0.045  -0.154   0.000   0.345   1.955
}

Note SiLU and GELU are negative for small negative inputs — they don’t clip. That non- monotonicity near zero is part of why they work better.

Also note: x=-2 gives SiLU -0.238 and GELU -0.045. Different functions, meaningfully different behavior on the negative side.


5. Technical explanation#

What models actually use#

Model familyFFN activation
BERT, original GPT-2GELU
GPT-3GELU
Llama 1/2/3, Mistral, QwenSiLU, in SwiGLU form
PaLM, GemmaGeGLU (GELU-gated)
Older CNNsReLU

Modern LLMs almost universally use a gated variant. The FFN becomes:

down( act(gate(x)) * up(x) )

three matrices instead of two, d_ff ≈ 8d/3 to keep parameters constant.

Cost analysis#

ReLU:     1 comparison, 1 select                     ~1 op
SiLU:     1 exp, 1 divide, 1 multiply                ~10-20 ops
GELU exact: 1 erf                                    ~20-40 ops
GELU tanh approx: 1 tanh, few mults                  ~15-25 ops

Sounds like a lot — but all of these are memory-bound anyway. For a (B,S,d_ff) tensor:

bytes = 2 · B·S·d_ff (read) + 2 · B·S·d_ff (write) = 4·B·S·d_ff
FLOPs ≈ 20 · B·S·d_ff  (for SiLU)

intensity = 20/4 = 5 FLOP/byte     ← still far below the ridge point of ~296

So an unfused activation costs you a full read+write of the tensor regardless of which function it is. The choice of activation is nearly free; the choice to fuse it is not.

Fused into the preceding GEMM’s epilogue, the activation costs approximately zero: the values are already in registers.

The tanh approximation of GELU#

# exact
gelu(x) = 0.5 * x * (1 + erf(x / sqrt(2)))
# tanh approximation (what most implementations use)
gelu(x) ≈ 0.5 * x * (1 + tanh(sqrt(2/pi) * (x + 0.044715 * x^3)))

Max error ~1e-3. Faster on hardware without a fast erf. These are not bit-identical, so a model trained with one and served with the other has a small distribution shift. Usually harmless; occasionally the explanation for “why does my ported model score slightly differently?”

Low-precision considerations#

FP16 max ≈ 65,504

Activations after a large FFN up-projection can be large. SiLU/GELU are roughly identity for large positive x, so they don’t amplify — but the product in SwiGLU (act * up) can. In FP16 this occasionally overflows to inf. Mitigations: compute the gate in FP32, use BF16 (huge range), or clamp. This is a real source of NaN bugs when quantizing (Section VII).


6. Under the hood#

In a fused kernel, the activation is in the GEMM epilogue:

... tensor core accumulation into registers ...
epilogue:
    acc = acc * scale + bias            # in registers
    acc = silu(acc)                     # in registers — FREE
    store acc to HBM                    # the only memory traffic

Compare to unfused:

GEMM: write acc to HBM
activation kernel: read from HBM, compute, write to HBM
→ 2 extra full-tensor traversals

For a (2048, 14336) FP16 tensor that’s 2×59 MB = 118 MB of avoidable traffic per matrix per layer. At 32 layers and 3 FFN matrices: several GB per prefill. This is why torch.compile, TensorRT, and every serious engine fuse epilogues.


7. Performance implications#

  • Unfused activations cost a full memory round trip. Always fuse.
  • The specific function barely matters for speed once fused.
  • Gated activations require an extra elementwise multiply and a third matrix — but the matrices are smaller, so total cost is similar.
  • Fusing gate+up into one GEMM (concatenate the weight matrices) turns two GEMMs into one and improves efficiency, then splits the output in the epilogue.

8. Production implications#

  • Verify your engine fuses activations. Check the Nsight kernel list: if you see a standalone silu_kernel between GEMMs, you’re leaving performance on the table.
  • Match the activation exactly to the reference implementation. GELU-exact vs GELU-tanh, or SiLU vs GELU, will silently shift outputs.
  • Watch for overflow in FP16 SwiGLU. BF16 avoids it.

9. Common mistakes#

Substituting GELU for SiLU (or exact for tanh). Numerically different; quality shifts.

Leaving activations unfused. A common 5-15% loss.

Assuming activations are compute-bound. They aren’t; they’re bandwidth.

Forgetting SwiGLU has three matrices when counting parameters or FLOPs.


10. Hands-on exercise#

A. Plot them. Plot ReLU, GELU (exact and tanh), SiLU on [-5, 5]. Plot the difference between exact and tanh GELU. Where is it largest?

B. Measure fusion. Time (x @ W); silu(result) as two ops vs a torch.compiled fused version. Report the speedup and explain it in terms of bytes moved.

C. Overflow hunt. Construct FP16 tensors where SwiGLU overflows to inf. At what input magnitudes? Repeat in BF16 and explain the difference.

D. Substitution experiment. Take a small model that uses SiLU, replace it with GELU, and measure perplexity change on a small text sample. How much does the wrong activation cost?


11. Interview questions#

  1. Why do neural networks need activation functions?
  2. What is SwiGLU and why do modern LLMs use it?
  3. Is the choice of activation function a significant performance decision? Explain.
  4. What is kernel fusion in the context of activations, and how much does it save?
  5. Why might FP16 SwiGLU overflow, and what would you do about it?
  6. What is the difference between exact and tanh-approximated GELU, and when does it matter?

12. Further reading#

  • [ESTABLISHED] Shazeer, “GLU Variants Improve Transformer” (2020)
  • [ESTABLISHED] Hendrycks & Gimpel, “Gaussian Error Linear Units” (2016)
  • [REFERENCE] PyTorch nn.functional activation docs; torch.compile fusion notes
  • Next: 06 — Embeddings

↑↓ navigate↵ openesc close