PidokuInfra

The Memory Model and Data Races

Intermediate Advanced 50 min Difficulty 4/5 Topic 05 of 06

Prerequisites 02, 04

The idea in one minute#

A data race is two goroutines accessing the same variable at the same time, at least one writing, with nothing ordering them. The Go memory model is the specification of what “ordering them” means: a read is guaranteed to see a write only if the write happens before the read through a chain of synchronization — a channel operation, a mutex, an atomic, a sync.Once, a goroutine start.

Without that chain, the compiler and CPU are free to reorder, cache in a register, or tear the access. There is no such thing as a benign data race. The race detector (-race) finds them at run time, and it should be on in every test run.

An analogy#

Two people editing the same paper document by post, with no agreement about whose copy is current. Each works from whatever arrived last. The result is not “the latest edit wins” — it is an unpredictable mix, and sometimes a page of one version stapled to a page of another. A synchronization point is a meeting where one hands the document to the other: after it, both agree on what the document says.

A picture#

flowchart LR
  subgraph G1["goroutine A"]
    A1["x = 42"] --> A2["ch <- struct{}{}"]
  end
  subgraph G2["goroutine B"]
    B1["<-ch"] --> B2["print(x)  sees 42"]
  end
  A2 -->|"a send happens before<br/>the matching receive completes"| B1
  subgraph R["No edge: a data race"]
    C1["goroutine C: y = 1"]
    C2["goroutine D: print(y)<br/>may see 0, may see 1,<br/>may never see the write"]
  end
  class A1,A2,B1,B2 compute
  class C1,C2 warn

How it really works#

Happens-before#

Within one goroutine, statements happen in program order. Between goroutines, only these events create an ordering:

EventGuarantee
go f()The go statement happens before f starts
Channel sendHappens before the corresponding receive completes
Unbuffered receiveHappens before the corresponding send completes
close(ch)Happens before a receive that returns because the channel is closed
mu.Unlock()Happens before the next mu.Lock() returns
once.Do(f)The single run of f happens before any Do returns
Atomic operationsBehave as if executed in one total order; a Load sees the latest Store
wg.Done() / wg.Wait()The final Done happens before Wait returns
Goroutine exitNo guarantee. Nothing is ordered after a goroutine merely finishing

If a write and a read of the same variable are not connected by such a chain, they are concurrent, and the program has a data race.

What actually goes wrong#

  • Lost updates. n++ is load, add, store. Two goroutines interleave and one increment vanishes (IV.04’s first counter).
  • Stale reads, forever. A loop for !done { } may be compiled to read done once into a register. The flag set by another goroutine is never seen; the loop never ends.
  • Torn values. An interface, slice or string is several words. A reader can observe the pointer of one value with the length of another — and then index past the end of memory. This is how a data race becomes memory corruption in a “memory-safe” language.
  • Broken invariants. A map being written while read can crash the runtime (concurrent map read and map write).
  • Reordering. The compiler and the CPU may reorder writes that have no dependency. On weakly ordered CPUs such as ARM64 — which is what many GPU servers and laptops now are — far more reorderings are visible than on x86. Code that “worked” on one may fail on the other.

The race detector#

Shell
go test -race ./...
go run -race .
go build -race -o app-race .

It instruments every memory access and tracks, per memory location, which goroutines touched it under which synchronization. When it observes two unsynchronized accesses it prints both stack traces and where the goroutines were created.

  • It reports only races that actually occur in that run: no false positives, but coverage depends on your tests and load.
  • Cost: roughly 5–10× CPU and memory. Run it in CI, in integration tests, and for a canary under real traffic if you can afford it — not usually fleet-wide.
  • A report is always a real bug. Fix it; do not argue with it.

Frequent races and their fixes#

RaceFix
A shared counter or flagatomic type
A map written from several goroutinesMutex or sync.Map
Appending to a shared sliceMutex; or give each goroutine its own index: results[i] = ...
Lazy initialization with an if x == nil checksync.Once / sync.OnceValue
Reading a config while another goroutine reloads itatomic.Pointer to an immutable value
A test goroutine calling t.Fatal after the test endedWait for goroutines before returning
Sharing rand.Rand, bytes.Buffer, a tokenizer with internal stateOne per goroutine, or a lock

Writing to different elements of a slice from different goroutines is not a race: results[i] for distinct i are distinct variables. That makes “preallocate, each worker fills its own slot” the simplest correct way to collect parallel results.

Double-checked locking#

Go
if instance == nil {            // unsynchronized read: a data race
    mu.Lock()
    if instance == nil { instance = build() }
    mu.Unlock()
}
return instance

This pattern is wrong in Go: the first read is not ordered after the write, so a goroutine may see a non-nil pointer to an object whose fields are not yet visible. Use sync.Once, which does the fast-path check with an atomic and is correct.

Things that are safe#

  • Reading from many goroutines data that nobody writes any more — provided it was published through a synchronization event (passed on a channel, stored before the goroutines started).
  • Immutable values: build completely, then share.
  • Confinement: data only one goroutine ever touches needs no synchronization.

Code#

Go
// races.go — a stale read that never ends (shown safely), lost updates, and correct publication.
package main

import (
	"fmt"
	"sync"
	"sync/atomic"
	"time"
)

type Config struct {
	Model   string
	Timeout time.Duration
}

func main() {
	// 1. Publication through a synchronization event: always correct.
	var cfg *Config
	ready := make(chan struct{})
	go func() {
		cfg = &Config{"large", time.Second} // written before the close...
		close(ready)
	}()
	<-ready // ...which happens before this receive returns
	fmt.Println("published via channel:", cfg.Model)

	// 2. Each goroutine writes its own slice element: no race, no lock.
	results := make([]int, 8)
	var wg sync.WaitGroup
	for i := range results {
		wg.Add(1)
		go func() {
			defer wg.Done()
			results[i] = i * i
		}()
	}
	wg.Wait() // the Done calls happen before Wait returns, so reading results is safe
	fmt.Println("disjoint slice elements:", results)

	// 3. A flag: the atomic version is guaranteed to be seen.
	var stop atomic.Bool
	spins := 0
	go func() {
		time.Sleep(5 * time.Millisecond)
		stop.Store(true)
	}()
	for !stop.Load() {
		spins++
	}
	fmt.Println("atomic flag: the loop ended after", spins, "spins")
	fmt.Println("   (with a plain bool the compiler may hoist the read out of the loop: an endless loop)")

	// 4. Lost updates on a shared int versus an atomic.
	var plain int64
	var safe atomic.Int64
	for g := 0; g < 8; g++ {
		wg.Add(1)
		go func() {
			defer wg.Done()
			for i := 0; i < 100000; i++ {
				plain++ // DATA RACE: run with -race to see the report
				safe.Add(1)
			}
		}()
	}
	wg.Wait()
	fmt.Printf("8 x 100,000 increments: plain=%d  atomic=%d\n", plain, safe.Load())

	// 5. Lazy initialization done right.
	loads := 0
	model := sync.OnceValue(func() *Config {
		loads++
		return &Config{"lazy", time.Second}
	})
	for g := 0; g < 16; g++ {
		wg.Add(1)
		go func() { defer wg.Done(); _ = model() }()
	}
	wg.Wait()
	fmt.Println("OnceValue: 16 goroutines asked, the loader ran", loads, "time")
}

Run it once normally, then with go run -race races.go and read the report for case 4.

Remember this#

  • A data race is unsynchronized concurrent access with at least one write. Its behaviour is undefined, not merely “sometimes stale”.
  • Ordering between goroutines comes only from channel operations, locks, atomics, Once, WaitGroup and goroutine start.
  • Run tests with -race. Every report is a real bug.
  • Distinct slice elements, immutable data, and confinement are safe without locks.

Try it#

  1. Run races.go with -race. Which lines does the report name?
  2. Replace the atomic flag in case 3 with a plain bool and build with optimizations. Does the loop end? (Try it on ARM and on x86 if you can.)
  3. Write the broken double-checked lock and a test that hammers it with -race.

Check yourself#

  1. Give three events that establish a happens-before edge between goroutines.
  2. Why is a race on an interface or slice value more dangerous than one on an int?
  3. Why is writing results[i] from goroutine i not a data race?

↑↓ navigate↵ openesc close