The idea in one minute#
When several goroutines touch the same data and at least one writes, access must be synchronized. Go gives you three tools at three price points. An atomic operation updates one word in a single indivisible CPU instruction: a nanosecond or two, no blocking. A mutex makes a stretch of code exclusive: a few nanoseconds when uncontended, and waiting goroutines are parked. A channel hands a value from one goroutine to another.
Pick by the shape of the problem: one number → atomic; a few fields that must change together → mutex; ownership moving between goroutines → channel.
An analogy#
A shared whiteboard. An atomic is a tally counter with a button — one press is one increment, nobody can interfere mid-press. A mutex is the single marker pen: whoever holds it writes, the rest wait their turn. A read-write lock lets any number of people read the board at once but clears the room when someone needs to write.
A picture#
flowchart TB
Q{"What is shared?"}
Q -->|"one counter, flag or pointer"| AT["sync/atomic<br/>atomic.Int64, atomic.Bool, atomic.Pointer[T]"]
Q -->|"several fields that must stay consistent"| MU["sync.Mutex<br/>lock, change, unlock"]
Q -->|"read very often, written rarely"| RW["sync.RWMutex,<br/>or atomic.Pointer to an immutable snapshot"]
Q -->|"initialize once"| ON["sync.Once / sync.OnceValue"]
Q -->|"wait for goroutines to finish"| WG["sync.WaitGroup"]
Q -->|"a value changes owner"| CH["channel"]
class Q queue
class AT,ON compute
class MU,RW memory
class WG,CH neutralHow it really works#
sync.Mutex#
type Cache struct {
mu sync.Mutex // guards the fields below it
items map[string]Item
hits int
}
func (c *Cache) Get(k string) (Item, bool) {
c.mu.Lock()
defer c.mu.Unlock()
it, ok := c.items[k]
if ok { c.hits++ }
return it, ok
}- The zero value is an unlocked mutex. Never copy one after first use (
go vet’scopylockscheck catches it); that is why types containing a mutex use pointer receivers. - A mutex is 8 bytes. Uncontended
Lock/Unlockis a single atomic compare-and-swap each — a few nanoseconds. Under contention a goroutine spins briefly, then parks. - It has a starvation mode: a goroutine that has waited more than 1 ms gets the lock handed to it directly, so no waiter is overtaken forever.
- Not reentrant: locking a mutex you already hold deadlocks.
- Keep critical sections short. No I/O, no channel operations, no calls into unknown code while holding a lock.
- When you need more than one lock, always acquire them in the same order.
sync.RWMutex#
Many concurrent readers (RLock) or one writer (Lock). It is not automatically faster:
every RLock still writes a shared counter, so with many cores and short critical sections the
cache line holding that counter becomes the bottleneck. It helps when reads are long or very
much more frequent than writes. For read-mostly configuration, an atomic.Pointer to an
immutable snapshot beats it: readers do one atomic load and never block.
sync/atomic#
var requests atomic.Int64
requests.Add(1)
n := requests.Load()
var cfg atomic.Pointer[Config]
cfg.Store(&Config{...}) // writer swaps in a whole new, never-modified value
c := cfg.Load() // readers get a consistent snapshotTyped atomics (Int32, Int64, Uint64, Bool, Pointer[T], Value) cannot be misused
with a plain read. CompareAndSwap(old, new) is the primitive behind lock-free algorithms:
“set it to new only if it is still old”.
An atomic protects one word. Two atomics updated one after another are not updated together; if an invariant spans several values, use a mutex.
sync.WaitGroup#
var wg sync.WaitGroup
for _, job := range jobs {
wg.Add(1) // before starting the goroutine, never inside it
go func() {
defer wg.Done()
process(job)
}()
}
wg.Wait()Go 1.25 added wg.Go(func() { process(job) }), which does the Add and Done for you and
removes the most common mistake.
sync.Once and friends#
once.Do(f) runs f exactly once, even when called from many goroutines, and every caller
waits until it has finished. sync.OnceValue(f) and sync.OnceValues(f) wrap a function so
its result is computed once and cached — the idiom for lazy initialization of a model, a
tokenizer or a client.
sync.Cond, sync.Map, sync.Pool#
Cond: “wait until a condition holds”, withWait,Signal,Broadcast. Rarely the best tool — a channel usually reads better — but right for a queue with complex wake-up rules.Map: a concurrent map specialized for keys written once and read often (II.03).Pool: reusable temporary objects (III.05).
golang.org/x/sync#
Outside the standard library but maintained by the Go team, and used everywhere:
| Package | Gives you |
|---|---|
errgroup | A WaitGroup that collects the first error and cancels a context; SetLimit(n) bounds concurrency |
semaphore | A weighted semaphore with context support |
singleflight | Collapse concurrent duplicate calls into one — the cure for a cache stampede |
What things cost#
Roughly, on a current CPU, uncontended:
| Operation | Time |
|---|---|
| Atomic add or load | ~1–5 ns |
| Mutex lock + unlock | ~2–20 ns |
| RWMutex RLock + RUnlock | ~5–20 ns |
| Buffered channel send + receive | ~15–60 ns |
| Park and wake a goroutine | ~200–500 ns |
Under contention these change by orders of magnitude, and the dominant cost becomes the cache line bouncing between cores (V.03). The fix for a contended lock is rarely a faster lock; it is less sharing: per-goroutine state merged at the end, sharding, or batching updates.
Code#
// counter.go — one shared counter, five ways: wrong, mutex, atomic, channel, and sharded.
package main
import (
"fmt"
"runtime"
"sync"
"sync/atomic"
"time"
)
const perGoroutine = 200000
func run(name string, workers int, inc func(worker int), total func() int64) {
var wg sync.WaitGroup
start := time.Now()
for w := 0; w < workers; w++ {
wg.Add(1)
go func() {
defer wg.Done()
for i := 0; i < perGoroutine; i++ {
inc(w)
}
}()
}
wg.Wait()
el := time.Since(start)
want := int64(workers * perGoroutine)
verdict := "correct"
if total() != want {
verdict = fmt.Sprintf("WRONG: lost %d updates", want-total())
}
fmt.Printf("%-22s %6.1f ns/op %s\n", name, float64(el.Nanoseconds())/float64(want), verdict)
}
func main() {
workers := runtime.GOMAXPROCS(0)
fmt.Println(workers, "goroutines incrementing a shared counter")
// 1. No synchronization: a data race. The result is simply wrong.
var racy int64
run("unsynchronized", workers, func(int) { racy++ }, func() int64 { return racy })
// 2. Mutex.
var mu sync.Mutex
var guarded int64
run("sync.Mutex", workers, func(int) { mu.Lock(); guarded++; mu.Unlock() }, func() int64 { return guarded })
// 3. Atomic.
var at atomic.Int64
run("atomic.Int64", workers, func(int) { at.Add(1) }, at.Load)
// 4. A channel and an owning goroutine.
ch := make(chan struct{}, 1024)
var owned int64
done := make(chan struct{})
go func() {
for range ch {
owned++
}
close(done)
}()
run("channel to an owner", workers, func(int) { ch <- struct{}{} }, func() int64 {
close(ch)
<-done
return owned
})
// 5. Sharded: each goroutine has its own padded counter; sum at the end. No sharing at all.
type padded struct {
n int64
_ [56]byte // keep each counter on its own cache line (V.03)
}
shards := make([]padded, workers)
run("per-goroutine shards", workers, func(w int) { shards[w].n++ }, func() int64 {
var s int64
for i := range shards {
s += shards[i].n
}
return s
})
}Remember this#
- Atomic for one word, mutex for an invariant over several fields, channel for handing over ownership.
- Never copy a mutex; keep critical sections short; lock in a consistent order.
RWMutexis not automatically faster; for read-mostly data, swap an immutable snapshot withatomic.Pointer.- The cheapest synchronization is not sharing: give each goroutine its own state and merge.
Try it#
- Run
counter.go. How many updates did the unsynchronized version lose? Run it again — is the number the same? - Remove the padding from the sharded version. How much slower is it, though the logic is unchanged? (Lesson V.03 explains.)
- Implement a config holder with
atomic.Pointer[Config]and a reload function. Show that readers never see a half-updated config.
Check yourself#
- Why can two atomics not protect an invariant between them?
- When is
RWMutexslower thanMutex? - What is the most effective way to fix a heavily contended lock?