Reading builds recognition. Building builds understanding. These six projects are all in Go with the standard library only, and none requires a GPU.
| # | Project | After module | Time |
|---|---|---|---|
| 1 | A memory explorer | II | 3–5 h |
| 2 | An arena allocator and an allocation hunt | III | 4–6 h |
| 3 | A cancellable pipeline with a leak test | IV | 4–6 h |
| 4 | A fast matrix multiply | V | 5–8 h |
| 5 | A tiny language model, end to end | VI.05 | 8–12 h |
| 6 | An embedding service | VI.08 | 8–12 h |
For each project, “done” means every box under Done when is ticked and you can explain the result to someone else.
1 — A memory explorer#
Goal. Make every drawing in module II something you can print.
Build. A package with Describe(v any) string that, using reflect and unsafe, prints
the layout of a value: for a struct, each field’s offset, size and padding; for a slice, its
pointer, length and capacity; for a string, its data pointer and length; for an interface, its
two words and dynamic type; for a map, its length. Add Shares(a, b []T) bool that reports
whether two slices overlap in memory.
Done when
- It prints the padding of a badly ordered struct and of its reordered version.
- It shows two slices of one array sharing memory, and stops showing it after an
appendthat reallocates. - It shows a typed nil pointer inside a non-nil interface.
- It suggests a field order that minimizes a struct’s size.
Stretch. Report the allocator size class each value would occupy on the heap.
2 — An arena allocator and an allocation hunt#
Goal. Practise removing allocations with measurement, not guesswork.
Build. Part one: a generic Arena[T] backed by chunked slices, handing out *T or indexes,
with Reset() to free everything at once. Part two: take a deliberately wasteful program — a
log parser that builds a map[string]*Record of a million entries with fmt.Sprintf keys —
and reduce its allocations and GC time step by step, recording each step in a table.
Done when
- A benchmark shows the arena allocating a million small objects with a handful of heap allocations.
- For the parser you have a table: change made,
allocs/op,B/op,ns/op, GC cycles. - Each change is justified by
-gcflags=-moutput or a heap profile, quoted in your notes. - The final version’s live heap contains no pointers in its largest structure.
Stretch. Compare GOGC=100, GOGC=400 and a GOMEMLIMIT on the original and final
versions.
3 — A cancellable pipeline with a leak test#
Goal. Write concurrent code that is correct under cancellation, and prove it.
Build. A three-stage pipeline — read lines, transform with a pool of N workers, write results in input order — with a bounded channel between stages, a context for cancellation and first-error-cancels-all behaviour. Then a test suite that cancels at random points and asserts that every goroutine has exited.
Done when
- Output order matches input order with N > 1 workers.
- Cancelling mid-stream returns promptly and leaks no goroutines (checked with
runtime.NumGoroutineor the goroutine-leak profile). - An error in one worker stops the others and is the error returned.
-
go test -race -count=50passes. - A slow consumer causes backpressure, not memory growth: you can show the queue length stays bounded.
Stretch. Rewrite the tests with testing/synctest so that timeouts are tested without real
waiting.
4 — A fast matrix multiply#
Goal. Apply modules III and V to the one kernel that matters most.
Build. MatMul(dst, a, b) for float32, improved in recorded steps: naive, flat storage,
loop reordering, bounds-check elimination, unrolling, cache blocking, parallel workers. Report
GFLOP/s for 256, 1,024 and 2,048 square at every step.
Done when
- You have a table of GFLOP/s per step and size, from benchmarks compared with
benchstat. -
-gcflags=-d=ssa/check_bce/debug=1reports no bounds checks in the inner loop. - The parallel version scales to at least half your cores and you can say what limits it.
- Results match the naive version within a stated tolerance.
- The benchmark loop performs zero allocations.
Stretch. With GOEXPERIMENT=simd, write a vector inner loop and add it to the table; or
call a BLAS library through cgo and report the gap.
5 — A tiny language model, end to end#
Goal. Join the tokenizer, the transformer and sampling into one working program.
Build. Train a byte-pair tokenizer on a small text file. Implement the forward pass from VI.05 with a KV cache. Load weights from a simple binary file — either random, or exported from a tiny model trained elsewhere. Add greedy, temperature, top-k and top-p sampling, and a command-line program that streams generated text as it is produced.
Done when
- The tokenizer round-trips arbitrary UTF-8 exactly.
- Generation with and without the KV cache produces identical tokens, and you report the speed-up at 64 and 256 tokens.
- The decode step performs zero allocations.
- Four goroutines generate concurrently from one shared model with no data race under
-race. - Streaming never prints half of a multi-byte character.
Stretch. Batch the prefill across positions as one matrix operation and measure time-to-first-token for a 500-token prompt before and after.
6 — An embedding service#
Goal. Build the kind of service Go is actually used for in AI systems.
Build. An HTTP service with POST /v1/embeddings (OpenAI-compatible request and response)
backed by a small deterministic embedder, behind a dynamic batcher; and POST /v1/search over
an in-memory vector index with exact and IVF modes. Include a bounded queue with 429 on
overload, request deadlines, graceful shutdown, /metrics and /debug/pprof.
Done when
- Under load, average batch size rises and p99 latency stays within a stated budget.
- At twice capacity the service rejects quickly and its memory stays flat.
- A disconnected client’s request is dropped before it reaches the model.
-
/metricsexposes request rate, errors, latency histogram, queue depth and batch size. - IVF recall against exact search is reported for at least three
nprobevalues. - SIGTERM drains in-flight requests and exits cleanly.
Stretch. Put the tool-calling loop from VI.09 in front of it: a tool that searches the index and a model client that answers from what it finds.
Finished all six? The projects in Inference Engineering continue from here: a full inference server, continuous batching, a KV-cache manager and a gateway.