PidokuInfra

Projects

Reading builds recognition. Building builds understanding. These six projects are all in Go with the standard library only, and none requires a GPU.

#ProjectAfter moduleTime
1A memory explorerII3–5 h
2An arena allocator and an allocation huntIII4–6 h
3A cancellable pipeline with a leak testIV4–6 h
4A fast matrix multiplyV5–8 h
5A tiny language model, end to endVI.058–12 h
6An embedding serviceVI.088–12 h

For each project, “done” means every box under Done when is ticked and you can explain the result to someone else.


1 — A memory explorer#

Goal. Make every drawing in module II something you can print.

Build. A package with Describe(v any) string that, using reflect and unsafe, prints the layout of a value: for a struct, each field’s offset, size and padding; for a slice, its pointer, length and capacity; for a string, its data pointer and length; for an interface, its two words and dynamic type; for a map, its length. Add Shares(a, b []T) bool that reports whether two slices overlap in memory.

Done when

  • It prints the padding of a badly ordered struct and of its reordered version.
  • It shows two slices of one array sharing memory, and stops showing it after an append that reallocates.
  • It shows a typed nil pointer inside a non-nil interface.
  • It suggests a field order that minimizes a struct’s size.

Stretch. Report the allocator size class each value would occupy on the heap.


2 — An arena allocator and an allocation hunt#

Goal. Practise removing allocations with measurement, not guesswork.

Build. Part one: a generic Arena[T] backed by chunked slices, handing out *T or indexes, with Reset() to free everything at once. Part two: take a deliberately wasteful program — a log parser that builds a map[string]*Record of a million entries with fmt.Sprintf keys — and reduce its allocations and GC time step by step, recording each step in a table.

Done when

  • A benchmark shows the arena allocating a million small objects with a handful of heap allocations.
  • For the parser you have a table: change made, allocs/op, B/op, ns/op, GC cycles.
  • Each change is justified by -gcflags=-m output or a heap profile, quoted in your notes.
  • The final version’s live heap contains no pointers in its largest structure.

Stretch. Compare GOGC=100, GOGC=400 and a GOMEMLIMIT on the original and final versions.


3 — A cancellable pipeline with a leak test#

Goal. Write concurrent code that is correct under cancellation, and prove it.

Build. A three-stage pipeline — read lines, transform with a pool of N workers, write results in input order — with a bounded channel between stages, a context for cancellation and first-error-cancels-all behaviour. Then a test suite that cancels at random points and asserts that every goroutine has exited.

Done when

  • Output order matches input order with N > 1 workers.
  • Cancelling mid-stream returns promptly and leaks no goroutines (checked with runtime.NumGoroutine or the goroutine-leak profile).
  • An error in one worker stops the others and is the error returned.
  • go test -race -count=50 passes.
  • A slow consumer causes backpressure, not memory growth: you can show the queue length stays bounded.

Stretch. Rewrite the tests with testing/synctest so that timeouts are tested without real waiting.


4 — A fast matrix multiply#

Goal. Apply modules III and V to the one kernel that matters most.

Build. MatMul(dst, a, b) for float32, improved in recorded steps: naive, flat storage, loop reordering, bounds-check elimination, unrolling, cache blocking, parallel workers. Report GFLOP/s for 256, 1,024 and 2,048 square at every step.

Done when

  • You have a table of GFLOP/s per step and size, from benchmarks compared with benchstat.
  • -gcflags=-d=ssa/check_bce/debug=1 reports no bounds checks in the inner loop.
  • The parallel version scales to at least half your cores and you can say what limits it.
  • Results match the naive version within a stated tolerance.
  • The benchmark loop performs zero allocations.

Stretch. With GOEXPERIMENT=simd, write a vector inner loop and add it to the table; or call a BLAS library through cgo and report the gap.


5 — A tiny language model, end to end#

Goal. Join the tokenizer, the transformer and sampling into one working program.

Build. Train a byte-pair tokenizer on a small text file. Implement the forward pass from VI.05 with a KV cache. Load weights from a simple binary file — either random, or exported from a tiny model trained elsewhere. Add greedy, temperature, top-k and top-p sampling, and a command-line program that streams generated text as it is produced.

Done when

  • The tokenizer round-trips arbitrary UTF-8 exactly.
  • Generation with and without the KV cache produces identical tokens, and you report the speed-up at 64 and 256 tokens.
  • The decode step performs zero allocations.
  • Four goroutines generate concurrently from one shared model with no data race under -race.
  • Streaming never prints half of a multi-byte character.

Stretch. Batch the prefill across positions as one matrix operation and measure time-to-first-token for a 500-token prompt before and after.


6 — An embedding service#

Goal. Build the kind of service Go is actually used for in AI systems.

Build. An HTTP service with POST /v1/embeddings (OpenAI-compatible request and response) backed by a small deterministic embedder, behind a dynamic batcher; and POST /v1/search over an in-memory vector index with exact and IVF modes. Include a bounded queue with 429 on overload, request deadlines, graceful shutdown, /metrics and /debug/pprof.

Done when

  • Under load, average batch size rises and p99 latency stays within a stated budget.
  • At twice capacity the service rejects quickly and its memory stays flat.
  • A disconnected client’s request is dropped before it reaches the model.
  • /metrics exposes request rate, errors, latency histogram, queue depth and batch size.
  • IVF recall against exact search is reported for at least three nprobe values.
  • SIGTERM drains in-flight requests and exits cleanly.

Stretch. Put the tool-calling loop from VI.09 in front of it: a tool that searches the index and a model client that answers from what it finds.


Finished all six? The projects in Inference Engineering continue from here: a full inference server, continuous batching, a KV-cache manager and a gateway.

↑↓ navigate↵ openesc close