PidokuInfra

Projects

Reading builds recognition. Building builds understanding. These five projects are all in Go, and only the last needs a GPU.

#ProjectAfter moduleNeeds a GPUTime
1A SIMT simulatorIINo3–5 h
2A roofline explorerIVNo2–4 h
3A fast matmul in pure GoVNo4–8 h
4A GPU monitor and capacity plannerVOptional4–6 h
5A CUDA kernel behind a Go serviceVIYes6–10 h

For each project, “done” means every box under Done when is ticked and you can explain the result to someone else.


1 — A SIMT simulator#

Goal. Make warps, masking and occupancy concrete by building them.

Build. A Go package that executes a tiny instruction set (load, store, add, mul, branch-if, halt) for a grid of threads. Threads are grouped in warps of 32 sharing one program counter; divergent branches are handled with an active mask and a reconvergence stack. Memory loads take a configurable number of cycles, during which the warp is not ready and the SM runs another.

Done when

  • SAXPY and a branchy kernel both produce correct results.
  • You can print total cycles, and cycles lost to masked (idle) threads.
  • Sorting the branchy kernel’s input measurably reduces cycles.
  • Raising the number of resident warps reduces idle SM cycles until it stops helping, and you can explain where it stops.

Stretch. Count memory transactions per warp load and show the effect of coalescing.


2 — A roofline explorer#

Goal. Predict performance from a data sheet.

Build. A CLI: roofline -gpu h100 -params 8e9 -bits 16 -batch 1,8,64,256. It prints intensity, regime, step time, total and per-user tokens/s, and KV-cache-limited concurrency for a given context length. GPUs are defined in a small JSON file.

Done when

  • Output matches the worked examples in IV.01 and V.03.
  • -plot writes an SVG of the roofline with your workload’s points on it (plain fmt.Fprintf of SVG elements is enough).
  • You have used it to answer one real question about hardware you use or want.

Stretch. Add launch overhead (IV.04) and tensor-parallel communication (VI.02) to the model.


3 — A fast matmul in pure Go#

Goal. Feel the gap between “correct” and “fast” and close some of it.

Build. MatMul(a, b, c []float32, n int) in four versions: naive, loop-reordered, tiled, and tiled + goroutines. A benchmark harness (IV.05’s rules) reports GFLOP/s for each at several sizes.

Done when

  • All versions agree to within floating-point tolerance.
  • You have a table of GFLOP/s by version and size, and can explain each jump.
  • You have measured your machine’s memory bandwidth (II.03) and placed each version on a roofline.

Stretch. Call an optimized BLAS through cgo and compare. Explain the remaining gap (SIMD).


4 — A GPU monitor and capacity planner#

Goal. Build the tool an on-call engineer actually wants.

Build. A Go service that samples GPUs (NVML via go-nvml, or nvidia-smi CSV; a fake sampler for machines without GPUs), exposes Prometheus metrics on /metrics, and serves a /plan endpoint that takes a model configuration and returns whether it fits and the expected concurrency (V.03).

Done when

  • Metrics include utilization, memory used/total, power as a fraction of limit, and temperature.
  • An alert condition flags “utilization high, power low” (IV.05).
  • /plan answers correctly for three models you looked up.
  • It runs without a GPU using the fake sampler, and its tests pass.

Stretch. Package it as a Kubernetes DaemonSet.


5 — A CUDA kernel behind a Go service#

Goal. Join every layer: Go → cgo → CUDA → GPU, served over HTTP.

Build. A CUDA kernel of your choice (batched matrix-vector multiply is ideal), wrapped in C with separate init, upload, run, download and close functions. A Go HTTP service keeps the weights on the device, batches concurrent requests into one launch, and returns results.

Done when

  • Weights are uploaded once; requests do not re-upload them.
  • Concurrent requests are combined into a single kernel launch, with a maximum wait.
  • You have measured throughput and p50/p99 latency at batch sizes 1, 8 and 64, with correct synchronization.
  • You can show the batching curve and mark the regime change against your roofline.

Stretch. Use pinned host buffers and a second stream to overlap upload with compute.

No GPU? Do the same project with the kernel replaced by a goroutine-parallel Go function. The batching service — the part most Go engineers will write professionally — is identical.

↑↓ navigate↵ openesc close