Everything so far, applied. This module builds the core pieces of an AI system in Go using nothing but the standard library: the arithmetic, a network that learns, a tokenizer, a transformer that generates, and the services around them.
None of it is meant to replace PyTorch or a production inference engine. It is meant to make those systems legible — and to show where Go genuinely is the right tool: the clients, servers, schedulers and pipelines around the model.
| # | Lesson | The question it answers |
|---|---|---|
| 01 | Go in the AI Stack | Where does Go fit, and what exists today? |
| 02 | Tensors and Matmul | How is a tensor stored, and how fast can pure Go multiply matrices? |
| 03 | A Neural Network from Scratch | How does a network compute, and how does it learn? |
| 04 | A Tokenizer | How does text become integers and back? |
| 05 | A Transformer Forward Pass | What does an LLM compute for each token, and what does the KV cache save? |
| 06 | Calling LLMs | How do I call a model API robustly: streaming, retries, structured output? |
| 07 | Embeddings and Vector Search | How do I find the nearest vectors among a million? |
| 08 | Serving Models from Go | How do I batch requests and apply backpressure in front of a model? |
| 09 | A Tool-Calling Loop | What is the loop at the heart of an agent? |
When you finish you can read an inference engine’s source and recognize every part, and you can build the Go services that surround a model in production.
Where a lesson stops, Inference Engineering continues: its projects build a full engine, continuous batching, a KV-cache manager and a gateway — also in Go.