PidokuInfra
06AI in Go

AI in Go

ExpertModule 069 topics~9h 15m

Topics, in order

01 Go in the AI StackAn AI system has two very different kinds of code. The numerical core — training loops and tensor kernels — runs on GPUs through C++, CUDA and Python, and Go is not the language for it. … Advanced 40 min 02 Tensors and MatmulA tensor is a block of numbers with a shape. In memory it is always the same thing: one flat []float32 plus the shape that says how to interpret it. Element [i][j][k] of a tensor of shape … Advanced 1h 5m 03 A Neural Network from ScratchA neural network is a function built from layers. Each layer multiplies its input by a matrix of weights, adds a bias, and applies a simple non-linear function. Running the layers in order … Advanced 1h 10m 04 A TokenizerA model consumes integers, not text. A tokenizer maps text to a sequence of token IDs and back. Modern LLM tokenizers use byte-pair encoding (BPE): start with the 256 possible bytes as the … Advanced 55 min 05 A Transformer Forward PassA large language model is a function from a sequence of token IDs to a score for every possible next token. Inside, each token is a vector that passes through a stack of identical layers. … Expert 1h 30m 06 Calling LLMsCalling a language model is an HTTP request with a JSON body and, usually, a streamed response. The mechanics are simple. Doing it well is four things: stream so users see output immediately … Advanced 1h 07 Embeddings and Vector SearchAn embedding is a vector — a few hundred to a few thousand float32s — produced by a model so that texts with similar meaning land close together. Search by meaning then becomes geometry: … Advanced 1h 08 Serving Models from GoA model is expensive to call and cheap to call wider: running sixteen inputs together costs little more than running one. A serving layer exploits that with dynamic batching — hold each … Expert 1h 5m 09 A Tool-Calling LoopA model can only produce text. It becomes able to do things when your program offers it tools — functions described by a name, a purpose and a JSON schema for their arguments — and runs a … Expert 50 min

About this module

Everything so far, applied. This module builds the core pieces of an AI system in Go using nothing but the standard library: the arithmetic, a network that learns, a tokenizer, a transformer that generates, and the services around them.

None of it is meant to replace PyTorch or a production inference engine. It is meant to make those systems legible — and to show where Go genuinely is the right tool: the clients, servers, schedulers and pipelines around the model.

#LessonThe question it answers
01Go in the AI StackWhere does Go fit, and what exists today?
02Tensors and MatmulHow is a tensor stored, and how fast can pure Go multiply matrices?
03A Neural Network from ScratchHow does a network compute, and how does it learn?
04A TokenizerHow does text become integers and back?
05A Transformer Forward PassWhat does an LLM compute for each token, and what does the KV cache save?
06Calling LLMsHow do I call a model API robustly: streaming, retries, structured output?
07Embeddings and Vector SearchHow do I find the nearest vectors among a million?
08Serving Models from GoHow do I batch requests and apply backpressure in front of a model?
09A Tool-Calling LoopWhat is the loop at the heart of an agent?

When you finish you can read an inference engine’s source and recognize every part, and you can build the Go services that surround a model in production.

Where a lesson stops, Inference Engineering continues: its projects build a full engine, continuous batching, a KV-cache manager and a gateway — also in Go.

↑↓ navigate↵ openesc close