PidokuInfra

Networking Fundamentals

Foundations Intermediate 1h Difficulty 2/5 Topic 08 of 13

Prerequisites 05


1. What is it?#

The network appears in inference in three distinct roles, with wildly different requirements:

1. Client ↔ service        HTTP/gRPC, streaming, ~10-200 ms RTT, small payloads
2. Service ↔ service       gateway → router → engine; ~0.1-1 ms, moderate payloads
3. GPU ↔ GPU across nodes  NCCL over InfiniBand/RoCE; ~2-10 µs, huge bandwidth

Role 3 gets its own treatment in Section IX. This file covers 1 and 2 and the general principles.


2. Why does it exist as a topic?#

Because a streaming, long-lived, high-token-rate workload stresses networking in ways ordinary request/response services do not — and because the most common “the model is slow” complaint turns out to be a proxy buffering the stream.


3. Simple analogy#

A phone call versus a letter. Ordinary HTTP is a letter: you write it, send it, get a reply. Streaming inference is a phone call: the connection stays open for minutes, information flows continuously, and anything in the middle that “helpfully” waits to hear the whole sentence before passing it on ruins the experience.

Most middleboxes were built for letters.


4. Tiny example#

The classic failure — a buffering proxy:

Engine:  emits token every 20 ms, starting at t=200 ms
Client:  should see first token at ~t=210 ms

With nginx default proxy_buffering on:
  nginx accumulates the response until its buffer fills or the response ends
  Client sees NOTHING until t=4,000 ms, then the whole thing at once.

TTFT measured at the server: 200 ms  ✓ dashboards green
TTFT experienced by the user: 4,000 ms  ✗ users furious

The fix:

nginx
location /v1/ {
    proxy_pass http://engine;
    proxy_buffering off;            # critical
    proxy_cache off;
    proxy_read_timeout 3600s;       # long generations
    proxy_set_header Connection '';
    proxy_http_version 1.1;
    chunked_transfer_encoding on;
}

Always measure TTFT from a real client through the real ingress path. Server-side metrics cannot see this class of bug, and it is extremely common.


5. Technical explanation#

Latency budget#

Same rack               0.05-0.2 ms
Same datacenter/AZ      0.2-1 ms
Cross-AZ                0.5-2 ms
Cross-region (same continent)  10-40 ms
Cross-continent         80-250 ms
TLS handshake           1-2 RTT (0-RTT with TLS 1.3 resumption)
TCP handshake           1 RTT
DNS (uncached)          10-100 ms

For a service with a 300 ms TTFT budget, a cross-continent client spends 200+ ms of it on physics. This is the entire argument for multi-region inference (Section XI.06) — you cannot optimize your way past the speed of light.

HTTP/1.1 vs HTTP/2 vs gRPC for streaming#

HTTP/1.1 + SSEHTTP/2gRPC
Streamingchunked encoding / SSEnative streamsnative
Multiplexing1 request per connectionmany per connectionmany
Head-of-line blockingper connectionper TCP connectionsame
Browser supportyes (EventSource/fetch)yesneeds grpc-web
Overhead per messagetext framingbinary HPACKprotobuf, compact
Typical usepublic APIinternalinternal, high-throughput

Practical guidance: SSE over HTTP/1.1 for public APIs (universal client support, works through everything, and it’s what the OpenAI-compatible ecosystem expects); gRPC internally between gateway and engines.

Server-Sent Events, precisely#

HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive
X-Accel-Buffering: no          ← tells nginx not to buffer

data: {"choices":[{"delta":{"content":"Hello"}}]}

data: {"choices":[{"delta":{"content":" world"}}]}

data: [DONE]

Note the blank line after each event — required by the spec. X-Accel-Buffering: no is a useful belt-and-braces header.

Connection management#

Streaming responses hold connections for the entire generation:

1,000 concurrent users × 30 s average generation
= 1,000 simultaneously open connections, held for 30 s each

Consequences:

  • File descriptor limits (ulimit -n) must be raised; 1024 is a common and fatal default.
  • Load balancers with connection limits per backend need tuning.
  • Idle timeouts must exceed your longest generation, or connections die mid-stream.
  • Ephemeral port exhaustion between gateway and engines at high connection churn — use connection pooling and keep-alive.

Bandwidth is rarely the issue (except when it is)#

Token stream: ~4 bytes/token + ~80 bytes SSE framing
5,000 tokens/sec × 84 bytes = 420 KB/s

Trivial. But:

  • Long prompts: a 100k-token prompt is ~400 KB of text per request. At 100 req/s that’s 40 MB/s inbound, plus JSON parsing cost.
  • Multimodal: images and audio are megabytes per request. This does saturate links.
  • KV cache transfer in disaggregated serving: gigabytes per request (Section XIII.06).

6. Under the hood#

Shell
# Connection states — look for TIME_WAIT buildup, SYN backlog drops
ss -s
ss -tan state established | wc -l

# Retransmissions and drops
netstat -s | grep -iE 'retrans|drop|overflow'
nstat -az | grep -i tcpext

# Per-connection latency
ss -tin | head -20        # shows rtt, cwnd, retrans per socket

# Is it the network or the server?
curl -w "@curl-format.txt" -o /dev/null -s https://your-endpoint/
# with curl-format.txt containing time_namelookup, time_connect,
# time_appconnect, time_starttransfer, time_total

time_starttransfer from curl is your true, client-measured TTFT including all network and proxy effects. Compare it to your server-side TTFT metric; the difference is your ingress tax.


7. Performance implications#

  • TTFT is the metric the network can ruin. ITL is usually safe (tokens are small and frequent), but a buffering middlebox turns a streaming API into a batch one.
  • TLS handshakes matter for short interactions. For agentic workloads making many short calls, connection reuse is worth 20-50 ms per call.
  • Nagle’s algorithm can add up to 40 ms delay to small writes. Set TCP_NODELAY on streaming sockets. Most frameworks do; verify.

8. Production implications#

  • Disable buffering at every layer: CDN, WAF, ingress controller, service mesh sidecar, application framework. Each is a separate place this bug hides.
  • Set generous read timeouts (longer than max generation time) but stall timeouts that are short (if no token for 30 s, something is wrong — kill it).
  • Raise fd limits and tune the listen backlog.
  • Terminate TLS at the edge, reuse connections internally.
  • Place inference close to users where TTFT matters; a cross-continent round trip can be half your budget.
  • Health checks must not count as load. A health check that runs a real generation on every probe wastes real GPU capacity.

9. Common mistakes#

Not testing through the real ingress. The #1 streaming bug.

Idle timeout shorter than generation time. Connections dropped at exactly 60 seconds is a classic signature.

Default ulimit -n of 1024. Fails at 1,000 concurrent streams.

Using HTTP/1.1 without keep-alive between gateway and engine. Handshake per request adds latency and burns ephemeral ports.

Compressing the SSE stream. gzip buffers by design; it defeats streaming. Disable compression on text/event-stream.

Ignoring client-side measurement. Server metrics cannot see ingress problems.


10. Hands-on exercise#

A. Reproduce the buffering bug. Put nginx in front of a streaming endpoint with default settings. Measure client TTFT with curl -N -w '%{time_starttransfer}'. Then set proxy_buffering off and measure again.

B. Map your latency budget. From your client location, measure DNS, TCP, TLS, and time-to-first-byte to your endpoint. What fraction of TTFT is network?

C. Connection limits. Write a client that opens N concurrent streaming requests. Increase N until something breaks. What broke — fds, backlog, memory, the LB? Fix it and repeat.

D. Compare protocols. Implement the same streaming endpoint with SSE and with gRPC server-streaming. Compare per-token overhead bytes and latency.


11. Interview questions#

  1. Why does a buffering proxy destroy an LLM API, and how would you detect it?
  2. What is SSE and why is it the default for LLM streaming APIs?
  3. Your users report the response arrives all at once after 5 seconds, but server TTFT is 200 ms. Diagnose.
  4. How do long-lived streaming connections change load balancer configuration?
  5. When would you use gRPC instead of HTTP+SSE for inference?
  6. What is a sensible timeout policy for an LLM API? Name at least two different timeouts.

12. Further reading#

  • [FUNDAMENTAL] Grigorik, High Performance Browser Networking — free online
  • [REFERENCE] MDN Server-Sent Events documentation
  • [REFERENCE] nginx proxy_buffering, X-Accel-Buffering
  • Next: 09 — PCIe, DMA, and interconnects

↑↓ navigate↵ openesc close