PidokuInfra

TensorRT and TensorRT-LLM

Intermediate Advanced 1h 15m Difficulty 4/5 Topic 09 of 14

Prerequisites IV.10, 05, 07


1. What is it?#

TensorRT is NVIDIA’s ahead-of-time inference compiler: it takes a model graph and builds an optimized, hardware-specific binary “engine.”

TensorRT-LLM is the LLM-specific layer on top: transformer-aware optimizations, paged KV cache, in-flight (continuous) batching, tensor/pipeline parallelism, and quantization support.

Model (HF checkpoint)
   → convert to TRT-LLM's representation
   → BUILD (minutes to hours): fuse, autotune, select kernels, plan memory
   → engine file (.engine, GPU-architecture-specific)
   → serve via the C++ runtime or Triton Inference Server

2. Problem → Why → Optimization#

PROBLEM   A JIT/eager stack leaves performance on the table: kernel selection
          is heuristic, fusion is incomplete, memory planning is conservative,
          and shapes are only known at runtime.
WHY       Because it must stay flexible.
OPTIMIZE  Give up flexibility. Fix the model, precision, parallelism, and
          shape ranges at build time, then optimize exhaustively — including
          benchmarking kernel candidates on the actual hardware.

TRADE-OFFS
  ✓ typically the fastest option for a fixed configuration (10-30% over vLLM
    on throughput benchmarks; more on some shapes)
  ✓ excellent FP8 support on Hopper
  ✗ build times of minutes to hours
  ✗ engines are architecture-specific — rebuild for each GPU type
  ✗ shape ranges baked in (max batch, max input len, max output len)
  ✗ steeper operational learning curve
  ✗ slower to support new model architectures than vLLM/SGLang

WHEN TO USE   Stable model, stable configuration, high volume, NVIDIA-only fleet,
              throughput or latency is worth operational complexity.
WHEN NOT TO   Rapidly changing models, many models, experimentation, heterogeneous
              hardware, or when the team's time is the scarce resource.

3. Simple analogy#

A bespoke suit versus off-the-peg.

Off-the-peg (vLLM): available today, fits most people well, easy to exchange.

Bespoke (TensorRT-LLM): measured to your exact body, fits better — but takes weeks, costs more, and if you gain weight (change the model or GPU) you start again.

Both are correct answers; the question is how stable your requirements are.


4. Tiny example — the workflow#

Shell
# 1. Convert a HuggingFace checkpoint to TRT-LLM format
python convert_checkpoint.py \
    --model_dir ./Llama-3-8B-Instruct \
    --output_dir ./tllm_ckpt \
    --dtype bfloat16 \
    --tp_size 1

# 2. Build the engine — this is where the optimization happens
trtllm-build \
    --checkpoint_dir ./tllm_ckpt \
    --output_dir ./engine \
    --gemm_plugin bfloat16 \
    --max_batch_size 64 \
    --max_input_len 4096 \
    --max_seq_len 8192 \
    --max_num_tokens 8192 \
    --use_paged_context_fmha enable \
    --kv_cache_type paged

# 3. Serve
python run.py --engine_dir ./engine --tokenizer_dir ./Llama-3-8B-Instruct
# or deploy via Triton Inference Server with the tensorrtllm_backend

The build parameters are commitments. max_input_len 4096 means a 5,000-token prompt fails. max_batch_size 64 caps concurrency. Choose them from your measured traffic distribution with headroom, and document them — they’re now part of your service’s contract.


5. What the build actually does#

1. GRAPH OPTIMIZATION
   constant folding, dead code elimination, layer fusion
   (attention + rope + kv-write into one plugin; gemm + bias + activation)

2. KERNEL SELECTION BY MEASUREMENT
   For each GEMM shape, benchmark candidate kernels ON THE ACTUAL GPU
   and pick the fastest. This is why the build is slow — and why it beats
   heuristic selection.

3. PRECISION ASSIGNMENT
   Apply the requested quantization; insert scale/dequant nodes;
   choose per-layer precision if using mixed precision.

4. MEMORY PLANNING
   Compute exact live ranges; allocate one workspace with buffer reuse.

5. PLUGIN SELECTION
   TRT-LLM ships hand-written plugins for attention (paged, fmha),
   GEMM, RoPE, all-reduce, quantized GEMMs.

6. SERIALIZATION
   Write the engine: kernels, weights, memory plan, all baked in.

Step 2 is the differentiator. cuBLAS uses heuristics to pick a kernel; TensorRT measures.


6. TensorRT-LLM’s LLM-specific features#

✓ Paged KV cache                       equivalent to vLLM's
✓ In-flight batching                   = continuous batching
✓ Chunked context (chunked prefill)
✓ Prefix caching (reuse of paged context blocks)
✓ Tensor parallelism, pipeline parallelism, expert parallelism
✓ FP8, INT8 (SmoothQuant), INT4 (AWQ/GPTQ), FP8 KV cache
✓ Multi-LoRA
✓ Speculative decoding (draft model, Medusa, EAGLE, lookahead)
✓ Multi-block attention for long context
✓ Custom all-reduce (faster than NCCL for small messages)

Feature parity with vLLM is close; the differences are in maturity per feature and in how quickly new model architectures are supported (vLLM/SGLang are usually faster there).


7. Performance#

Published and reproduced benchmarks vary by workload; representative ranges:

Llama-3-70B, 8×H100, throughput benchmark:
  vLLM (BF16)             baseline
  TRT-LLM (BF16)          1.05-1.25x
  vLLM (FP8)              1.7-1.9x
  TRT-LLM (FP8)           2.0-2.4x

Low-latency (batch 1-4):
  TRT-LLM often 1.2-1.4x ahead (better kernel selection, tighter fusion,
  custom all-reduce)

Caveats worth stating plainly:

  • Benchmarks are usually run by the vendor of the winning system.
  • The gap narrows every release as both improve.
  • The gap is often smaller than the gap between a well-tuned and a badly-tuned deployment of either.
  • Operational complexity is a real cost that benchmarks don’t show.

Practical guidance: if you’re serving one or two stable models at high volume on NVIDIA hardware, TRT-LLM is worth evaluating. Otherwise the extra 15% rarely justifies the rebuild treadmill.


8. Production implications#

  • Engines are (GPU architecture × TRT version × model × config)-specific. Your build matrix grows quickly. Automate it in CI.
  • Store built engines as artifacts, versioned and keyed by that tuple. Never build at deployment time.
  • The shape limits are a service contract. max_input_len exceeded = hard failure. Enforce the same limit at your gateway with a clear error.
  • Plan the upgrade path. A TRT-LLM version bump usually requires rebuilding every engine and re-validating.
  • Triton Inference Server is the usual serving front end. It adds its own operational surface (model repository, config.pbtxt, ensemble models).
  • Validate quality on the built engine, not on the source checkpoint. The build changes numerics.

9. Common mistakes#

Building at deployment time. Minutes to hours of downtime.

max_input_len too small. Production requests fail hard.

Deploying an engine built for a different GPU architecture. Fails to load.

Not automating the build matrix. It becomes unmanageable at 3 models × 2 GPU types × 2 precisions.

Assuming the benchmark gap transfers to your workload. Benchmark your own shapes.

Choosing TRT-LLM for a rapidly-changing model portfolio. The rebuild cost dominates.

Not re-validating quality after a rebuild.


10. Hands-on exercise#

A. Build and serve. Convert and build a small model (1-8B). Time each phase. Serve it and verify correctness against HuggingFace.

B. Benchmark honestly. Compare TRT-LLM and vLLM on the same workload with the same precision: throughput, p50/p95 TTFT, p50/p95 ITL. Use a realistic length distribution, not fixed lengths. Report all numbers, including the ones that don’t favor your preferred system.

C. Shape limits. Build with max_input_len 2048. Send a 3,000-token prompt. Observe the failure mode. Design the gateway-side guard.

D. Build matrix. Write a script that builds engines for {2 precisions} × {2 TP degrees} and stores them as versioned artifacts. Time the whole matrix.

E. Quality check. Compare outputs from the TRT-LLM engine and the HF model on 100 prompts. Measure top-1 agreement and any perplexity difference.


11. Interview questions#

  1. What does TensorRT’s build step do that a JIT compiler cannot?
  2. Why must engines be rebuilt per GPU architecture?
  3. What are the build-time shape commitments and how do you choose them?
  4. When would you choose TensorRT-LLM over vLLM, and when not?
  5. How would you manage engines in a CI/CD pipeline?
  6. Why is the reported benchmark gap often smaller in practice than in vendor benchmarks?
  7. What operational costs does TRT-LLM add?

12. Further reading#

  • [REFERENCE] TensorRT-LLM documentation and examples/ directory
  • [REFERENCE] NVIDIA Triton Inference Server documentation
  • [REFERENCE] TensorRT-LLM performance documentation (docs/source/performance/)
  • Next: 10 — FlashAttention

↑↓ navigate↵ openesc close