The idea in one minute#
A trace of an AI request has the usual HTTP and database spans plus three new kinds: a model call (which model, how many tokens in and out, why it stopped), a tool call (the model asked for a function to be run), and an agent invocation (a loop of model and tool calls working toward a goal).
OpenTelemetry’s GenAI semantic conventions give these spans standard names and attributes
(gen_ai.*), so one instrumentation works across providers and backends. As of October 2026
every one of those attributes is still marked Development: adopt them, pin the version, and
expect renames.
An analogy#
A detective’s case file. Each entry records who was asked (the model), what was asked and answered (the messages), what was looked up as a result (tool calls), and how one lead led to the next (the parent-child structure). Without the file you have a verdict and no idea how it was reached.
A picture#
flowchart TB ROOT["POST /v1/tasks<br/>SERVER span"] --> AG["invoke_agent research-agent"] AG --> C1["chat model-large<br/>input 1,820 tok, output 96, finish: tool_calls"] AG --> T1["execute_tool search_docs"] T1 --> R1["retrieval / HTTP / DB spans"] AG --> C2["chat model-large<br/>input 6,410 tok, 5,900 cached, output 412, finish: stop"] C1 -.->|"propagated traceparent"| SRV["Inference server spans<br/>queue, prefill, decode"] C2 -.-> SRV class ROOT neutral class AG queue class C1,C2 compute class T1,R1 io class SRV memory
How it really works#
The span types#
Operation (gen_ai.operation.name) | Span name | Represents |
|---|---|---|
chat | chat {model} | One call to a chat or completion model |
embeddings | embeddings {model} | One embedding call |
execute_tool | execute_tool {tool} | Running a tool the model requested |
invoke_agent | invoke_agent {agent} | One agent run: a loop of the above |
create_agent | create_agent {agent} | Agent set-up |
| retrieval | (added in 2026) | Fetching context for the model |
The attributes that matter#
| Attribute | Example | Why you want it |
|---|---|---|
gen_ai.provider.name | openai, anthropic, aws.bedrock, or your own | Slice by provider. Replaced the older gen_ai.system |
gen_ai.request.model / gen_ai.response.model | requested vs the exact one that answered | Aliases and routing change the answer |
gen_ai.usage.input_tokens / gen_ai.usage.output_tokens | 6410 / 412 | Cost and load. Cached-token counts are recorded separately where the provider reports them |
gen_ai.response.finish_reasons | ["stop"], ["length"], ["tool_calls"] | length means truncated output |
gen_ai.request.temperature, .max_tokens, .top_p | Reproducing behaviour | |
gen_ai.response.id | provider’s response ID | Joining evaluations and feedback back to the call |
gen_ai.conversation.id | Grouping the turns of a session | |
gen_ai.agent.name, gen_ai.tool.name, gen_ai.tool.call.id | Agent and tool identity | |
error.type | timeout, rate_limit_exceeded | Failure classification |
Tool calls made over the Model Context Protocol carry MCP attributes (mcp.method.name,
mcp.session.id) on the same execute_tool span instead of a duplicate span.
The metrics#
The conventions also define metrics, so client-side and server-side dashboards line up:
| Metric | Type | Side |
|---|---|---|
gen_ai.client.token.usage | Histogram, by token type | Caller |
gen_ai.client.operation.duration | Histogram | Caller |
gen_ai.server.request.duration | Histogram | Model server |
gen_ai.server.time_to_first_token | Histogram | Model server |
gen_ai.server.time_per_output_token | Histogram | Model server |
Content: prompts and completions#
The text of prompts and responses is the most useful thing for debugging quality and the most
dangerous thing to store. The conventions therefore make content opt-in:
gen_ai.input.messages, gen_ai.output.messages and gen_ai.system_instructions are recorded
only when you enable content capture.
A sensible policy:
- Off by default. Token counts, models and finish reasons are always on.
- When on: redact in the Collector, sample, restrict who can read, and set a short retention.
- For large payloads, store the content elsewhere (an object store) and put a reference on the span.
- Decide per tenant; some customers’ contracts forbid retention outright.
Status, honestly (checked 3 October 2026)#
- Every
gen_ai.*span, attribute, metric and event is at Development stability. None is stable. - In June 2026 the conventions moved out of the main semantic-conventions repository into a
dedicated one,
open-telemetry/semantic-conventions-genai. - There have been breaking renames along the way (
gen_ai.system→gen_ai.provider.name;prompt_tokens/completion_tokens→input_tokens/output_tokens; message content moved from events to attributes). - Instrumentation libraries let you opt into the newest form with
OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental; without it, many still emit an older shape.
What to do about it: use the conventions anyway — every major backend reads them, and a standard that moves is still better than a private schema. Pin instrumentation versions, keep attribute names in one place in your code, and normalize old names to new in the Collector.
Who produces these spans#
| Source | Notes |
|---|---|
| OpenTelemetry instrumentation packages for model SDKs | The reference implementations |
| OpenLLMetry (Traceloop), OpenInference (Arize), and similar | Broader framework coverage; converge on OTel conventions |
| Agent frameworks | Many emit OTel spans natively |
| Gateways and proxies (LiteLLM, Envoy AI Gateway, …) | One instrumentation point for every provider |
| Inference engines | vLLM and others can emit OTLP spans for a request’s server-side phases; llm-d traces across router and engine |
A gateway is the highest-leverage place: every model call passes through it, so one piece of middleware yields uniform spans, token metrics and per-tenant accounting.
Connecting caller and server#
When you run the model yourself, propagate traceparent from the application through the
gateway into the engine. One trace then shows the agent’s chat span and underneath it the
engine’s queue, prefill and decode — which tells you whether a slow step was the model, the
queue or the network. With a hosted API the trace stops at your client span; record the
provider’s response ID so their support can find it.
Tracing agents#
- One
invoke_agentspan per task, with every model and tool call as children. - Put the task outcome on the root (succeeded, gave up, hit the step limit) and the number of steps.
- Watch for the patterns that only a trace shows: loops that repeat the same tool call, contexts growing every turn, a retry storm on a failing tool, one slow tool dominating.
- Cost per task is the sum of token usage over the subtree — compute it in the backend.
Code#
A hand-built agent trace with GenAI attributes, summarized the way a backend would: cost and latency rolled up over the tree.
// genai.go — an agent run as spans with gen_ai.* attributes, and the roll-ups that matter.
package main
import (
"fmt"
"strings"
)
type Span struct {
ID, Parent, Name string
Millis float64
Attr map[string]any
}
func main() {
spans := []Span{
{"a", "", "invoke_agent research-agent", 9840, map[string]any{
"gen_ai.operation.name": "invoke_agent", "gen_ai.agent.name": "research-agent"}},
{"b", "a", "chat model-large", 1310, map[string]any{
"gen_ai.operation.name": "chat", "gen_ai.provider.name": "self-hosted",
"gen_ai.request.model": "model-large", "gen_ai.usage.input_tokens": 1820,
"cached_input_tokens": 0, "gen_ai.usage.output_tokens": 96,
"gen_ai.response.finish_reasons": "tool_calls"}},
{"c", "a", "execute_tool search_docs", 640, map[string]any{
"gen_ai.operation.name": "execute_tool", "gen_ai.tool.name": "search_docs"}},
{"d", "a", "chat model-large", 2950, map[string]any{
"gen_ai.operation.name": "chat", "gen_ai.request.model": "model-large",
"gen_ai.usage.input_tokens": 6410, "cached_input_tokens": 1900,
"gen_ai.usage.output_tokens": 180, "gen_ai.response.finish_reasons": "tool_calls"}},
{"e", "a", "execute_tool run_query", 1180, map[string]any{
"gen_ai.operation.name": "execute_tool", "gen_ai.tool.name": "run_query"}},
{"f", "a", "chat model-large", 3620, map[string]any{
"gen_ai.operation.name": "chat", "gen_ai.request.model": "model-large",
"gen_ai.usage.input_tokens": 9870, "cached_input_tokens": 6500,
"gen_ai.usage.output_tokens": 412, "gen_ai.response.finish_reasons": "stop"}},
}
// Illustrative prices per million tokens.
const inPrice, cachedPrice, outPrice = 2.00, 0.20, 8.00
fmt.Println("span ms in cached out finish")
var in, cached, out int
var modelMs, toolMs float64
steps := 0
for _, s := range spans {
indent := ""
if s.Parent != "" {
indent = " "
}
i, _ := s.Attr["gen_ai.usage.input_tokens"].(int)
c, _ := s.Attr["cached_input_tokens"].(int)
o, _ := s.Attr["gen_ai.usage.output_tokens"].(int)
fin, _ := s.Attr["gen_ai.response.finish_reasons"].(string)
fmt.Printf("%-30s %6.0f %6d %8d %6d %s\n", indent+s.Name, s.Millis, i, c, o, fin)
in, cached, out = in+i, cached+c, out+o
switch s.Attr["gen_ai.operation.name"] {
case "chat":
modelMs += s.Millis
steps++
case "execute_tool":
toolMs += s.Millis
}
}
cost := (float64(in-cached)*inPrice + float64(cached)*cachedPrice + float64(out)*outPrice) / 1e6
fmt.Println(strings.Repeat("-", 72))
fmt.Printf("task: %d model calls, %.1f s in models, %.1f s in tools, %.1f s unaccounted\n",
steps, modelMs/1000, toolMs/1000, (spans[0].Millis-modelMs-toolMs)/1000)
fmt.Printf("tokens: %d input (%.0f%% cached), %d output → $%.4f per task\n",
in, 100*float64(cached)/float64(in), out, cost)
fmt.Println("input tokens grow every step: context accumulation is the cost driver in agents.")
}Remember this#
- New span kinds: model call, tool call, agent invocation — standardized as
gen_ai.*. - Always record model, token counts and finish reason. Capture content only by opt-in, with redaction and short retention.
- The conventions are widely supported and not yet stable: pin versions and normalize names.
- Propagate trace context into your own engine to see queue, prefill and decode under each call.
- For agents, the task is the unit: roll up cost and latency over the whole subtree.
Try it#
- Run
genai.go. Make the agent loop one extra step with the same tool. How much does the task cost change, and where would you see the loop in a trace viewer? - Design a content-capture policy for a product with consumer and enterprise tenants.
- Instrument a small program that calls any model API with an OTel GenAI instrumentation
library and send the spans to a Collector’s
debugexporter. Which attribute names does your version emit?
Check yourself#
- Which three span kinds are specific to AI applications?
- Why is content capture opt-in?
- What does “Development stability” oblige you to do in practice?