Stop Inventing Your Own Telemetry: OpenTelemetry's GenAI Conventions, And Wiring Evals Into The Same Pipeline
For two years every team instrumented their LLM calls differently, which meant every observability vendor built a different integration and nobody could compare anything. That ended in 2026: OpenTelemetry graduated from the CNCF in May, its GenAI semantic conventions stabilised the attribute schema for LLM spans, agent traces and AI metrics, and the major platforms - Google Cloud, AWS, Azure, Datadog - adopted them. For financial teams this is more than tidiness. A standard schema is what lets cost, latency, token use, tool calls and quality evaluation land in one pipeline, which is the only realistic way to answer the question a supervisor eventually asks: what did this system do, and how do you know it was right?
AlchmAI Engineering14 min read
May 2026
OpenTelemetry graduated from the CNCF, with its GenAI semantic conventions stabilising the attribute schema for LLM spans, agent traces and AI metrics
6 layers
Covered by the conventions: LLM client spans, agent spans, events, metrics, agent orchestration and MCP tool calling, plus quality evaluation
Experimental
The status of most GenAI conventions as of early 2026 - dual emission via OTEL_SEMCONV_STABILITY_OPT_IN is how you adopt safely
Every PR
Touching a prompt, model version or retrieval config should trigger an eval run against the golden dataset, and block on regression
Ask a financial engineering team what their AI system cost last month and you will usually get a number from a provider dashboard. Ask what it cost per completed reconciliation, or which workflow drove the increase, or whether the p95 latency regression last Tuesday came from the model, the retrieval layer or a slow MCP server, and the answers get vague. Not because the teams are careless, but because until recently everyone instrumented this differently. Each observability vendor built a bespoke integration, every framework emitted its own attribute names, and comparing two systems - or two weeks of the same system - meant comparing two schemas.
That is over, and the change is more consequential than it sounds. OpenTelemetry graduated from the CNCF in May 2026, and its GenAI semantic conventions stabilised the attribute schema for LLM spans, agent traces and AI metrics. The conventions began by covering LLM client spans, agent spans, events for capturing prompt and completion content, and metrics; they have since expanded to cover agent orchestration, MCP tool calling, content capture and quality evaluation - six layers in all. Google Cloud, AWS, Azure and Datadog have adopted them, with Datadog's agent observability supporting them natively. For once the boring answer is the right one: adopt the standard, do not invent a schema.
What To Instrument, And What The Attributes Are Called
The value of a convention is that you stop making decisions. Here is a financial agent step instrumented to the GenAI conventions, with the finance-specific attributes added in our own namespace rather than by bending the standard ones.
from opentelemetry import trace, metrics
tracer = trace.get_tracer("alchmai.agent")
meter = metrics.get_meter("alchmai.agent")
# Metric names come from the convention - do not invent your own, or you lose
# every dashboard and alert the ecosystem already ships.
token_usage = meter.create_histogram("gen_ai.client.token.usage", unit="{token}")
op_duration = meter.create_histogram("gen_ai.client.operation.duration", unit="s")
async def assess_alert(alert, cfg):
# Span naming convention: "{operation} {model}". Tooling relies on it.
with tracer.start_as_current_span(f"chat {cfg.model_snapshot}") as span:
span.set_attributes({
# --- GenAI semantic conventions (stable names) ---
"gen_ai.operation.name": "chat",
"gen_ai.provider.name": cfg.provider,
"gen_ai.request.model": cfg.model_snapshot,
"gen_ai.request.temperature": cfg.temperature,
"gen_ai.request.max_tokens": cfg.max_tokens,
# --- Our namespace, for the questions the convention cannot know ---
# Unit economics: cost per business outcome, not per call.
"alchmai.workflow": "surveillance.triage",
"alchmai.tenant": alert.desk_id,
# The decision fingerprint that makes behaviour reconstructible.
"alchmai.decision_fingerprint": cfg.fingerprint,
"alchmai.prompt_version": cfg.prompt_version,
"alchmai.policy_bundle": cfg.policy_bundle_version,
# Point-in-time basis for any market or reference data used.
"alchmai.as_of": alert.as_of.isoformat(),
})
result = await model.complete(...)
span.set_attributes({
"gen_ai.response.model": result.model, # may differ from request
"gen_ai.response.finish_reasons": result.finish_reasons,
"gen_ai.usage.input_tokens": result.usage.input,
"gen_ai.usage.output_tokens": result.usage.output,
"alchmai.abstained": result.parsed.get("abstained", False),
})
common = {
"gen_ai.operation.name": "chat",
"gen_ai.request.model": cfg.model_snapshot,
"alchmai.workflow": "surveillance.triage",
}
token_usage.record(result.usage.input, {**common, "gen_ai.token.type": "input"})
token_usage.record(result.usage.output, {**common, "gen_ai.token.type": "output"})
return resultAdopting Without Breaking Your Dashboards
Most GenAI conventions were still experimental as of early 2026, which sounds like a reason to wait and is not. The convention ships a migration mechanism precisely for this: the OTEL_SEMCONV_STABILITY_OPT_IN environment variable enables dual emission of both legacy and new attribute names, so you can move without a flag day.
- 01Turn on dual emission and leave it on for a full reporting cycle. Old dashboards keep working while new ones are built against the stable names.
- 02Rebuild alerts against the convention names first, because alerts are what wake people up and are the thing you least want silently broken.
- 03Pin the convention version in your build, as you would any dependency, and treat a version bump as a reviewed change rather than a transitive upgrade.
- 04Only then remove the legacy emission, and confirm nothing queries the old names by searching your dashboard definitions rather than by asking the team.
The Metrics That Actually Get Watched
A convention gives you attribute names, not judgement about what matters. Across the financial AI systems we operate, these are the measures that have earned a place on a dashboard with a named owner - and the ones that have caught real problems.
- Cost per completed business outcome, not cost per call. Pounds per triaged alert, per reconciled break, per drafted note. This is the only number a business owner can act on, and it is computable only if you tag spans with the workflow.
- Tokens per task, trended. A slow rise here is almost always context creep - a prompt that grew, a retriever returning more, a history that stopped being compacted. It is the earliest signal that unit economics are drifting.
- Abstention and refusal rate. When a model declines to assess, that is signal. A rise means something changed upstream - new document formats, a new instrument type, a degraded retriever - usually well before output quality visibly degrades.
- Tool call error and retry rates by server. With MCP in the picture, a slow or flapping server shows up as agent latency and is otherwise very hard to attribute.
- Deterministic check failure mix. For any system with a policy or risk layer below the model, the distribution of rejection reason codes is the best drift detector available, and it costs nothing to compute.
- p95 end-to-end latency split by stage. Model, retrieval, tool, and your own code. Without the split, every latency conversation is speculation.
“If you cannot say what a completed task costs and which stage its latency came from, you do not have an observability problem - you have an economics problem that will surface at renewal.”
Evals Belong In The Same Pipeline, And In CI
The conventions now extend to quality evaluation, and this is the part we think most teams are still missing. Observability tells you what happened; evals tell you whether it was right. Keeping them in separate systems means nobody can answer the combined question - did the change that improved latency also degrade quality? - which is the question that actually decides whether to ship.
The practice that has become standard among teams shipping reliable LLM systems in 2026 is simple to state and takes discipline to hold: every pull request that touches a prompt, a model version or a retrieval configuration triggers an eval run against a golden dataset, and a PR that regresses quality past a threshold does not merge. The teams building the most reliable applications are not the ones with the fanciest models; they are the ones who invested early in evaluation infrastructure and treat the eval dataset with the same care as production code.
name: evals
on:
pull_request:
paths:
- "prompts/**"
- "config/models.yaml"
- "retrieval/**"
- "policy/**"
jobs:
golden-set:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Run evals against the golden dataset
env:
# Traces from the eval run go to the SAME collector as production,
# tagged as an eval. One schema, one pipeline, comparable numbers.
OTEL_EXPORTER_OTLP_ENDPOINT: ${{ secrets.OTLP_ENDPOINT }}
OTEL_RESOURCE_ATTRIBUTES: "deployment.environment=eval,alchmai.pr=${{ github.event.number }}"
OTEL_SEMCONV_STABILITY_OPT_IN: "gen_ai_latest_experimental"
run: uv run python -m evals.run --suite surveillance --out results.json
- name: Gate on regressions
run: |
uv run python -m evals.gate results.json --baseline main --hard-fail confirmed_ranked_low=0 --hard-fail invalid_citations=0 --no-regress analyst_agreement --no-regress benign_demoted --budget cost_per_task_gbp=0.04 --budget p95_latency_ms=9000Building The Golden Set From Production, Not Imagination
One point that connects the two halves of this playbook. The golden dataset should be built from real production failures rather than synthetic examples, and a standardised telemetry pipeline is what makes that practical. When every decision is traced with a fingerprint, a workflow tag and a link to the eventual human disposition, promoting a production failure into the eval suite becomes a query and a click rather than a reconstruction project.
- 01Every human override or disagreement is a candidate. Capture the override at the point it happens, linked to the trace, and make adding it to the golden set a one-click action in the reviewer's own interface.
- 02Every abstention with an unclear cause is a candidate. These are the cases where the system knew it was out of its depth, which is exactly where the boundary needs defining.
- 03Every incident gets a regression case, permanently. This is the ratchet that stops the same failure recurring, and it is the single habit that most improves a system's reliability over a year.
- 04Keep a held-out slice that is used once, at major version changes, and publish the result whatever it says. A golden set that is optimised against stops measuring quality and starts measuring fit.
The Bottom Line
The instrumentation question for AI systems has a boring, correct answer now, and adopting it is one of the higher-leverage afternoons available to a financial engineering team. OpenTelemetry graduated from the CNCF in May 2026 and its GenAI conventions stabilised the schema for LLM spans, agent traces, metrics, agent orchestration, MCP tool calling and quality evaluation, with the major cloud and observability platforms adopting them. Use the standard names, add your finance-specific attributes in your own namespace, capture content references rather than content, and migrate with dual emission rather than a flag day. Then put the measures that matter on a dashboard with an owner - cost per completed outcome, tokens per task, abstention rate, tool error rate by server, deterministic check mix, and p95 split by stage - and run evals through the same pipeline, gating every PR that touches a prompt, a model version or a retrieval config against a golden set built from real production failures. Do that and the question a supervisor eventually asks - what did this system do, and how do you know it was right - has an answer that takes minutes rather than weeks. That is the production engineering we do for financial AI in London, and it is the least glamorous work with the highest return in the whole stack.
References & Further Reading
- OpenTelemetry - Semantic conventions for generative AI. opentelemetry.io/docs/specs/semconv/gen-ai
- OpenTelemetry - Inside the LLM call: GenAI observability with OpenTelemetry. opentelemetry.io/blog/2026/genai-observability
- Greptime - How OpenTelemetry traces LLM calls, agent reasoning and MCP tools. greptime.com/blogs/2026-05-09-opentelemetry-genai-semantic-conventions
- Datadog - Agent observability natively supports OpenTelemetry GenAI semantic conventions. datadoghq.com/blog/llm-otel-semantic-convention
- MLflow - OpenTelemetry GenAI semantic conventions. mlflow.org/docs/latest/genai/tracing/opentelemetry/genai-semconv
- OpenObserve - OpenTelemetry for LLMs: complete SRE guide for 2026. openobserve.ai/blog/opentelemetry-for-llms
- Braintrust - Best AI eval tools for CI/CD pipelines (2026 review). braintrust.dev/articles/best-ai-evals-tools-cicd-2025
- Confident AI - Best CI/CD tools for testing AI agents before production in 2026. confident-ai.com/knowledge-base/compare/best-ci-cd-tools-testing-ai-agents-before-production-2026
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information