You cannot audit an agent from its final answer. The record that matters sits underneath it: which tool ran, with what arguments, and what came back. OpenTelemetry defines those fields in a vendor-neutral specification, which means a two-person team can capture the same trace a large platform would, without buying a platform.
Why does the final answer tell you so little?
An agent's output is a summary of a process you did not watch. When it is right, you learn nothing about why. When it is wrong, the visible error is almost never where the fault is. A bad answer at step ten usually traces back to a tool call at step three that returned an empty result the model then papered over.
This is the part that surprises operators who have run ordinary software for years. Conventional monitoring answers one question: is the service up. Agents fail while every service is up. The API returns 200, the model responds, the latency looks normal, and the agent still quietly did the wrong thing, because it called the wrong tool, passed a malformed argument, or accepted a result it should have rejected.
We covered the specific shapes those failures take in what AI agents actually break on in a real back office. The pattern that matters here is simpler. The failures are invisible to infrastructure monitoring because they are decisions, not outages.
What should a small team actually record?
The honest answer is five fields, and they are already named for you.
OpenTelemetry's GenAI semantic conventions were split into their own repository in May 2026, and they define the tool-call surface explicitly:
gen_ai.tool.name, which tool was invokedgen_ai.tool.call.id, the identifier tying a call to its resultgen_ai.tool.call.arguments, what the model actually passedgen_ai.tool.call.result, what came backgen_ai.tool.type, the kind of tool it was Alongside those, the spec names five operations worth wrapping in a span:create_agent,invoke_agent,execute_tool,chat, andgenerate_content. Agent-level attributes covergen_ai.agent.id,gen_ai.agent.name,gen_ai.agent.version, andgen_ai.conversation.id, which is what lets you replay a whole session instead of a single hop.
That list is the entire answer to "what do I watch." It is not a product. It is a field list. If you are running two or three agents, you can emit those attributes from your own code into whatever store you already run, and you will be able to answer the only question that matters after an incident: which tool ran, with what, and what did it return.
Recording gen_ai.agent.version is the one teams skip and later regret. Without it you cannot tell whether a behavior change came from your own prompt edit or from a model update underneath you.
The spec also defines usage attributes, including gen_ai.usage.cache_read.input_tokens and gen_ai.usage.cache_write.input_tokens. Those let you attach a cost figure to a specific decision rather than to a monthly invoice, which is how you find the step that is quietly retrying an expensive call.
Why is the standard still marked Development?
Because it is not finished, and pretending otherwise would set you up badly.
Every gen_ai attribute in the registry carries a Development badge. The stable attributes appearing in the same span definitions are the ones inherited from the core OpenTelemetry conventions, pinned at v1.44.0: error.type, server.port, and their peers. In practice that means the GenAI names can still change, and anything built tightly against them will need maintenance.
That is an argument for owning the emission, not for waiting. The field list is a good description of what to capture regardless of what the keys are eventually called. Capture the data under the current names, keep the mapping in one place, and a rename becomes a small edit rather than a re-instrumentation project.
Do you need a platform for this?
Not to start. The open-source tooling is real. Langfuse, which integrates with OpenTelemetry, sits above 34,000 stars on GitHub and can be self-hosted, so it is a genuine option for a team with no monitoring contract.
The sequencing matters more than the tool. A platform is useful once you are already emitting structured traces. It is close to useless if your agents emit nothing but final answers, because it will faithfully store the same opaque output you already had. Instrument first, then decide whether you want somebody else to host the storage and the dashboards.
The practitioner conversation reflects the same unease. A thread in r/AI_Agents asking whether agents are actually better than deterministic workflows drew 65 upvotes and 68 comments, and the recurring argument is not that agents fail to work. It is that a workflow you can read beats a decision you cannot inspect. Instrumentation is what closes that gap. It makes the nondeterministic path legible after the fact, which is the property the deterministic workflow had for free.
Where the measurement layer pays for itself
Build the measurement layer before you expand the agent's authority. Every increase in what an agent is allowed to do multiplies the number of ways it can be quietly wrong, and the cost of a silent failure scales with how much you trusted it.
The order we recommend to clients is unglamorous. Instrument the tool calls. Run the agent narrowly for a few weeks with a human reading the traces. Only then widen its scope, using the traces to decide which steps have earned autonomy. That sequence also produces the evidence you need to justify the spend, which is the same discipline we apply when we design AI revenue systems rather than bolting an agent onto an existing funnel and hoping.
An agent you cannot inspect is not an asset yet. It is an experiment you have put in front of customers.