AUDR — Agent Usage Detail Record — is an open standard for recording who initiated an agent run and how much each cost, across every system a run passes through. It defines a common JSON schema that any harness, router, or billing system can emit and ingest, so a single agent run can be represented through records that share a common structure. AUDR was drafted at Chargebee, is licensed under Apache 2.0, and is stewarded by Chargebee, with the stated goal of moving cost governance to an independent foundation as adoption grows. It is useful anywhere you need a reliable record of what an agent run consumed and who or what it was associated with.
The problem AUDR addresses is that a single agent run touches multiple systems. The application knows the customer and the feature. The router knows the tokens and the cost. The tools know what they executed. As the project describes it, a run can be fully observable at every individual layer and still leave you without a single end-to-end record of who ran it and what it cost. Without a shared way to join these observations, usage data is orphaned from the business context that gives it meaning. The telecom industry solved an analogous problem with the Call Detail Record, an open standard carriers converged on so a call's attributes could be captured and exchanged in a common format, independent of any single carrier's systems. AUDR is built on the same principle: a common record for agent runs that any harness, router, or billing system can emit and ingest to help businesses make sense of the economics at the run level.
AUDR works through three rules. The first is a shared run ID, minted by the harness, passed to the router in request metadata, and echoed back, so that every system that touches the run carries the same ID. The second is clear authority per field: the harness owns attribution — customer, environment, initiator — while the router owns usage — tokens, provider. Each fact has exactly one source. A record carries the raw counts that drive cost, such as tokens, tool calls, and seconds of compute, alongside the business context that says whose cost it is: customer, feature, environment. Every layer keeps reporting what it already reports, and AUDR adds the rules that let those reports come together into one record.
The third rule is strict merge rules. The sink assembles records sharing a run and span ID, and no component rewrites another's block. Conflicts are rejected, and a correction is a new record, never a mutation. The documentation illustrates this with a sample record in which run.run_id is "run_8f2a1c" (minted by the harness) and span_id is "span_4b91"; attribution includes a customer_id of "acme-corp" and an initiator of "end_user", both sourced from the harness; usage includes llm input_tokens of 1204 and output_tokens of 318, sourced from the router; and the emitter component is "router". One record, one authoritative source per field.
Adapters capture records from the runtime you already use. The Core SDK builds, validates and delivers records straight from your own code, available in Python (audr) and TypeScript (@openaudr/audr), and every adapter and sink builds on it. NVIDIA NeMo Relay records completed LLM and tool scopes, with attribution read from the root scope's metadata (Python, audr-adapter-nemo-relay). LiteLLM registers as a callback on the SDK or Router and records completion, Responses API, embedding and rerank calls (Python, audr-adapter-litellm). Merge Gateway wraps the native SDK client and records every response, streamed or not, using the gateway's own token and cost report (TypeScript, @openaudr/audr-adapter-merge-gateway). Vercel AI SDK registers as an AI SDK 7 telemetry integration and records model, tool, embedding and rerank calls (TypeScript, @openaudr/audr-adapter-vercel-ai). Mastra registers as an observability exporter and records model, embedding and tool calls (TypeScript, @openaudr/audr-adapter-mastra). Adapters read identifiers, usage and timings, never prompts or outputs, and every package is Apache 2.0 and published to PyPI or npm.
Sinks deliver records to your destination. The flow is runtime to adapter to core client to sink to destination. The Chargebee sink delivers records to a Chargebee site's usage-ingest batch endpoint for usage-based billing (Python audr-sink-chargebee and TypeScript @openaudr/audr-sink-chargebee), and the Lago sink delivers records to Lago's batch event endpoint for usage-based billing (TypeScript @openaudr/audr-sink-lago). Running something else? The core SDK emits records directly from your own code, and any destination can be reached with a new sink. To try it, you register an adapter with the runtime you already use and get a usage record for every model and tool call, including the customer it belongs to; you can write the records to a local file to start, with no account, hosted backend, or pricing configuration needed.
AUDR is designed to sit on top of OpenTelemetry, not compete with it. OTel's GenAI semantic conventions provide the foundation for describing model calls and usage, and AUDR reuses them: an AUDR record can be emitted as an OTel span, and the OTel collector is a first-class sink. What OTel does not define is the set of rules needed when usage becomes a durable record — which attributes are required, how attribution is handled when it is missing, how retries remain idempotent, or how corrections are made. Observability can tolerate a dropped span; a usage record cannot, which is why AUDR adds those requirements and delivery semantics on top. FOCUS solves a different part of the same problem: it standardizes the billing data you receive from providers so costs from AWS, Azure and others can be represented in a common schema, while AUDR standardizes the usage you emit when an agent run happens, before that usage is priced. The two are complementary, and AUDR records can be rated by any backend and mapped into FOCUS-compatible cost data, completing the upstream half of an existing standard.
The practical benefit is being able to answer concrete questions about agent economics at the run level. Wrap your router, emit the records, and AUDR can help you answer questions such as: How much does this agentic feature cost? What does this customer's agent usage look like, and how much does it cost? What are the unit economics and margins per customer for my agentic features? Which workflows or models are driving our costs? Which power users are driving our costs? AUDR adds nothing in the normal request path: it emits records asynchronously and out of band, so recording usage does not add synchronous work to inference. The one exception is optional pre-flight budget gating, which would make a single check before a run starts.
Because the spec carries no prices or rating logic and the SDK has no concept of plans, invoices, or how a customer should be charged, AUDR records what happened and who it happened for, leaving what you do with that data up to you. You can point the records at Chargebee, a competing rating engine, your own, or a warehouse for analytics, and AUDR works the same way. You do not need a billing system to use it: records can be stored locally, sent to your warehouse, fed into an observability system, or used for internal cost analysis or future projections. A billing system is just one possible consumer of the record.
Today five adapters, two sinks and the core SDK are published, in Python, TypeScript or both: adapters for NVIDIA NeMo Relay, LiteLLM, Merge Gateway, Vercel AI SDK and Mastra, and sinks for Chargebee and Lago. Support for OpenRouter is in development. The three rules at the core of AUDR are stable — one run ID across every layer, one authoritative source per field, and strict merging with no silent overwrites — and will not change without a major version, while the field set will continue to grow as providers introduce new things to measure. The project invites involvement: read the spec for the full schema, field ownership rules and delivery semantics; write an adapter for a harness or router not yet reached, which the project describes as roughly 200 lines against the shared fixtures; write a sink for a warehouse, ledger or billing system you already deliver usage to; or open an issue with a specific account of where a design decision breaks. Questions can be sent to audr@chargebee.com.
In short, AUDR is an open, Apache 2.0 standard that turns fragmented per-layer observability into one joined record of who initiated an agent run and what it cost, giving teams building and monetizing agents a neutral, vendor-independent foundation for understanding agent economics and cost governance.