> Connecting your agentic system to Emet: what Emet checks and what you get back, connecting your data, capturing a turn over OpenTelemetry or JSON, building the exporter, edge cases, going live and handling errors. _INTEGRATION GUIDE_ # Connecting your agentic system to Emet This guide details the path from an Emet account to a verified agent in production. It covers what Emet checks and what you get back, connecting your data, what you build, the data you send, edge cases, going live, and handling errors. > **The general guide, and yours** — This is Emet's general integration guide. Once we align on a pilot scope with your team, our onboarding process — questionnaires and a walk-through of your systems — produces an integration plan specific to your agents and data sources. Your team can review that plan at any time inside your Emet dashboard. **[Request access →](https://emet.so/contact)** _SECTION 1_ ## At a glance ### What Emet does for your agent Your agent answers questions with figures, comparisons and trends drawn from its tools. Emet examines each answer, recomputes its figures and calculations in deterministic code, and checks every claim against the evidence your agent actually received and, with your data connected, against the source itself. No model ever judges its own output. Every response your agent makes receives an Emet Accuracy Report: an assessment of its accuracy and completeness. Inaccurate figures or claims are logged, and Emet provides a clarification or correction based on the available evidence. This way you find out which answers hold up, which ones do not, and exactly which claim, tool result, and model were responsible. ### What you get back Every checked answer produces an Emet Accuracy Report in the Emet dashboard, where everyone on your team can log in and review it at any time. Each report shows every claim in the answer with its verdict and reason, the evidence behind it, and the calculation behind every derived figure. Each claim receives one of four verdicts, with its reason shown alongside it. | Verdict | What it means | | --- | --- | | Verified | The claim matches the evidence. | | Clarified | The claim is right but missing context a reader needs, such as its scope or time window. Emet adds it. | | Corrected | The claim is wrong, and the evidence establishes the right answer. Emet auto-corrects it and shows the correction against the original wording. | | Flagged | The claim cannot be supported or corrected from the evidence. Emet flags it rather than guess (or hallucinate an answer). | ### Example: right data, wrong reading The question was “How did CTV perform against display last week?” The agent's SQL was correct and returned four rows: | Channel | Week of | Conversions | CPA | | --- | --- | --- | --- | | CTV | 31 August | 1,248 | $50.00 | | CTV | 7 September | 1,340 | $43.96 | | Display | 31 August | 1,030 | $40.00 | | Display | 7 September | 1,075 | $41.67 | > **The agent answered** — “CTV conversions rose 7% week over week to 1,340, the most of any channel. Its CPA fell to $43.96, now below display. The lift was driven by the new creative.” The Accuracy Report showed: | Claim | Verdict | Why | | --- | --- | --- | | “CTV conversions rose 7% week over week to 1,340” | Verified | The result shows 1,248 rising to 1,340, an increase of 7.4%. | | “the most of any channel” | Clarified | True of the two channels the query covered. Emet adds the scope: “the most of the two channels queried.” | | “Its CPA fell to $43.96, now below display” | Corrected | Display's CPA was $41.67 the same week, so CTV's is still above it. Emet corrects it to “Its CPA fell to $43.96, still above display's $41.67.” | | “The lift was driven by the new creative” | Flagged | Nothing in the result describes creative, so the cause can be neither supported nor corrected from this evidence. | > **What this shows** — The query was right, and every number in the answer came from it. One claim still needed context, one was wrong, and one had no support. Accurate data does not guarantee an accurate reading, and that gap is what Emet checks. ### What is AI observability? AI agents are non-deterministic: the same question can take a different route each time, so you can't test every behavior before launch. Traditional monitoring tells you an agent is running, what it cost and how many tokens it used. AI observability goes further and shows you what your agent actually did in production: what it was asked, which tools it called, what came back, and what it finally said. Neither tells you whether the answer was right. This works through instrumentation, which means adding a small amount of logging to your agent, typically through an SDK or an integration with your existing framework. After that, every run is captured automatically as structured data, organized in three levels: - **Conversation** — The full multi-turn exchange with a user. - **Trace** — A single turn, meaning one question, the agent's work, and its response. - **Span** — One step within a trace, such as a sub-agent or a tool call and its result. Because Emet sees the tool results your agent received at every step, and can query your connected data sources directly, it checks each claim in the final answer against both the evidence your agent saw and the source itself. Observability shows you what your agent did. Emet tells you whether what it said is true. ### What you build Connect your agent to Emet over OpenTelemetry (OTLP). If it already emits OpenTelemetry traces, point its exporter at Emet with a configuration change. No OpenTelemetry yet? Send each turn as a simple JSON log using our reference exporters for TypeScript and Python ([section 5](https://emet.so/docs#building-the-exporter)). ### What you do not have to touch **Nothing in your request path changes** - **No proxy required.** — Your model calls go straight to your provider, exactly as today. - **One SDK, at most.** — Already on OpenTelemetry? It's a config change. If not, one reference exporter for your language sends plain HTTPS and JSON, under your control. - **No inbound access to your agent.** — Emet never calls into your application. Traces go one way, from your exporter to Emet. The only thing Emet connects to is the data source you choose to connect: read-only, scoped by you, revocable at once ([section 3](https://emet.so/docs#connecting-your-data)). - **No added latency.** — Exports run in the background, after your user already has the answer. - **No production risk.** — An export that fails is logged and dropped. It never reaches the request that served your user. ### How it fits Your agent works exactly as it does today. Emet only enters the picture after your user already has the answer, and reads your connected data only to check it. _Figure: Sequence: the user asks a question; the agent calls its tools with arguments and gets results; the agent answers the user. Only then does the agent export the turn to Emet over OTLP and receive an acceptance. Emet runs read-only queries against your data, gets the source results, checks each claim, and produces the Accuracy Report._ ### The path to a first report _Figure: Six steps in order: 1. Set up, 2. Connect, 3. Instrument, 4. Export, 5. Check, 6. Roll out._ | Step | What happens | Who | Section | | --- | --- | --- | --- | | 1. Set up | Create a team, register the application, mint a key | Whoever owns the Emet account | [2](https://emet.so/docs#before-you-start) | | 2. Connect | Connect your data warehouse in the Emet dashboard, read-only | Whoever administers your warehouse | [3](https://emet.so/docs#connecting-your-data) | | 3. Instrument | Confirm your agent emits OpenTelemetry GenAI traces with content capture on, or choose the JSON route | The engineer who owns the agent | [4](https://emet.so/docs#capturing-a-turn) | | 4. Export | Point your OTLP exporter at Emet, or hook in a JSON exporter after each turn | The same engineer | [4](https://emet.so/docs#capturing-a-turn) and [5](https://emet.so/docs#building-the-exporter) | | 5. Check | Send real turns and read the first reports | The same engineer | [7](https://emet.so/docs#going-live) | | 6. Roll out | Move from a trial application to full traffic | Your team | [7](https://emet.so/docs#going-live) | Steps 3 and 4 happen in your codebase. With OpenTelemetry already in place, they are a configuration change. _SECTION 2_ ## Before you start Two short lists: the answers your team needs before writing code, and the three things to create in the Emet dashboard. > Already in a pilot? Your integration plan in the Emet dashboard answers these questions for your systems — start there. ### Questions to answer first Every one of these is about your own system, and each one maps to a field Emet needs. | Question | Why it matters | Becomes | | --- | --- | --- | | Does your agent already emit OpenTelemetry traces? If not, which framework runs it? | If it does, sending to Emet is a configuration change. If not, most frameworks already hold the whole turn in a message history, which makes a JSON exporter short. | OTLP, or which reference exporter to start from | | Where does a turn finish in your code? | On the JSON route, that is where the exporter hooks in, after the answer has gone out. | The place you call the exporter | | Which identifier tells you whose work a turn is? Usually a workspace, account or client id your agent already holds. | Emet attributes every turn to a team by it. It cannot be added to traces already sent. | `tenant_ref` | | Which identifier groups the turns of one conversation? | Emet shows a turn next to the ones around it. | `conversation_id` | | Can you read each tool call's arguments as well as its result? | A claim can only be traced back to the query that produced it if the arguments were kept. This cannot be added later either. | The arguments on each tool span | If your agent already records OpenTelemetry GenAI spans with content capture on, most of this data is already flowing, and [section 4](https://emet.so/docs#capturing-a-turn) shows how to point it at Emet. ### Set up the account All three steps are in the Emet dashboard under Applications. - **01 · Create a team** — One for each part of your company whose work should be reported separately. Give each one the external ID your agent will send as `tenant_ref`. - **02 · Register the application** — Meaning the agent system that will send traces. A name and a short description are enough. - **03 · Mint a key** — For the application, leaving the team blank, so each trace is attributed by its own `tenant_ref`. The key is shown once; store it the way you store any other production credential. ### One application per environment Register a separate application for each environment that will send traces, such as “Nova Analyst (staging)” and “Nova Analyst (production)”, each with its own key. Trial traffic and real traffic stay apart on every report, rate limits are counted per application so a load test never throttles production, and revoking a staging key never touches production. Keys can carry an expiry date, rotate in pairs without downtime, and can be revoked at once. ### What to hand your engineer - The Emet API host that comes with your account, and a key for the trial application. - The external ID of at least one team. - The answers to the questions above. _SECTION 3_ ## Connecting your data Your traces tell Emet what your agent saw. A data connection lets Emet check it against the source. With your warehouse connected, Emet runs its own read-only queries to confirm that the results your agent received match your data, that the query asked the right question of it, and that each claim in the answer holds up against the source, not only against the trace. Emet connects to Databricks and other major data warehouses. You connect yours in the Emet dashboard, which walks you through it. ### What the connection is - **Read-only.** — Emet runs read queries only. It never writes, alters or deletes. - **Scoped by you.** — Emet sees only the data you grant it access to. - **Revocable at once.** — Remove Emet's access in your warehouse, and it stops immediately. - **Visible in your own logs.** — Every Emet query runs through the connection you create, so it appears in your warehouse's query history alongside everything else. ### How Emet learns your data During onboarding, Emet maps your definitions: which tables hold which metrics, your attribution windows, fiscal calendar, currencies and time zones. That way its queries ask the same question your agent did, by your rules. ### Before you connect | Topic | Detail | | --- | --- | | Compute | Emet's queries run on the warehouse compute you connect, so you control where they run and what they cost. | | Credentials | Stored encrypted and used only for verification queries. You can rotate or revoke them at any time. | | What Emet reads | Only the data behind the claims in your agent's answers. Emet re-runs the numbers; it does not browse or copy your tables. | | Query results | Kept with the report they support, as the evidence behind each verdict. | _SECTION 4_ ## Capturing a turn Emet takes a turn in one of two ways. OTLP, OpenTelemetry's standard protocol, is the primary one. Agents without OpenTelemetry can send each turn as a simple JSON log instead. Both carry the same evidence: every model call and tool call in the turn, with what went in and what came out. ### Sending over OTLP Point your OpenTelemetry exporter's traces endpoint at Emet and pass your application key in the exporter's headers as `Authorization: Api-Key `, the same way you would send traces to any other tracing backend. Emet reads the OpenTelemetry GenAI semantic conventions, so an agent that already records them connects with a configuration change. Emet reads the GenAI attributes below. Each one maps to a field in the JSON envelope, so both routes carry the same evidence: | OpenTelemetry attribute | Emet field | | --- | --- | | `gen_ai.input.messages`, `gen_ai.output.messages` | `input_messages`, `output_messages` on a model span | | `gen_ai.system_instructions`, `gen_ai.tool.definitions` | `system_instructions`, `tool_definitions` on a model span | | `gen_ai.tool.call.arguments`, `gen_ai.tool.call.result` | `arguments`, `result` on a tool span | | `gen_ai.retrieval.query.text`, `gen_ai.retrieval.documents` | `query`, `documents` on a retrieval span | | `gen_ai.memory.query.text`, `gen_ai.memory.records` | `query`, `records` on a memory span | | `gen_ai.conversation.id` | `conversation_id` | #### What your spans need - **Content capture on.** — Without it, messages, tool arguments and results are never recorded, and Emet has no evidence to check. - **The turn's team.** — Each turn carries the team's external ID from [section 2](https://emet.so/docs#before-you-start), so its report lands with the right team. - **The turn's conversation.** — Set `gen_ai.conversation.id`, so Emet can show a turn next to the ones around it. - **Tool arguments as well as results.** — A claim can only be traced back to the query that produced it when the arguments are there. - **Your system prompt stays private.** — Standard instrumentation records it as text. To keep it private, hash or remove `gen_ai.system_instructions` before export, as the JSON route does. To confirm it works, send a few real questions through the trial application and check that each one appears as a report in the dashboard, with every claim naming its evidence ([section 7](https://emet.so/docs#going-live)). ### Sending JSON directly Without OpenTelemetry, your exporter sends Emet one JSON document per turn, called the envelope. It names whose work the turn was and carries the turn as a small tree of spans. Every model call and every tool call is its own span, and all of them hang off the turn: a turn with five tool calls across three model calls has nine spans. Every field and limit below is what the endpoint enforces today. #### The request | Part | Value | | --- | --- | | Method and path | `POST /api/v1/traces/` | | Authentication | `Authorization: Api-Key`, followed by your application key | | Body | The envelope, as JSON | _Posting an envelope_ ```bash curl -X POST "https:///api/v1/traces/" \ -H "Authorization: Api-Key $EMET_APPLICATION_KEY" \ -H "Content-Type: application/json" \ --data @envelope.json ``` #### The envelope | Field | Required | Meaning | | --- | --- | --- | | `tenant_ref` | Yes | The team's external ID. A value that matches no team is accepted and marked unattributed. A request without it is rejected. | | `conversation_id` | Yes | Groups the turns of one conversation. | | `trace_id` | Yes | Identifies this turn. Resending the same `trace_id` adds to the same trace. | | `ingestion_mode` | Yes | `"sdk"` when your exporter builds the envelope from your agent's own state, as the reference exporters do. `"telemetry"` when it is converted from OpenTelemetry spans. | | `spans` | Yes | The spans of the turn, 1 to 1,000 per request. | | `end` | No | `true` on the last request of a turn, so Emet starts checking at once. A request that carries the turn span has the same effect. | | `semconv_version` | No | Which vocabulary your payload keys follow, stored with the trace so it stays readable as that vocabulary changes. An exporter that builds the envelope from its own state sends `"emet-genai-1"`. One that converts OpenTelemetry spans sends the GenAI conventions version its instrumentation emits. | | `session_ref` | No | A non-secret id that groups the turns of one agent session. Never a login token or cookie. | | `resource` | No | Describes the sending service: its name, version and deployment environment. | #### Rules the endpoint enforces - `conversation_id`, `trace_id` and `span_id` accept letters, digits and the characters `_ . : -` only, and an id made only of dots is rejected. They may be up to 255, 128 and 64 characters long. - Every timestamp carries a UTC offset, and no span ends before it starts. - A `span_id` appears once per request, and only one span per trace is the final output. - A request carries at most 1,000 spans and 32 MB. `attributes` and `resource` hold small metadata: 16 KiB each. - Text must not contain a NUL character, and numbers must be finite. #### Span fields | Field | Required | Meaning | | --- | --- | --- | | `span_id` | Yes | Unique within the trace, such as `model.0` or `tool.call_1`. | | `parent_span_id` | No | Empty for the turn, the turn's `span_id` for everything else. | | `name` | Yes | A readable name, such as `"chat"` or `"execute_tool get_spend"`. | | `kind` | No | `turn`, `model`, `tool`, `retrieval` or `memory` for most agents. | | `started_at`, `ended_at` | Yes | ISO 8601 timestamps with a UTC offset. | | `status` | No | `ok` or `error`. Set `error` on a tool that failed. | | `is_final_output` | No | `true` on the one model span that produced the answer. | | `model_name`, `provider`, `tokens`, `finish_reason` | No | For model spans. | | `tool_name`, `tool_call_id` | No | For tool spans. `tool_call_id` links the span to the model's call. | | `error_message` | No | What went wrong, on a span with status `error`. | | `attributes` | No | Small key-value metadata of your own, such as which retry attempt a span belongs to. | | `payload` | No | The content Emet verifies. See below. | #### The payload The payload is the evidence, and it is the part of the envelope that matters most. Its keys depend on the span kind: | Span kind | Payload keys | Holds | | --- | --- | --- | | `tool` | `arguments`, `result` | What the tool was called with and what it returned. Each is a JSON object or a string. | | `model` | `input_messages`, `output_messages` | The conversation the model saw, in order, and what it produced: text, tool calls, or both. | | `model` | `system_instructions`, `tool_definitions` | Optional. A hash of your system prompt (never its text), and the tools the model was offered. | | `retrieval` | `query`, `documents` | What was searched for and what came back. | | `memory` | `query`, `records` | What was read from or written to memory. | Each message is a role (`user`, `assistant` or `tool`) and a list of typed parts: `text` for anything the user or the model wrote, `tool_call` for the model asking for a tool (id, name, arguments), and `tool_call_response` for the tool's answer as the model saw it (id, response). ### A worked example One question, one tool call, one answer: “What did we spend on paid search last week?” _Figure: The turn span carries timing only. Under it hang three spans: model.0, which calls a tool; tool.call_1, which carries the arguments and the result; and model.1, which writes the answer and is marked as the final output._ | Span | Kind | What it carries | | --- | --- | --- | | `turn` | `turn` | The start and end of the whole turn. No payload. | | `model.0` | `model` | Input: the user's question. Output: a call to `get_spend` for paid search, 7 to 13 September. | | `tool.call_1` | `tool` | Arguments: channel paid search, 7 to 13 September. Result: spend 48,210.55 USD. | | `model.1` | `model` | Input: the question, the tool call and its result. Output: “Paid search spend last week was $48,210.55.” Marked as the final output. | > **Result** — Emet splits the answer into one claim, “spend was $48,210.55”, finds it in the tool result, and marks it Verified. The span below is the one that carries the evidence. The arguments say which question was asked of the data; the result is what came back. _The tool span from this example_ ```json { "span_id": "tool.call_1", "parent_span_id": "turn", "name": "execute_tool get_spend", "kind": "tool", "status": "ok", "tool_name": "get_spend", "tool_call_id": "call_1", "started_at": "2026-09-14T09:12:05.020Z", "ended_at": "2026-09-14T09:12:06.410Z", "payload": { "arguments": { "channel": "paid_search", "start": "2026-09-07", "end": "2026-09-13" }, "result": { "spend": 48210.55, "currency": "USD" } } } ``` _SECTION 5_ ## Building the exporter This section is for the JSON route. If your agent emits OpenTelemetry, your existing exporter already does this work ([section 4](https://emet.so/docs#capturing-a-turn)). The exporter is one function in your codebase. It turns a finished agent turn into the envelope and posts it. ### Three steps, any stack - **01 · Collect during the turn.** — For each model call keep the messages in and out, the model name, token counts and timing. For each tool call keep the name, the arguments, the result and timing. Most agent frameworks already hold all of this in the run's result once the turn completes. - **02 · Build the envelope when the turn completes.** — One model span per model call and one tool span per tool call, all under one turn span. Mark the model span that produced the answer as the final output. - **03 · Send it in the background.** — After the answer has gone out, post the envelope with a short timeout. ### Design rules | Rule | Why | | --- | --- | | Export after the answer is sent, never before. | Your users see no added latency. | | Never let the export throw into the request. | A failed export is a missed report, not a failed answer. Log it and move on. | | Use a short timeout, around 8 seconds, and retry only 429 and 5xx, briefly. | A slow network never holds a worker open, and a passing hiccup does not cost a report. | | Take `tenant_ref`, the conversation and its history from your own authorization and storage, never from the request body. | Whoever controls `tenant_ref` decides which team's reports a turn lands in, and history taken from the browser could plant evidence the model never saw. | | Keep tool arguments, not only results. | A claim can only be traced back to its query when the arguments are there. | | Send a hash of the system prompt, not the text. | It answers “did these turns run under the same prompt?” without sending the prompt itself. | | Send no end-user identifiers, credentials or connection strings. | Emet needs the figures, not your secrets. Strip them from tool arguments, results and error messages. | | Replace files and images with a short descriptor: media type, size and digest. | Emet verifies figures, not binary content, and the request stays small. | ### The reference exporters We maintain two reference exporters that apply every rule above. Your team can adapt them rather than write one from scratch; we share the source files with you. | Framework | Tested with | Reads the turn from | Where it runs | | --- | --- | --- | --- | | Vercel AI SDK (TypeScript) | `ai` 6.0.116, strict type checking | Each step of `result.steps`. The fields it reads mean the same in AI SDK 6 and 7. | In Next.js, inside `after()`, which runs once the response has been sent. | | Pydantic AI (Python) | `pydantic-ai` 2.51.0 | The run's message history: this turn's messages plus the earlier turns the model saw. | In FastAPI, as a background task, which runs after the response has been sent. | | Any other framework | — | The run's own record of the turn: every framework keeps one. | Any background task or queue that runs after the reply. | In a Next.js route, the only change is to note when each step finishes, then hand the finished run to the exporter inside `after()`, which runs once the response is out. _Hooking into a Next.js route (Vercel AI SDK)_ ```typescript const result = await generateText({ model, system, tools, messages, onStepFinish: () => { stepEndedAt.push(new Date()); }, }); after(() => sendToEmet(buildEnvelope({ tenantRef: workspace.id, // from your auth, never the request body conversationId: thread.id, traceId: crypto.randomUUID(), startedAt, endedAt: new Date(), system, messages, steps: result.steps, // every model call, tool call and result stepEndedAt, }))); ``` In FastAPI, the agent runs as today. The finished run goes to a background task, which FastAPI starts only after the answer has been returned. _Hooking into a FastAPI endpoint (Pydantic AI)_ ```python thread = await load_thread(user, question.thread_id) # history from your storage result = await agent.run(question.text, message_history=thread.messages) await save_thread(thread, result.all_messages()) background.add_task(export_turn, result, user.workspace_id, thread.id, started_at) return {"answer": result.output} ``` ### Test before you ship Point the exporter at your trial application and run a handful of real questions, including one where a tool fails and one that takes several tool rounds. Each post should come back `202` with no capture gaps. ### What the reference exporters handle | Situation | What the exporter does | | --- | --- | | A turn with several rounds of tool calls | Rebuilds, for every model call, exactly the conversation the model saw at that point. | | Earlier turns in the same conversation | Includes them in the history the model saw, and emits spans for the current turn only. | | Parallel tool calls | One tool span per call, each linked to the model's call by its id. | | A tool that fails | Marks the span as an error, keeps the arguments, and sends no result. Emet does not count it as a capture gap. | | A tool that returns nothing | Sends an explicit empty result, so the call is not mistaken for missing evidence. | | Lists, numbers, decimals, dates or ids | Converts them to plain JSON that Emet accepts. | | Files and images returned by a tool | Replaces them with media type, size and SHA-256 digest. | | Structured output (Pydantic AI) | Sends the structured answer as the answer's text, so its figures are checked. | | Tools that need approval (Pydantic AI) | Keeps the original call's arguments when the approved call runs in a second pass, and sends a denied call as a failed one. | | The system prompt | Sends a SHA-256 hash that stays the same across every call and turn under the same prompt. | | Identifiers from your own systems | Replaces characters Emet does not accept and keeps them within the length limits. | | A 429 or a 5xx from Emet | Tries up to three times, waiting at least as long as Emet asks. Gives up on a wait longer than ten seconds so it never holds your worker. | > **One case is not covered** — With Vercel AI SDK tool approvals, an approved call runs before the first step of the next call and appears in no step, so the reference exporter does not send it. _SECTION 6_ ## Edge cases Real agents retry, stream, call tools in parallel and hand work to sub-agents. None of it needs special support from Emet, but each case has one right way to show up in the envelope. | Case | How to send it | | --- | --- | | A tool call that fails | Send the span anyway, with status `error` and what went wrong in `error_message`. Keep the arguments and leave the result out. A span with no result and no error status is reported as a capture gap, because Emet cannot tell a failure from a result that was never recorded. | | Parallel tool calls | One tool span per call, each with its own `tool_call_id`, all under the turn. | | Retries and fallback models | One model span per attempt, each with its own model name, timing and status. Only the attempt that produced the answer is the final output. A tool call made by a superseded attempt still happened, so send it. | | Streaming | Stream to your user as usual and export once the stream has finished, from the finished steps or message history rather than the text chunks: the chunks carry prose, the steps carry the evidence. | | Long multi-step loops | One span per call, all under the same turn. A request carries up to 1,000 spans. | | Sub-agents | Add a span of kind `agent` for the sub-agent and point its model and tool spans at it. If the sub-agent's work is its own conversation with its own answer, export it as its own turn. | | Long conversations | The history grows with every turn, and that is useful: an earlier tool result can be the evidence for a later claim. The only ceiling is 32 MB per request. | | Very large tool results | Send them whole; they are evidence. If one turn truly exceeds 32 MB, split its spans across several requests with the same `trace_id`. | | One turn in several requests | Post batches with the same `trace_id`; order does not matter and duplicates are stored once. Emet starts checking as soon as a request carries the turn span or sets `end`, so send those only with the last request. Otherwise it checks 30 seconds after the last span arrives, and anything later is stored but not checked again. | | A turn that belongs to no team | It is accepted and marked unattributed, and filed under your organization. Worth an alert on your side: it usually means a new client whose team does not exist yet in Emet. | A failed tool call carries its status, the message and the arguments. No result, and no capture gap. _A failed tool call (fragment: name and timing omitted)_ ```json { "span_id": "tool.call_2", "kind": "tool", "status": "error", "tool_name": "get_budget", "error_message": "budget API timed out", "payload": { "arguments": { "channel": "paid_search" } } } ``` _SECTION 7_ ## Going live How to know the integration is right before it carries real traffic, and what to watch once it does. ### Check the first turns Send a handful of real questions through the trial application, including one that makes a tool fail, and check each one. The first four checks read the JSON endpoint's response. Over OTLP, confirm your exporter logs each export as successful, then use the last three. | Check | Where | Healthy | | --- | --- | --- | | The post is accepted | Your exporter's log | `202` | | Tool arguments and results arrived | The `gaps` in the response | Both lists empty | | The turn found its team | `unattributed` in the response | `false` | | The turn was complete when checked | `late` in the response | `false` | | The report exists | Reports in the dashboard | One row per turn, named by the question | | Claims name their evidence | The open report | Each claim names the tool its evidence came from | | Teams are right | The report's team line | Changes with who asked, never the same team for everyone | _A healthy response_ ```json { "trace": "5f0c2b8e-1d4a-4c7e-9a31-2f6b8d0e4c19", "conversation": "a3e91f47-6b2d-4f08-8c5e-7d1a9b3c6e20", "accepted": 4, "late": false, "unattributed": false, "gaps": { "missing_tool_arguments": [], "missing_tool_results": [] } } ``` ### Roll out in stages - **01 · Trial application.** — Real questions, trial key, until every check above passes. - **02 · Production application, part of the traffic.** — A new application and key for production. Export a share of turns first if you want to watch volume on your side. - **03 · All traffic.** — Export every turn. Reports accumulate per turn and stay available in your dashboard. ### What to monitor | Signal | What it usually means | | --- | --- | | Responses other than `202` | A mapping bug (`400`), a key problem (`403`), or a burst over the rate limit (`429`). | | Non-empty `gaps` | A tool whose arguments or result stopped being captured, often after a framework upgrade. | | `unattributed: true` | A client whose team does not exist yet in Emet, or an identifier that changed. | | `late: true` | Spans sent after the turn was already being checked. | The rate limit is 10,000 requests an hour per application. If your production traffic will run above that, talk to us before you switch on step 3. ### Keys over time Keys rotate in pairs, so a rotation never needs a coordinated deploy: mint the new key, deploy it at your own pace, then revoke the old one. Revoking is immediate, which also makes it the fastest way to stop an export. ### After a framework upgrade Agent frameworks change how they record a turn between major versions. After upgrading yours, send a few turns through the trial application and repeat the checks above before the upgrade reaches production. _SECTION 8_ ## Handling errors Every response Emet can send, what it means, and what your exporter should do about it. OpenTelemetry exporters handle these on their own, so this section matters most on the JSON route. | Status | Meaning | What to do | | --- | --- | --- | | `202` | Accepted. The spans are stored and the turn will be checked. | Read `gaps`, `unattributed` and `late`, and log anything unexpected. | | `400` | The envelope did not validate. The body names every field at fault. | Log the body and fix the mapping. Retrying unchanged fails the same way. | | `403` | The key is missing, unknown, expired or revoked. | Check the key. Every key problem gets the same response on purpose. | | `413` | The request is over 32 MB. | Split the turn across several requests with the same `trace_id`. | | `429` | Over 10,000 requests an hour for this application. | Wait for the `Retry-After` header, then retry. | | `5xx` | Emet could not take the request. | Retry with backoff. Duplicate spans are dropped, so a retry never double-counts. | ### What an error body contains | Field | Present on | Holds | | --- | --- | --- | | `error` | Every error | A readable message, such as “Invalid input.” | | `code` | Every error | A stable code, such as `invalid`, `throttled` or `request_too_large`. | | `status_code` | Every error | The HTTP status. | | `details` | `400`, `403`, `429` | For a `400`, the field errors, with each span's position in the `spans` list. | | `limit_bytes` | `413` | The request size limit. | Below, the first span was sent with a timestamp that has no timezone. `details` points at span 0 and names the field. _A real 400_ ```json { "error": "Invalid input.", "code": "invalid", "status_code": 400, "details": { "spans": { "0": { "started_at": ["Timestamps must carry a timezone offset."] } } } } ``` ### Typical field errors | Where | Message | | --- | --- | | `tenant_ref` | This field is required. | | `spans`, position 0, `started_at` | Timestamps must carry a timezone offset. | | `spans`, position 2, `payload` `result` | Expected an object or a string. | | `trace_id` | Identifiers may only contain letters, digits, and the characters `_ . : -` and may not be dots alone | | `spans` | Only one span per trace may be the final output | ### Troubleshooting | Symptom | Likely cause | Fix | | --- | --- | --- | | `400` on a tool result or arguments | A tool returned a list or a bare number | Serialize it to a JSON string, or wrap it in an object. | | `400` on a timestamp | A timestamp without a UTC offset | Send ISO 8601 with `Z` or an explicit offset. | | `400` on an identifier | An upstream id with spaces, slashes or other characters | Replace characters outside the allowed set before sending. | | `202`, but `gaps` lists tool spans | Tool arguments or results were not captured | Read them from the framework's step or message history, not from the streamed text. | | `202`, but no report to open | The turn had nothing to check, or only metadata arrived | Make sure the payloads carry messages and tool results. | | Every report lands on the same team | `tenant_ref` is hardcoded in the exporter | Send the identifier of whoever the turn was for. | | `unattributed: true` on every turn | The team's external ID does not match what the agent sends | Compare the team's external ID in the dashboard with the value in the envelope. | | `late: true` | Spans arrived after the turn was already being checked | Send the turn span, or `end`, only with the last request of a turn. | | Claims without a verdict, data not captured | The tool result behind the claim is missing | Check that every tool span carries its result. | > **Your exporter should never fail loudly** — Whatever Emet answers, the export must not change what your user sees. The reference exporters retry briefly on a 429 or 5xx and otherwise log the error and return. Nothing they do can reach the request that served your user, and that is the pattern to keep. --- ## Stop babysitting AI. Start trusting it. Emet delivers deterministic verification for enterprise AI data analysis — in seconds, with a receipt. **[Book a Demo →](https://emet.so/contact)** --- Emet · Truth infrastructure for AI data systems · https://emet.so