> Observability shows you what your agent did. Emet shows whether what it said is true — it asks your data the same question and re-derives every number in code, with the math. Nothing in your request path. _PLATFORM_ # Code checking AI. Not another model grading the answer. Observability shows you what your agent did. Emet shows whether what it said is true — it asks your data the same question and re-derives every number in code, with the math. **[Book a Demo →](https://emet.so/contact)** [Read the integration guide →](https://emet.so/docs) _THE LOOP_ ## The answer goes out. The check begins. - **Your agent answers your user** - **The turn is exported over OpenTelemetry, after the answer is sent** - **Emet re-runs the query on your data (read-only)** - **Every number re-derived in code** - **Every sentence checked against the result** - **Accuracy Report + Slack alert** _Never in the request path · Read-only access you control · No model grades its own output_ _ANATOMY OF A CATCH_ ## One answer. Five checks. _QUESTION "What was our Q2 return on ad spend on Meta?" · AI ANSWER "Q2 ROAS was 3.46x — our strongest quarter yet. Increase spend."_ - **What your agent did.** — Emet reads the finished turn over OpenTelemetry — every tool call, the arguments it was called with, and the raw result that came back. - **Audit the query.** — A perfectly grounded answer built on a wrong query is still wrong. Emet checks the question your agent asked of the data, not just the number it got back. - **Re-derive from your data.** — Emet holds its own read-only access to your data and asks your warehouse the same question, correctly scoped. It never asks the AI what the data said. - **Check every sentence.** — Every factual claim gets a verdict and a reason. Code decides the numbers. Then the wording of each claim is checked on its own by an independent model from a different AI provider than the one that wrote the answer — it never scores the answer as a whole, and it can't overrule the math. - **The receipt.** — Your user gets the answer at full speed. Moments later the Accuracy Report lands — and if a number was wrong, the owner gets the fix in Slack before it reaches a deck. _SPAN TREE / FINISHED TURN_ ``` turn → model.0 → tool.call_1 get_meta_insights → model.1 (final output) ``` _TOOL SPAN · JSON_ ```json { "span_id": "tool.call_1", "kind": "tool", "tool_name": "get_meta_insights", "payload": { "arguments": { "metric": "roas", "date_start": "2026-06-16", "date_end": "2026-06-22", "attribution": "7d_click_1d_view", "time_zone": "America/New_York", "currency": "USD" }, "result": { "roas": 3.46 } } } ``` | THE AI'S QUERY | Value | Check | | --- | --- | --- | | Metric | ROAS = conversion value ÷ spend | ✓ | | Currency | USD | ✓ | | Attribution | 7-day click, 1-day view (your definition) | ✓ | | Time zone | account time zone, America/New_York | ✓ | | Date range | Jun 16–22 · 1 week | ✗ Should be Apr 1 – Jun 30 (all of Q2) | The AI checked the wrong dates. _READ-ONLY QUERY · SQL_ ```sql -- Emet · read-only · on your SQL warehouse SELECT fiscal_quarter, SUM(spend_usd) AS spend, SUM(conv_value_7dc_1dv_usd) AS conversion_value, SUM(conv_value_7dc_1dv_usd) / SUM(spend_usd) AS roas FROM acme.paid_media.meta_ads_daily -- Meta data landed in your warehouse · account time zone WHERE date BETWEEN '2026-01-01' AND '2026-06-30' GROUP BY fiscal_quarter ORDER BY fiscal_quarter; ``` | fiscal_quarter | spend | conversion_value | roas | | --- | --- | --- | --- | | Q1 | 33,052.17 | 161,412.00 | 4.8836 | | Q2 | 37,084.09 | 90,110.63 | 2.4299 | _90,110.63 ÷ 37,084.09 = 2.43x_ _161,412 ÷ 33,052.17 = 4.8836 → 4.88x_ - **CORRECTED · "Q2 ROAS was 3.46x" → 2.43x.** — The AI queried one week, not the quarter. - **CORRECTED · "our strongest quarter yet" → Q1 ROAS was 4.88x.** — Q2 is down 50%. - **NOT SCORED · RECOMMENDATION · "Increase spend."** — Emet verifies what the data says, not what you should do. Campaign data cleared after 2 corrections · safe to report Q2 ROAS was 2.43x, down from 4.88x in Q1. [See a full Accuracy Report →](https://emet.so/report/evr-2026-q2-001) **The AI said strongest quarter yet. Your data said down 50%.** _WHAT EMET CATCHES_ ## The answer can look right. The source tells another story. | The error | Example | How Emet catches it | | --- | --- | --- | | Wrong date range | ROAS for one week reported as the quarter | Query audit | | Wrong attribution, time zone or currency | Last-click on a data-driven account; a date range off by one in the account's time zone; spend summed across currencies | Query audit | | Wrong number | "Campaigns were ~26.8% of email volume." Actual: 98.7%. | Re-derivation | | Bad math | A period-over-period change computed against the wrong prior period | Re-derivation | | Overgeneralized | "Email was over 96% of revenue in August and YTD." True for August only. | Claim check | | Invented cause | "The $13,867 gap is due to a 5-day attribution window." Nothing in the data says why. | Claim check | Examples from a one-week client case study at a multi-brand retailer. _Interpretive claims — opinions and recommendations — are deliberately not scored._ _EMET VS. LLM EVALUATION_ ## A different root, not a better judge. Eval tools check faithfulness to context. Emet checks truth against the source. | | LLM evaluation tools | Emet | | --- | --- | --- | | Checks the answer against | The context the model saw, or a test set you maintain | Your live source data, re-queried with its own read-only access | | How | An LLM judge, or a scorer you write yourself | Deterministic re-derivation built in, per data source; an independent model checks the wording | | Checks the query | Only if you write that check | Built in: dates, windows, time zones, currency, attribution | | Output | A score per response or per test | A verdict and a reason per claim, with the math | Keep your tracing. Langfuse, Arize Phoenix, LangSmith and Logfire keep working — add Emet as a second OpenTelemetry destination. [See integrations →](https://emet.so/integrations) _ONBOARDING_ ## Mapped to your definitions. - **Source-of-truth validation** — At onboarding we confirm every data source is connected and Emet can query each one, so it knows exactly what it's checking against. - **Your metrics** — ROAS, pipeline, margin — as your business defines them, not a generic default. - **Your calendar** — Fiscal quarters, attribution windows, time zones and currencies, mapped once. _DEFINITIONS_ ``` roas = conversion_value (7-day click, 1-day view) ÷ spend fiscal_q2 = Apr 1 – Jun 30 account_tz = America/New_York currency = USD · daily fx ``` _WHERE VERDICTS GO_ ## The right people get the receipt. - **Slack** — one alert per issue, not per report - **Emet dashboard** — every report, every team - **Export** — CSV [See the Slack integration →](https://emet.so/integrations/slack) _SECURITY_ - **Read-only, scoped by you, revocable at once** - **Never in the request path — exports run after your user has the answer** - **Stores the trace, the claims, the verdicts and the query results behind them — never a copy of your tables.** [Security →](https://emet.so/security) _FAQ_ ## Questions worth asking. **Q: Isn't this just AI checking AI?** Numbers are re-derived in deterministic code against your source. A model only checks the wording; it comes from a different provider than the one that wrote the answer, and it can't overrule the math. **Q: Won't better models fix this?** On frontier models in live client data, 1 in 5 claims carried an error. And a model can't tell you which answer is the wrong one — if it could, it wouldn't have handed it to you. **Q: We built our own checks. Why Emet?** An agent checking an agent can say an answer looks consistent. It can't say it's right without its own access to the source. Emet's root is your data, not a smarter judge. **Q: Does Emet slow our agent down?** No. The turn is exported after your user has the answer. A failed export is logged and dropped. **Q: What about the answer my user already saw?** It went out at full speed. The report and any correction follow moments later, in the dashboard and in Slack, routed to whoever owns that agent or data source. **Q: Who checks Emet?** Every verdict ships with its evidence and its equation, so anyone can check it by hand. **Q: Do we have to replace our tracing?** No. Point a second OTLP exporter at Emet. Langfuse, Arize Phoenix, LangSmith and Logfire keep working. **Q: What does Emet need access to?** Your agent's traces, and read-only access to the data sources you choose. Nothing else. --- ## Stop babysitting AI. Start trusting it. Emet delivers deterministic verification for enterprise AI data analysis — in seconds, with a receipt. **[Book a Demo →](https://emet.so/contact)** --- Emet · Truth infrastructure for AI data systems · https://emet.so