# Hallucinations in data analysis aren't a text problem _September 18, 2026 · Paul, Head of Engineering_ > Why evaluating AI data analysis requires a fundamentally different approach than standard LLM observability. If you ask an LLM for the capital of France and it says London, that is a hallucination. If you ask an AI data agent for last quarter's revenue and it says $4.2M when the actuals are $3.8M, that is a calculation error. The industry uses the word "hallucination" for both, but treating a wrong number like a semantic error is why most enterprise AI deployments fail to cross the trust threshold. ## The wrong tool for the job Standard hallucination detection is built to catch errors in text. It checks whether an answer is grounded in the context the model was given — by semantic similarity, embedding distance, or another model acting as judge. But you cannot evaluate a math problem with semantic similarity. If the source data says `$1,000` and the AI says `$10,000`, the semantic distance is negligible, but the business impact is massive. [[Catch was="$10,000" now="$1,000"]] ## How to actually verify data To evaluate AI data analysis, you have to treat it like code execution, not text generation. 1. **You need the trace:** You have to know exactly what the AI did. What queries did it run? What endpoints did it hit? (This is why Emet reads OpenTelemetry). 2. **You need the source:** You cannot just look at the AI's output. You have to query the actual source system (Stripe, Salesforce, Snowflake) to get the ground truth. 3. **You need deterministic re-derivation:** You have to take the source data and recalculate the AI's claim using standard, deterministic code. If the AI claims ROAS is 3.46x, you don't ask another model if 3.46x sounds right. You calculate ROAS and compare the numbers. ## Trust requires proof An enterprise cannot run on vibes. If an AI system is making decisions about resource allocation, budget, or strategy, the numbers it produces must be correct. Stop treating them as a text problem. Treat them as bugs that need to be caught, flagged, and corrected before they land in a board deck. --- **[Book a Demo →](https://emet.so/contact)** Emet · Truth infrastructure for AI data systems · https://emet.so