Build the suite.
A set of questions and their known-correct answers, drawn from your model and how people really query it.
how canary works
AtScale Canary is an AI answer accuracy verification product launching in fall 2026. It checks the definitions, calculations and filters in an AI-generated answer against governed definitions in the AtScale semantic layer and flags anything that doesn’t match.
With Canary, you know the number is right before anyone acts on it.
THE PROBLEM
Most teams don’t test AI accuracy at all, because there’s been no framework to do it.
Data teams test their pipelines, dashboards and code before they ship. AI answers usually go out after a spot check, since there hasn’t been a standard way to test them on a company’s own questions and data. Without a measurement, there’s nothing to attest to. A wrong answer that looks right gets used, and agents act on it without a person reading it first.
Canary is the first accuracy attestation framework we know of. It turns your questions and the SQL you know is right into a test suite, a score you can report and a dated record you can show an auditor.
how canary works
Canary runs each test prompt through a harness based on BIRD-Interact, compares the answer with the result of your gold SQL and records why each miss happened. A simulated user answers Ava’s clarifying questions, so Canary scores ambiguous prompts too.
Canary accuracy diagram — coming soon
WHAT IT DOES
List the questions your business asks. Each test case pairs a question with test data, the gold SQL and the validated result. Every test case passes or fails, and the goal is 100%.List the questions your business asks. Each test case pairs a question with test data, the gold SQL and the validated result. Every test case passes or fails, and the goal is 100%.
Canary runs every prompt and its gold SQL and reports a score, such as 184 of 200 correct.
Canary records why each answer missed, such as an ambiguous prompt or a missing metric definition, and what the model needs to fix it.
AtScale Semantic Modeling drafts the fix, Canary runs the suite again and a person approves the change. The test case stays in the suite, so you’ll catch it if the same miss comes back.
Run the suite on a schedule and before a launch or an audit. A falling score shows drift before it shows up in a board pack.
A set of questions and their known-correct answers, drawn from your model and how people really query it.
Canary runs the suite against your stack and scores every answer against the gold standard.
You get the failures and the reasons, plus a dated record you can hand an auditor or a regulator.
An agent doesn’t stop to sanity-check a number before it acts on it. Before an agent workflow goes live, turn the questions it’ll ask into Canary test cases. Promote only the model versions that pass, and run the suite again whenever the model or the agent’s LLM changes. Agents then call those attested metrics through the AtScale MCP server.
It compares the AI’s answer with the validated result of your gold SQL. Canary runs each test prompt through a harness based on BIRD-Interact, and every test case passes or fails.
Yes. A simulated user answers Ava’s clarifying questions during the run, so Canary scores ambiguous prompts too.
It’s known-correct SQL for a known question, written and reviewed by the people who own the metric.
A second model gives you another guess. Canary compares every answer with a result your team already knows is right.
Yes, and AtScale recommends starting there. A test case is a prompt, test data, gold SQL and a validated result. Public benchmarks such as BIRD-Interact can run alongside your own cases.
Canary records why the answer missed, such as an ambiguous prompt or a missing metric definition. AtScale Semantic Modeling drafts the fix, Canary reruns the suite and a person approves the change before it’s promoted.
Yes. Canary can score any answer, including answers from AI running without AtScale.
It’s a dated record of each prompt, the answer that came back, the correct result and the score. You can hand it to an auditor.
Yes. Canary compares your results against a baseline run without AtScale, so you can show what the semantic layer adds on your own questions.
Yes, with help from AtScale services at launch. You can query test results, and our services team connects your acceptance criteria to CI.
Yes. Turn the questions an agent will ask into Canary test cases and promote only the model versions that pass. Agents then call those attested metrics through the AtScale MCP server.