Book a demo

PRODUCT : semantic layer : CANARY (this fall)

Verify AI answeraccuracy.

how canary works

How does AtScale Canary verify AI answer accuracy?

AtScale Canary is an AI answer accuracy verification product launching in fall 2026. It checks the definitions, calculations and filters in an AI-generated answer against governed definitions in the AtScale semantic layer and flags anything that doesn’t match.


With Canary, you know the number is right before anyone acts on it.

THE PROBLEM

You can’t attest to accuracy you never measured.

Most teams don’t test AI accuracy at all, because there’s been no framework to do it.

Data teams test their pipelines, dashboards and code before they ship. AI answers usually go out after a spot check, since there hasn’t been a standard way to test them on a company’s own questions and data. Without a measurement, there’s nothing to attest to. A wrong answer that looks right gets used, and agents act on it without a person reading it first.

Canary is the first accuracy attestation framework we know of. It turns your questions and the SQL you know is right into a test suite, a score you can report and a dated record you can show an auditor.

how canary works

Canary scores every answer against SQL you know is right.

Canary runs each test prompt through a harness based on BIRD-Interact, compares the answer with the result of your gold SQL and records why each miss happened. A simulated user answers Ava’s clarifying questions, so Canary scores ambiguous prompts too.

Canary accuracy diagram — coming soon

Score the answers, and find out which ones failed

WHAT IT DOES

Define your questions

List the questions your business asks. Each test case pairs a question with test data, the gold SQL and the validated result. Every test case passes or fails, and the goal is 100%.List the questions your business asks. Each test case pairs a question with test data, the gold SQL and the validated result. Every test case passes or fails, and the goal is 100%.

Run

Canary runs every prompt and its gold SQL and reports a score, such as 184 of 200 correct.

Diagnose

Canary records why each answer missed, such as an ambiguous prompt or a missing metric definition, and what the model needs to fix it.

Track it over time

AtScale Semantic Modeling drafts the fix, Canary runs the suite again and a person approves the change. The test case stays in the suite, so you’ll catch it if the same miss comes back.

Diagnose

Run the suite on a schedule and before a launch or an audit. A falling score shows drift before it shows up in a board pack.

Ask it questions you already know the answers to

Build the suite.

A set of questions and their known-correct answers, drawn from your model and how people really query it.

Run and score.

Canary runs the suite against your stack and scores every answer against the gold standard.

Report.

You get the failures and the reasons, plus a dated record you can hand an auditor or a regulator.

Attest the metrics before agents act on them.

An agent doesn’t stop to sanity-check a number before it acts on it. Before an agent workflow goes live, turn the questions it’ll ask into Canary test cases. Promote only the model versions that pass, and run the suite again whenever the model or the agent’s LLM changes. Agents then call those attested metrics through the AtScale MCP server.

FAQs

How does Canary decide an answer is right?

It compares the AI’s answer with the validated result of your gold SQL. Canary runs each test prompt through a harness based on BIRD-Interact, and every test case passes or fails.

Can Canary score ambiguous prompts?

Yes. A simulated user answers Ava’s clarifying questions during the run, so Canary scores ambiguous prompts too.

What counts as a gold standard?

It’s known-correct SQL for a known question, written and reviewed by the people who own the metric.

Why not have another model check the answer?

A second model gives you another guess. Canary compares every answer with a result your team already knows is right.

Can we use our own benchmark?

Yes, and AtScale recommends starting there. A test case is a prompt, test data, gold SQL and a validated result. Public benchmarks such as BIRD-Interact can run alongside your own cases.

What happens when a test fails?

Canary records why the answer missed, such as an ambiguous prompt or a missing metric definition. AtScale Semantic Modeling drafts the fix, Canary reruns the suite and a person approves the change before it’s promoted.

Does it work on answers ACE didn't compute?

Yes. Canary can score any answer, including answers from AI running without AtScale.

What does an attestation look like?

It’s a dated record of each prompt, the answer that came back, the correct result and the score. You can hand it to an auditor.

Does Canary measure accuracy with and without AtScale?

Yes. Canary compares your results against a baseline run without AtScale, so you can show what the semantic layer adds on your own questions.

Does Canary run in our CI pipeline?

Yes, with help from AtScale services at launch. You can query test results, and our services team connects your acceptance criteria to CI.

Can we test our agents before they go live?

Yes. Turn the questions an agent will ask into Canary test cases and promote only the model versions that pass. Agents then call those attested metrics through the AtScale MCP server.

Hear about it first