Book a call

PRODUCT : CANARY

Is your AI accurate? Canary knows.

Attests the answer is right (Coming this fall)

AtScale Canary

AtScale is the only vendor that ships the BIRD accuracy benchmark. Run it and you see what AtScale does to accuracy, where the gaps are, and whether you closed them. We put our code where our mouth is.

BIRD proves it in public. Canary runs the same test on your data, your metrics and the questions your business actually asks, and names the answers that came back wrong before an executive acts on one.

THE PROBLEM

Accuracy is the one claim in this market with no scoreboard

The claim is everywhere and the evidence isn’t. Nobody tests it in your environment, against your metrics, on the questions your business actually asks.

An answer that looks right and is wrong is worse than no answer, because it gets used. And nothing in the stack raises its hand when it happens.

A file of questions, a score, and a list of what broke

The input is a file of real questions, each paired with the SQL that produces the right answer. Canary asks every one of them through your semantic layer and compares what comes back. You get a score you can report and a list of failures you can act on.

Canary accuracy diagram — coming soon

Scores the answers, and shows you which ones failed

WHAT IT DOES

Tests against a gold standard

Answers are scored against known-correct SQL, not against another model’s opinion of them.

Shows the failures

Not just a score. Which questions failed, what the right answer was, and whether the model or the definition broke.

Tracks it over time

Accuracy that falls is drift, and it shows up in a score before it shows up in a board pack.

Ask it questions you already know the answers to

Build the suite.

A set of questions and their known-correct answers, drawn from your model and how people really query it.

Run and score.

Canary runs the suite against your stack and scores every answer against the gold standard.

Report.

You get the failures and the reasons, plus a dated record you can hand an auditor or a regulator.

ACE computes and Canary proves. Attestation depends on having computed the number, so a catalog or a context layer can’t offer it. Ava makes sure the right metric was chosen; Canary is the evidence that it was.

What counts as correct, and who decides

The questions engineers ask first

What counts as a gold standard?

Known-correct SQL for a known question, written and reviewed by the people who own the metric.

Does it work on answers ACE didn't compute?

It can score any answer, and that comparison is the interesting one.

How often does it run?

On a schedule, and on demand before a launch or an audit.

What does an attestation look like?

A dated record of what was asked, what came back, what was correct, and the score.

FAQs

What is the BIRD benchmark, and why does AtScale publish it?

BIRD is a public, open benchmark for text-to-SQL accuracy, and AtScale is the only vendor that ships it rather than just citing it. Run BIRD yourself and you see what AtScale does to accuracy, where the gaps are, and whether you’ve closed them. Canary (coming this fall) takes that same approach and points it at your own data, your own metrics, and the questions your business actually asks, not a public leaderboard.

How is Canary different from trusting an AI's own confidence or having another model check the work?

Canary doesn’t grade an answer against another model’s opinion of it. It scores every answer against a gold standard, known-correct SQL for a known question, written and reviewed by the people who own the metric. That’s the difference between “this looks right” and “this is right.”

What happens when Canary finds a wrong answer, do I just get a score?

No, you get the score and the diagnosis. Canary shows which questions failed, what the right answer should have been, and whether the model or the definition broke, so you know what to fix instead of just that something’s wrong.

Can Canary catch accuracy that degrades over time, not just a one-time test?

Yes. Canary tracks accuracy over time, and a falling score is drift showing up in the data before it shows up in a board pack. Running the suite on a schedule, not just once before launch, is what turns Canary from a one-time check into an early warning.

How does Canary relate to ACE and Ava?

ACE computes and Canary proves; attestation depends on having already computed the number, so a catalog or a context layer can’t offer it. Ava makes sure the right metric was chosen, and Canary is the evidence that it was. ACE and Ava are available today; Canary is coming this fall.

Hear about it first