Pick a question. Each has a certified answer, and each is asked five different ways, the same five on both rows.
Scorecard
All 4 questions, five times per side, on gpt-oss-20b.
Ask your AI the same business question five times and count the answers. More than one number means it’s guessing. It can see every table you own, but nothing tells it what the columns mean, so each attempt fills the gap a little differently.
Each person words it differently. A small, inexpensive open-weight model answers all five against a fictional retailer’s tables. The second row gets the same model, the same tables, the same data and the same five wordings. The only difference is that those tables have been described: what each column means, which units it uses, which rows to leave out. Nothing was renamed or migrated.
These are the demo’s real definitions, word for word what the second row receives. Without them, the misses above have ordinary causes: cents read as dollars, returned items counted as sales, a retired copy of the table used because its name looked right.
Your warehouse and your Cognos and Power BI models already hold years of business logic, most of it correct and none of it readable by an AI. Drafting the definitions takes AI tools hours now. What takes judgment is deciding which definition is right when three systems disagree about net revenue, and naming an owner for each one, which is the part we do with your people.
Your most-used reports become the test questions, and their certified figures become the answer key. We ask each question five times on every model you care about, before and after, run the query behind each answer and compare the result to the report. Scoring is deterministic, so anyone on your team can audit it, and no AI grades another AI.
The test is allowed to report no improvement. One that can’t fail wouldn’t tell you anything.
The scorecard in the demo works this way on a fictional retailer. In the sprint it runs on your data and also tracks time and cost per answer, because described data lets a smaller, cheaper model do the daily work. The demo’s model is one of those.
Read-only access. We ask your AI your questions today and score the answers.
AI drafts definitions across your sources and flags every conflict.
Your owners settle the contested definitions. We publish them to your AI tools.
Answers live in the AI you already own, with a before-and-after scorecard.
Start with the domain whose questions you most need answered, often finance or sales. Data quality fixes and permission clean-up sit outside the sprint; we flag what we find and name who should own it. Afterward you can stop there and keep everything, or add the next domain at its own fixed price. You can also have us re-run the scorecard every quarter, so you know when a system change has started to shift the answers.
Tell us how your data and AI are set up today. A PMsquare consultant reads every report before it goes out, so expect yours within one business day. It covers where your AI is most likely to guess wrong, which domain to start with, and the first questions a sprint would test.