SolutionsApplied AI

Every tool can put AI on your data. None of them tell you which answers to trust.

The difference is a governed layer underneath: your definitions, your business rules, and a flag on every answer saying whether it is signed off like the dashboard or directional only. We build that layer in the stack you already own, and we keep it right as the business changes.

The problem, concretely

One question, four confident answers.

Ask a question every operator asks, 60-day repeat rate, and the stack returns numbers points apart, each delivered with total confidence. The bar operators set for this is the right one: an AI answer has to be as reliable as the dashboard, or it is not more efficient.

"Why is 60-day repeat down for the summer cohorts?"Same warehouse · four answers

What the tools say, no flag on any of it

Analytics platform's AI
24.1%
Excludes exchanges and gift orders entirely. A template definition, and nothing tells you so.
AI straight on the warehouse
31.7%
Counted any second order row as a repeat, subscription renewals and exchanges included.
The dashboard
27.4%
Right, and refreshed weekly, and silent on why it moved.

The governed layer, drilling to the why

60-day repeat is 27.4%, down ~2pts vs June last year
Definition used: second completed DTC order within 60 days, exchanges excluded, the one finance signed. Matches the dashboard because it is the same definition.
Governed · tested nightly
Follow-up: "driven by what?" Mix, not behavior. Core-line cohorts repeat at last June's rates; the dip comes from a higher share of new customers starting with a small accessory-only order, a segment that has always repeated lower. Cohort LTV tracks ~4% under plan on the same shift.
Governed · same definitions, cohort and LTV skills
And one more thing it noticed: the change log shows an accessories entry offer went live on paid landing pages on 6/14, which lines up with the mix shift. That campaign table is not governed yet. The likely join is the utm_content field, pulled and attached, but confirm with the growth team before this drives a decision.
Not governed yet · handed off with its homework done
Nothing here is a black box. Every answer shows the definition it used and the query it ran, so a skeptic can take it apart.

Illustrative example.

The part the AI demos skip

Every way of doing this has a failure mode, and the demo hides all of them.

AI straight on the warehouse
A frontier model connected directly to a clean warehouse answered 21% of business questions correctly in Anthropic's testing. Not because the model is weak, but because an ambiguous question has several defensible readings, and the model picks one and commits. An analyst would ask what you meant.
Analytics platforms' AI
The better tools give right-ish answers, but nothing tells you which answers are governed and which are guesses. We have watched one produce two different LTVs for the same brand, with no way to tell which to sign off. And the definitions are rented: built for the median brand, gone when you leave.
Building it in-house
The strongest version we have seen: a sharp analyst builds a working AI data skill in a season. It is right until a definition moves, and it leaves when they do. Even teams with people on AI full time rarely end up with a governed layer, because the layer is maintenance, not a build.
A governed layer, honestly
It only covers what has been encoded, and outside that the honest behavior is a flag, not a guess. And accuracy is maintained, not installed: an unmaintained system decays from accurate to confidently wrong without failing loudly.

The real risk is not a wrong number, it is unexamined confidence. Most people do not question an answer that looks credible, and an answer you cannot take apart is not one you can sign off.

The playbook

The layer gets built once, in your stack, and everything else runs on it.

You do not need all of it, and almost nobody starts with more than two. Everything lives in your repo and your warehouse, so it is yours whether or not we are around.

01
Definition audit & metric contract
Where the numbers disagree today, why, and which version finance will stand behind. Nothing answers from a definition nobody signed.
02
Governed semantic layer, in your repo
Your definitions in code, versioned and tested: multi-brand rollups, subscription logic, the custom parts a template cannot hold.
03
Business rules & guardrails
The context that keeps a confident wrong read from reaching a decision: which orders count, which comparisons mislead, which cuts are too thin to trust. The person asking does not need to know any of that, because the layer does.
04
AI skills for your actual questions
Topline by business line, cohorts, marketing: guided asks that run the checks an analyst would run, built as skills your team can read and edit.
05
The governed flag on every answer
Signed off like the dashboard, or directional and marked. The flag is what makes self-serve safe for people who are not data people.
06
Accuracy testing & maintenance
Known-answer questions run on a schedule, and definitions ship in the same pull request as model changes, so drift is caught before leadership is.
07
Analyst work on top of the layer
Creative fatigue reads, new-product incrementality, experiment design and readouts: automated once the numbers underneath hold, with the reasoning shown.
08
Rollout with a concierge, not a login
We sit with the first users, watch what they actually ask, and encode what is missing. Self-serve nobody trusts is a login nobody uses.
Why maintained beats installed

Accuracy is maintained, not installed.

The 95% is what that study measured across every question thrown at the system. It is not the ceiling: inside the governed boundary an answer matches the dashboard by construction, because it is the same tested definition. The flag exists so you always know which side of the boundary an answer came from.

The number that should worry you is the second one. Definitions drift as the business changes: new lines, new channels, a metric quietly redefined in one team's sheet. An unmaintained system does not fail loudly; it degrades from accurate to confidently wrong.

Measured accuracy decays from 95% to 65% within a month when definition maintenance stops.

21%95%+
answer accuracy, a frontier model on the raw warehouse versus the same model over a governed definition layer
Anthropic Data Science & Engineering, internal results, June 2026. Consistent with our client implementations.
After the foundation

Once the answers hold, the layer starts doing analyst work.

None of this is AI improvising an analysis. Each one starts by agreeing how the read should be done, the way your best analyst would do it: what counts, what gets excluded, which comparisons mislead. That method gets encoded once, and then it runs the agreed way every time, across more cuts than a person has time to check. Without it, self-serve means everyone doing their own version at whatever quality they bring.

Creative
Creative fatigue reads
The team agrees once what fatigue means and which cuts are too thin to trust. Then every angle gets read against it continuously, and briefs change before budget burns instead of when someone has a spare afternoon.
Launches
New-product incrementality
The rules for separating added demand from demand moved off existing lines get set before the launch, not argued after it. The next launch decision runs on lift, read the same way every time.
Experiments
Experiment design & readouts
Sample sizes and test lengths set before the test, results read against them after, so an inconclusive result gets called inconclusive instead of becoming a win because someone wanted one.

All of it cites the definitions and queries it used. A read that cannot be taken apart does not ship.

Who this is usually for

A data function you could not otherwise justify.

The brands this fits best are usually somewhere between $5M and $50M: big enough that a wrong number costs real money, not big enough to carry a data team. Operators at that size keep describing the same dead end to us.

Hire in-house

~$200k for one person who needs 4 skill sets, takes months to find, and walks out with everything they learned about your business.

Buy the platform

The pitch lands at a similar number, the definitions are the template's rather than yours, nothing flags what is governed, and leaving means starting over.

Build the layer

From a few thousand a month, a senior team builds the governed layer in the stack you already own. Everything stays yours, and the answers hold whether or not we are around.

Where it runs

Live or going in now.

Not every implementation is a written study yet, so here is the honest list: what was there, and what runs now.

ClientWhat was thereRunning now
Noovo
Luxury camper vans · DTC
What was thereData Studio reports nobody opened, over a funnel nobody could trace a prospect through.
Running nowA rebuilt prospect-to-customer funnel, with governed chat and reporting on top of it.
amp
Fitness SaaS
What was thereThe old self-serve fallacy: every team maintaining its own drifting dashboard copy, each with its own answer.
Running nowOmni + Claude. Users publish their own dashboard variations, all backed by the same code-defined definitions. One set of numbers.
Home & furnishings brand
DTC · client unnamed
What was thereRoutine questions queued behind a data team.
Running nowThe layer behind the 2.8x growth study answers the day-to-day. A data team of 2 covers what would normally take 5.
Arena Club
Sports collectibles marketplace · in progress
What was thereHundreds of Metabase reports straight on the production database, each with its own SQL, plus ungoverned Claude on top.
Running nowFoundations rebuilt so answers come from definitions instead of improvisation, and the whole thing can scale.
Marketplace advertising business
Amazon ecosystem · client unnamed
What was thereA Tableau report queue between a long-established operation and its numbers.
Running nowGoverned chat and dashboards, so decisions stop waiting on report requests.
The measurement case studies

Get in touch.

Tell us what you are trying to figure out. If we can help, we will say how. If we are not the right fit, we will say that too, and point you somewhere better.

01
Free work first
We start by doing a real piece of work on your data, free, so the first thing you judge is output, not a pitch.
02
First build, fixed scope
A trusted core metric layer in weeks, not quarters, at a fixed price. Small enough to be safe, real enough to matter.
03
Scale when it earns it
Month to month from there, from a few thousand a month. You own every model, dashboard and definition.