News & Views
Confidently Wrong – Why You Need Context Governance
Writing down why your numbers move fixes most wrong answers, for people and for AI agents. Then the business changes, or the code does, and the notes stay where they were. This is what governance looks like at a 40-person company, and why a good domain model leaves less to govern.
Table of content
We’re just back from Compass in Budapest where I made the case that to be great at agentic analytics, we need to hold to three principles:
-
- An agent needs curated context to avoid confidently wrong answers.
-
- Context evolves, so governance is key. Someone has to own that context and keep it true as the business changes.
-
- We need to model the business before you model the data. A warehouse that describes the business well needs far fewer notes in the first place.
The first point is the subject of our earlier article, How to Build a Context Layer (Because You Cannot Buy One): what to write down, where to keep it, and why an agent can’t retrieve what nobody recorded.
In this article, we’ll pick up the two remaining points: once the context exists, how do you keep it right and relevant? And how do you need less of it?
As before, we are using the example of a fictional company (but fully equivalent to the kinds of setups we see at Tasman). We’ll do the theory in plain text, and have worked examples for Kelder Koffie in framed boxes like the one below.
KELDER KOFFIE
Kelder Koffie is a fictional coffee subscription business on an Amsterdam canal: about 40 people, one analyst called Sanne, four data sources (Shopify, Recharge, Klaviyo and the ad platforms) and a tidy dbt project. We built it so that we know the right answer to every question.
In March 2026 Kelder moved its billing from Shopify to Recharge. The import had no “paused” state, so 1,903 paused subscriptions arrived as cancelled. So if you do a naive count in the tables, then raw March churn reads 9.4%; but the real figure is actually 3.3%. With the analyst’s notes, an AI agent got that right every time. Without them, it never did as it lacked that context.
The repo is here in our GitHub, and every screenshot and number in this piece comes from it.
Context evolves, so you need to govern it.
When you write context down, you tie it to the business at that particular point in time. A note on a column says what a field really means, but only until the column changes again. A decision record says why a number is calculated the way it is, but only until finance needs a different FX calculation. A pinned query says what the right answer was, but only for a particular report in a particular moment.
All of those keep moving. The business changes its rules, switches systems, runs a campaign, hires a new CMO. The code can change underneath for good reasons — eg. a refactor or a clean-up, usually by someone who wasn’t really involved with the original analytics decisionmaking. When the code or definitions move, but the note doesn’t, then that note turns into a wrong instruction. A person might notice and correct it, but an agent follows it very literally.
I call this semantic drift, or semantic rot: the pipeline keeps running while the meaning of the numbers it produces is slowly drifting away from the business. All tests pass, nothing fails, reports keep running — but the meaning changes nonetheless.
KELDER KOFFIE
Kelder changed its churn rule twice in eight weeks: once on purpose, once by accident.
The deliberate change shows what good looks like. From 1 May a failed card payment only counts as churn if it isn’t recovered within 30 days of retries (that retry period is called dunning). Sanne recorded it in one go: a row in the metric changelog, a business events row, and decision record 0009, agreed with Femke in finance. Its rejected alternative is worth stealing: “Silently restating all history under v2, because board decks had already quoted v1 figures.” The board saw v1, so v1 stays visible.
The accidental change is the rest of this section.
Drifting Commits
Drift often starts with an ordinary code change that’s more housekeeping then anything else – well intended. Typically also reviewed by someone who can see the SQL, knows its important, but does not remember the reasons behind it. The best way to show this is in the example of Kelder.
KELDER KOFFIE
On 26 June someone tidies the churn model. It has two near-identical blocks of SQL, one per churn definition, which look like duplication. The commit, “Fix churn logic”, deletes a six-line block and changes one join.

kelder-dbt/models/marts/metrics_subscriber_churn_monthly.sqlThe pull request description is one line, and none of the four impact boxes is ticked. It is reviewed and merged. Every dbt test passes.

Kelder keeps a restated churn series (past months recalculated under today’s definition, so this year and last year compare fairly). After the commit, that series counts the paused subscriptions as cancellations again, and the year-on-year comparison flips.

Churn really fell, from 2.85% to 2.45%. After the commit it appears to rise to 3.50%, the opposite of what happened.
Every note still described the business correctly and was now wrong about the code. The caveat on churned_at still said the import rows were excluded. Decision 0007 still explained why the unadjusted series exists: finance reconciles against Recharge exports.

kelder-dbt/models/marts/_marts__models.yml

kelder-dbt/context/decisions/0007-recharge-migration-churn-artifacts.md
Nobody did anything wrong on purpose. The reason was written down, just not anywhere the person making the change would look.
The tests don’t flag this.
Three things let drift through and tests are a critical failure point.
First, data tests check superficial shape of the numbers. A test that a rate sits between 0 and 1, or that each key is unique, passes happily on a good-looking number that has bad reasoning behind it. Second, safeguards typically concentrate on the headline numbers, like last month’s revenue. The figure everyone quotes gets a reconciliation (as it would make a board deck look bad otherwise); but the underlying data doesn’t — and that is where drift hides. Third, most readers of a number never interact with the code or logic that produces it. A dashboard, a saved query or a board pack shows whatever the column says.
KELDER KOFFIE
One of Kelder’s tests asserts that every churn rate is a share between 0 and 1. The broken restated figure for March is 8.87%, a perfectly valid share. March’s headline number had a pinned answer, and it survived the commit because it reads a different column. Then look at who reads the restated series:
| Reader | Reads the SQL? | What it said after the commit |
|---|---|---|
| The agent with notes | Yes | Improved, 2.45% against 2.85%, and flagged the bug (3 of 3 runs) |
| The dashboard | No | Worse |
| A saved query pinned in Slack | No | Worse |
| The board pack | No | Worse |
The agent read the caveat, read line 107, saw they disagreed and recomputed: “Don’t use the verified query for this one. The cause is a bug in the governed model.”
An agent that catches this is impressive, but it checks once per question, for one reader. Governance can’t depend on which reader happens to look.
What governance means at 40 people
“Governance” usually means a council, a policy page and a data catalogue (a searchable inventory of tables and owners). Those record decisions; none can stop a commit. Michael Madson put it well this summer: committees, stewards and catalogues “can create the appearance of governance without the capacity to govern”. Gartner predicted in 2024 that 80% of data and analytics governance programmes would fail by 2027 for lack of a real or manufactured crisis.
My definition is smaller and stricter: someone owns each note, and a change in the business, or the code, forces the notes to change too. A test for your own team: if a change to a number can merge without anyone being asked whether that number moves, the number isn’t governed, however good the documentation is.
In practice that comes down to four properties:
- Owned. A named person, not a team. If your owner field says “Data”, nobody owns it. (dbt lets you put an owner on a group of models, but treats it as a label and enforces nothing. Pair it with a
CODEOWNERSfile and branch protection, so GitHub requires that person’s review.) - Versioned. Each definition has a version and a start date, and each series has one stated use.
- Change-controlled. A change to what a number means cannot merge until someone has said whether it moves, and the known answers still hold.
- Co-authored. The people who use a number agree its definition with the people who build it.
KELDER KOFFIE
Every business event row names a person: Sanne for the migration, Joris for the pause campaign, Lotte for the courier strike. Churn has three series with one use each: as reported (what the board saw), restated (for comparing years) and raw (for reconciling with Recharge only).
Three checks that catch the drift
Change control is the key bit here – just like with any data generating system. I also think it is actually the cheapest to add, particularly if you compare against what not fixing it would cost. Three small checks run in CI (the automated checks that run whenever someone opens a pull request). Each one turns a sentence in a note into something that can stop a merge.
- Pin the answer. A verified query is a question people ask often, the SQL that answers it correctly, and the correct answer, stored where an editor or an agent can’t quietly change it. CI reruns every query on every pull request and compares the result with the stored answer.
- Ask one question in every pull request. Does this change move a metric? The template offers four boxes: no metrics move; metrics move and the changelog is updated; metrics move and a business event is added; a decision record is added or superseded.
- Test the answer instead of trusting it. If the change touches a metric model, exactly one impact box must be ticked. “No metrics move” is checked against the verified queries, so you can’t tick it and be wrong. “Metrics move” requires an edit to the changelog or events file in the same pull request. The decision box requires a link to the record.
KELDER KOFFIE
All three checks live in the repo, in under 300 lines of Python. The year-on-year question is one of seven verified queries:

kelder-dbt/context/verified_queries.ymlHere is what the June commit gets:

When a number should move, move it on purpose, in one pull request: add or supersede the decision record (never edit an old one), add the changelog row, update the stored answer. And never update a stored answer just to make a failing check pass. The governance meeting now happens inside every pull request that touches a metric. It takes minutes and leaves a record.
Model the business before you model the data
Checks catch drift after someone writes it. A better model stops some of it from being possible at all. Most of the notes a data team writes are really patches over a gap in how the data describes the business.
I learned this the expensive way. At Peak, the brain training app, we sent auditors different lifetime value numbers during due diligence. I had built subscription numbers straight off a raw billing table that logged every payment retry as a transaction, so failed retries showed up as extra renewals. Better notes would not have saved me. A better model would.
Why one join could break churn
Most warehouses are organised around their source systems: a Shopify table and a Recharge table, joined and summed into marts. That works until a source’s idea of a record differs from the business’s. The warehouse accepts whatever the source calls a cancellation, and a note has to correct it. Every correction is one more thing a refactor can undo.
KELDER KOFFIE
Kelder’s marts follow its source systems. The raw import rows sat one join away from every churn number, held back by a flag and a filter someone had to remember.

churned_at, status, is_legacy_pause_restore) exist only because the model took the source system’s word for it. kelder-dbt/docs/erd_marts.svgThe domain model
A domain model describes what the business is made of before anyone writes SQL against a source: the objects, what counts as each one, and how they relate. It comes with an ontology, the agreed answer to “what exists here, and what counts as each thing?”. This is an old idea. Bill Inmon argued for a business-first model of the warehouse decades ago. What is new is that an agent now reads your model literally, so the cost of a sloppy one shows up straight away.
The payoff is that business rules become structural. Instead of a note saying “these rows aren’t really cancellations”, the model can’t call them cancellations. Sources map into the domain layer once, so a source changing underneath only changes the mapping.
KELDER KOFFIE
Kelder has five objects: customer, subscription, status change, delivery and payment attempt. Its ontology fits in two sentences. A pause is a state, never a cancellation. A gift is a type of subscription with no recurring revenue.

kelder-dbt/docs/domain_model_logical.svgOne rule does most of the work: a cancellation needs an initiator (the customer, the dunning process or operations) and a reason. The 1,903 import rows have neither, so the domain layer refuses to call them cancellations and holds them for review.

held_for_review. kelder-dbt/docs/domain_model_erd.svgThe June commit worked by pointing a reporting model at the unadjusted events. In this design there are no unadjusted cancellations to point at. And when billing moved from Shopify to Recharge, only the mapping into the domain layer would have changed; every consumer, agent included, keeps reading “subscription” as before.
This is design reasoning, not a trial result. The domain model is a proposal in the repo’s docs folder; Kelder’s dbt project does not implement it, and we did not run agents against it.
Four rules, and how to make them stick
- Model the business before you touch the data. Start from subscribers and pauses, not from a Shopify table and a Recharge table. One object may draw on several sources; that’s fine.
- Each thing has exactly one definition. This means talking to people. Marketing, when you say lead, what do you mean? Why is that different from sales?
- Downstream tools read only from the domain layer. Dashboards and agents never query raw billing data. In dbt you can enforce this with model access: mark source-shaped models
privateorprotected, and a model in another group that references them fails when the project is parsed, before anything runs. - Business logic lives in one place. A rule like a 30-day dunning grace is written once, in the domain layer. A model contract (dbt’s
contract: {enforced: true}, which checks column names and types before the model builds) keeps that layer’s shape fixed while sources move underneath.
What is left to govern
With a domain model, most of what used to be context notes becomes structure. What remains are the reasons behind decisions and the outside events no table can hold.
KELDER KOFFIE
Six things happened at Kelder in the first half of 2026. Only one is a true context note.
| What happened | Where it belongs | Left to govern |
|---|---|---|
| Billing migration, 12 March | Domain model, plus decision 0007 | One record: how the 142 genuine cancellations during the freeze were confirmed |
| Pause-instead-of-cancel campaign, 28 March | Domain model | Nothing: a pause is a status |
| Klaviyo connector outage, 9 April | Data foundations | A completeness check on the load |
| Churn v2, 1 May | Versioned metric definition | Changelog row, decision 0009, stored answers |
| PostNL courier strike, 19 May | Context note | One event row, owner Lotte |
| Father’s Day gift bundle, 14 June | Domain model | Nothing: a gift is its own subscription type |
The context layer shrinks to something one named person can keep current. Governance gets cheaper as the model gets better.
What to do on Monday
There is no tool to buy. For the data team, in this order:
- Name a person for each number on the board slide. A person, not a team. Put the names in your model groups and in
CODEOWNERS. - Pin the answers. Store the correct answer for those numbers, plus one comparison people make (year on year), as verified queries. Rerun them in CI.
- Add the impact question to your pull request template, and check it, so “nothing moves” is tested rather than taken on trust.
- Ship the next definition change on purpose: a new version with a start date and a decision record, in one pull request.
- Write down what a customer, a subscription and a cancellation are before your next model touches a new source. Start with one team’s questions, not the whole warehouse.
Skip the council, the catalogue and the policy document for now. Come back to them when something fails for the wrong reason.
Your notes were right on the day you wrote them. Governance keeps them right the day after.
v1.2, last updated 2026-10-08