Data governance for AI: preparing your organization to scale
Overview
This article is for anyone about to put AI to work on their customer experience data: heads of CX, journey management practice owners, and the data or IT counterparts they need alongside them. It sets out what has to be true about your data before AI can be trusted with it, which layers to get right in which order, and who owns each one. It is a readiness model rather than a click-by-click guide, so read it before you scale AI across your practice, not after.
The short version: models are easy to get hold of now, and the tools are mature. Readiness is the part nobody hands you.
AI does not create a governance problem, it exposes one
Most organizations have lived with fragmented ownership, inconsistent definitions, and knowledge that only exists in a few people's heads. It has rarely felt urgent, because humans quietly absorbed the cost. When two dashboards disagreed, someone knew which one to trust. When a term meant something different in one team than another, someone translated. When a number looked wrong, someone remembered why.
AI does none of that. It cannot tell that one source is stale, that "churn" means two different things in two departments, or that a survey was run on a segment that no longer exists. It reflects what you give it, at speed and at scale.
That is why we say insight quality is a data management problem, not an AI problem. If AI gives you answers you do not trust, the fix is almost never a better prompt. It is upstream, in what the AI can see, and in how well that material is structured, owned, and defined.
Note: The uncomfortable part of an AI pilot is often the most useful part. When the answers come back inconsistent, you have just been handed a precise map of where your governance is weak.
You probably do not have a data shortage
Nearly every organization we work with already has what it needs: research archives, survey exports, support logs, BI dashboards, product analytics, strategy decks, and a great deal of institutional knowledge. Very few have a shortage of material.
What they have is a context problem. The material exists, but nothing tells you what it means, where it belongs in the customer experience, whether it is still true, or who is accountable for keeping it that way. Volume is not the constraint. Meaning is.
This is the gap TheyDo is built to close. Evidence and numbers both attach to a specific place in the customer experience, the step where something actually happens, which gives two very different kinds of data a shared coordinate. A customer's words and a performance metric can then describe the same moment and be reasoned about together. That structure is what makes AI answers traceable instead of merely fluent. See How everything connects in TheyDo for how the model fits together.
The layers, in order
Readiness is cumulative. Each layer depends on the one beneath it, and skipping a layer does not save you time. It moves the failure later, where it costs more to fix.
1. Know what you actually have
Start with an honest inventory: which evidence exists, where it lives, how it flows in, and how old it is. In TheyDo this is the Data Hub, where interviews, surveys, support logs, and feedback data arrive either as manual uploads or through connected providers, each source carrying its type, its origin, and a processing status you can see. The point of the inventory is not completeness. It is knowing what you are working with.
2. Make it trustworthy
Trustworthy means you can answer three questions about any piece of evidence: where did this come from, who is accountable for it, and is it still current. Quotes point back to the source they were taken from, insights point back to the quotes beneath them, and the chain stays walkable in both directions.
It also means nothing disappears quietly. When a data source is deleted, its quotes are not silently lost: the row is prefixed [Deleted source], the source name is dimmed with a "This source has been deleted" tooltip, and you can list every orphaned quote using the None option on the Source filter of the Quotes tab. That is the behavior you want from a system you intend to trust, because it lets you find the evidence whose provenance has gone missing instead of quietly inheriting it.
3. Give it structure and meaning
Raw evidence is not AI-ready evidence. What makes it usable is structure: insights anchored to the step where the friction happens, opportunities framed as patterns that recur across journeys, metrics placed alongside the qualitative evidence for that same step, and a shared vocabulary layered over all of it through your taxonomy (statuses, types, groups, and tag groups).
Taxonomy is the least glamorous layer and the highest leverage one. The same idea tagged three different ways splits into three piles that never line up. AI tag matching can help you apply an agreed vocabulary consistently, but it cannot decide what your vocabulary should be, and that decision is the part that matters.
Note: Editing taxonomy needs taxonomy-edit permission
4. Assign ownership and decision rights
Every domain that matters needs a name against it. In practice that means filling in the owner property on the building blocks that carry weight rather than leaving it blank, and using workspace roles deliberately. Alongside the three system roles (Workspace Admin, Workspace Editor, Workspace Viewer), TheyDo ships four ready-made custom roles that map neatly onto governance responsibilities: Building Block Editor, Product Manager, Research Owner, and Taxonomy Owner. A named Taxonomy Owner is one of the cheapest governance decisions available to you.
Ownership also needs to be visible, not just recorded. Statuses show where work stands, and the journey activity log shows who changed what and when. See What are roles and permissions? and How to use the journey activity log.
Tip: Run an unowned-items pass before you scale AI usage. In any building block library, set the Owner filter to
Noneto list everything with no owner assigned, then work through it. Save it as a view and it becomes a standing check rather than a one-off cleanup.
5. Build AI literacy across the team
People need to understand how AI works and, more importantly, how it fails. Literacy here is role-specific rather than general. A researcher needs to know why a thin evidence base produces confident-sounding nonsense. A practice lead needs to know which questions the model is genuinely good at. An executive needs to know what an answer's provenance looks like, and how to ask for it. This is closer to infrastructure than to training, so budget for it accordingly. Getting better answers with the TheyDo Agent is a reasonable starting point for practitioners.
6. Set the guardrails
Guardrails are policies plus the controls that enforce them. The ones worth deciding on explicitly:
- Sensitive data. PII removal is an organization-level AI setting that strips personally identifiable information from new file uploads processed by AI features. Three things matter about it. It is not retroactive, so anything uploaded before you switched it on is unaffected. When it is active the toggle reads "Enabled since" with a date, which is how you tell what predates it. And it lives in the AI settings section of the organization details page, visible only to non-freemium organizations, to members holding the
Journey AIorganization permission. Decide on it before your first bulk upload, not after. See What is PII stripping in Journey AI? - Who can change what. Permissions are your real guardrail, and for AI writes the specific control is Edit mode: an organization admin enables it per workspace from the AI permissions table in organization settings, and it gates whether the TheyDo Agent can write to your building blocks at all. Within a conversation, actions that need confirmation surface as approval cards, and Doc edits arrive as a diff you approve or reject. Once a change has been applied there is no undo, which is exactly why who holds edit rights is a governance decision rather than a convenience setting. See Editing with the TheyDo Agent.
- Where AI is switched on. Both the organization-level and the workspace-level AI toggles have to be on before any AI write action can happen. That gives you a deliberate rollout: start in the workspace where the data is in good shape, and expand from there.
- An audit trail. The activity log gives you the record for a single journey, and the All activity tab in the Updates panel gives you the chronological stream across a workspace or the whole organization. Agree up front who reviews it and how often. Two things worth knowing: merge events are recorded but currently render as generic entries, and building blocks created and deleted within fifteen minutes by the same person are pruned from the feed. It is a strong trail, not a forensic one.
7. Then automate, augment, and transform
Only now does the interesting part pay off. With the layers beneath it in place, AI stops being a demo and starts compounding: mining insights from research at volume, revealing opportunities across journeys, summarizing for stakeholders, drafting Docs, and taking on the structural work that used to eat weeks. See What is the TheyDo Agent? and Mine insights with AI for what is available today.
Where to start
Momentum matters more than completeness here. The common failure is a governance program that spends two quarters producing a framework nobody uses. Pick something small, valuable, and visible instead.
- Choose one high-value journey where the outcome matters to someone senior and the evidence base is reasonably good.
- Inventory the evidence for it: what sources exist, how current they are, and what is missing.
- Name an owner for that journey and an owner for your taxonomy. Two names, written down.
- Agree the handful of definitions that keep causing arguments, and record them where people will actually see them.
- Tidy the taxonomy for that one journey: consistent statuses, consistent tags, unused tags retired, duplicate insights merged.
- Point AI at that journey and nothing else, then compare what it returns against what your team already knows to be true.
- Write down what broke. That list is your governance backlog, grounded in evidence rather than theory.
Then repeat on the next journey. Each pass gets faster, because the taxonomy and the definitions carry over.
What to align with your data and IT teams
Most of the above sits with the CX or journey management practice. A few things do not, and they tend to surface late and awkwardly. Worth raising early:
- Where the source of truth lives for each metric. If a number in TheyDo and a number in the BI tool disagree, someone needs to have already decided which one wins and why. This is a definitional question, not a technical one.
- How data arrives, and how often. Manual exports drift out of date quietly. Connected sources close that gap: insight mining can run on a daily, weekly, or monthly schedule for Qualtrics and Medallia sources, though manual uploads cannot be scheduled, and a schedule still needs someone to approve its runs. Decide who that someone is, because unapproved automations are removed after fifteen runs.
- Lineage expectations. Your data team may already have standards for documenting where a field comes from. Reuse them rather than inventing a parallel vocabulary for experience data.
- Sensitive data handling. Confirm that PII removal and your organization's own data-classification policy agree with each other, and that whoever holds the
Journey AIorganization permission knows they hold it. - Access and identity. Single sign-on, role assignment, and offboarding are governance controls, not just IT hygiene. A departed colleague who still owns forty insights is a governance gap.
- What AI is permitted to touch. Your organization may already have an AI policy. Map TheyDo's AI settings and Edit mode onto it explicitly, so the answer to "is this approved" exists on paper before someone asks.
One framing that tends to land well with data and IT stakeholders: this is not a request for a new data warehouse. It is a request to give experience data the same treatment structured data already gets, namely documented meaning, a named owner, and a governance process. Data management is how you make AI trustworthy, and that is usually a conversation your data team is pleased to be invited into.
A readiness checklist
Use this as a pulse check rather than a gate. Any question you answer "no" to is a starting point.
- Can you list every evidence source feeding your journeys, and say how current each one is?
- For any insight, can you get to the quote and the source behind it?
- Does every journey and every significant building block have a named owner?
- Is there one person accountable for taxonomy?
- Are the definitions that cause the most arguments written down anywhere?
- Do you know who holds the
Journey AIpermission and who has Edit mode available, and is that deliberate? - Is PII removal set the way your policy requires, and do you know what predates it?
- Could you explain to an executive where a given AI answer came from?