← Back to list

You Can’t Govern AI Without Governing Data First

Most organisations think they have an AI problem. They are worried about bias, hallucinations, explainability, and whether their models are…

Dr Victoria Holt · 2026-05-25 14:21 · 0 claps · 4.4 min read
#ai-governance #data-governance #data-governance-tools #microsoft-purview
Open on Medium ↗
Wiki topics: SAF · Safety & Alignment

You Can’t Govern AI Without Governing Data First

Most organisations think they have an AI problem. They are worried about bias, hallucinations, explainability, and whether their models are behaving the way they expect. To fix this, they build AI governance frameworks, create policies, form steering committees, and define principles for responsible AI. Yet very quickly, something doesn’t quite work. AI programmes stall, confidence drops, and risk teams start asking questions that can’t be answered while business teams push forward anyway because the value is too obvious to ignore. Underneath all of it sits a problem that is rarely acknowledged properly: most organisations are trying to govern AI without first governing the data that AI depends on.

Governing AI without governing the data

Governing AI without governing the data

The idea that AI governance is a separate discipline has become widely accepted. It is spoken about as something new, something layered on top of existing capabilities, and something focused on models rather than data. In practice, that separation does not really exist. Recent research from KPMG highlights this exact disconnect, revealing that 62% of organisations cite a lack of data governance as the main barrier to AI success. This statistic underscores a fundamental reality that the industry is finally waking up to, which is that AI governance is not parallel to data governance but is built directly on top of it.

Every AI system is entirely dependent on the data it is given. This extends beyond training data to include operational data, contextual data, reference data, and increasingly, real-time data flowing through complex environments. The model may generate the outcome, but the data shapes the decision. Analysts at Gartner and McKinsey have long pointed out that data quality and robust data architecture are the true differentiators in enterprise technology deployments, and this is magnified tenfold with artificial intelligence. If the data is poorly governed, inconsistent, incomplete, or misunderstood, the AI system has nothing stable to operate on. It doesn’t matter how well the model has been designed or how carefully the ethics have been considered because the outcome is already compromised.

What is striking is how often governance conversations stay focused strictly at the model layer. There is a lot of energy spent on ensuring models are fair, transparent, and explainable. These are important concerns, but they are not where most problems originate. The more fundamental issues sit much earlier in data that cannot be trusted, in ownership that is unclear, in access that is not fully understood, and in lineage that cannot be traced end‑to‑end. When AI is introduced into that environment, it does not fix those issues but instead accelerates them. A dataset with quality issues does not quietly produce unreliable reports anymore; it starts driving decisions at scale. Data that was previously accessed by a small group of analysts is suddenly embedded in workflows, models, and automated processes that reach far beyond its original scope. AI exposes how well data is governed, and in many cases, the answer is not very well at all.

For a long time, data governance has been positioned as a compliance activity. It has been about documenting data, defining policies, assigning ownership on paper, and ensuring that regulatory expectations can be met when required. In some organisations it has been effective, but in many it has remained disconnected from how data is actually used. AI changes that dynamic completely. Governance can no longer sit to the side as a supporting function. It has to move into the flow of decision-making itself, shifting from something that describes data to something that actively controls how data is used. This is where the real transition happens, moving from static governance to operational governance, or from knowing to controlling.

That shift is not easy because AI systems do not operate in neat, isolated domains. They connect across applications, data platforms, and business processes, combining datasets in ways that were never originally intended. Trying to govern that with manual processes and centralised review boards is not realistic because the scale and speed simply do not allow it. This is why so many organisations find themselves in the same position where they understand the need for governance and have frameworks in place, but when pushed, they cannot answer basic questions regarding what data is actually being used by their AI systems, who owns it, how it is being accessed, and what decisions it is influencing.

These are fundamentally data governance questions, which is why the concept of non-invasive data governance is becoming critical. Heavy, process-driven governance models that rely on gatekeeping and manual enforcement struggle in modern environments because they introduce friction and slow down access. What is needed instead is governance that works with the data landscape rather than against it, driven through metadata rather than the physical movement of data. This approach enables visibility without restricting use unnecessarily, applying controls dynamically as data flows between systems rather than relying on static rules defined at a single point in time.

Even that is not enough without automation. AI operates at a scale and speed that manual governance cannot match, with decisions being made continuously and new use cases emerging faster than human processes can be updated. If governance is going to keep up, it has to become part of the system itself through automated classification, continuous lineage tracking, and real-time monitoring of access. Policy enforcement must become something that is applied automatically and consistently rather than relying on human intervention, marking the difference between governance that reacts and governance that operates.

Technology plays a vital role here, and platforms such as Microsoft Purview are becoming important because they enable this shift from theory to operation. By bringing together discovery, classification, lineage, and access governance into a connected layer that sits across the data estate, such tools allow organisations to implement governance in a way that scales. Microsoft leaders frequently emphasise that secure, well-governed data is the ultimate prerequisite for safe AI deployment, and integrating these capabilities helps bridge the gap between data management and model oversight.

The presence of a platform does not solve the problem on its own if governance is still treated as mere documentation. The value comes when the approach changes first, and the technology is used to enable it. As organisations move further into AI, particularly toward more autonomous and agentic systems, this dependency only becomes more visible. The conversation is shifting away from whether a model is accurate and toward whether the organisation can trust what is happening when that model is used in practice. That trust is built at the data layer, and the organisations that succeed with AI will be the ones that recognise data governance and AI governance are not separate initiatives, but two inseparable parts of the exact same system.


메타데이터
post_id
732ac754d21e
slug
you-cant-govern-ai-without-governing-data-first-732ac754d21e
url
https://medium.com/@holt.victoria/you-cant-govern-ai-without-governing-data-first-732ac754d21e
canonical_url
https://medium.com/@holt.victoria/you-cant-govern-ai-without-governing-data-first-732ac754d21e
author_url
https://medium.com/@holt.victoria
status
ok
fetched_at
2026-06-09 15:37:30